Skip to content
AI Inside the WorkflowShippedAI

Three stages to one plan, with the brief cached once

Stage 1 writes a methodology brief with no client data, cached by content hash. Stages 2 and 3 match and synthesize in a background job that streams a preview.

  • Supabase
  • Custom

The problem

One plan took over three minutes of sequential generation with nothing on screen. The expensive part, turning a practitioner's research into a usable brief, was redone for every client even though it contains no client data at all.

What we built

Stage 1 writes an educational brief per practitioner methodology, one parallel call per bundle, with no client information in it. Stage 2 matches each brief to the client's intake and lab summary, again in parallel. Stage 3 synthesizes the drafts into one plan, and its rule is to resolve conflicts decisively: pick one recommendation, never present alternatives, preferring the one safer alongside the client's listed medications, then the one better supported by the source, then the more practical. Single-bundle plans skip synthesis. Practitioners per plan are capped at three.

The brief is cached on the practitioner record, keyed by a content hash of the prompt version, the methodology, and its research rows, with one variant per anonymize setting. There is no expiry clock; any edit to the research or the prompt changes the hash and the brief is rewritten. Concurrent plans on a shared practitioner read-merge-write so neither clobbers the other, deterministic fallbacks are never cached, and the cached text is position-independent so it slots into any plan. A warmer pre-filled every active practitioner: 120 of 120, 0 failures.

Generation runs as a background job. The request returns a job id immediately, synthesis text streams into a preview as it is written, and stalled jobs are reaped. The per-bundle drafts stay in the version snapshot for audit, the compliance scrubber runs before anything is stored, and edits and regeneration pass through the same gate.

Where AI does the work

Stage 1 writes an educational brief per practitioner methodology from its research, with no client data. Stage 2 matches each brief to the client's intake and labs, one parallel call per bundle. Stage 3 synthesizes the drafts into one plan and resolves conflicts by choosing one recommendation, the one safer with the client's listed medications first.

Where rules do the work

The brief cache is keyed by a hash of the prompt version, the methodology, and its research rows, so any edit invalidates it and nothing expires by clock; fallbacks are never cached. A background job returns immediately and is polled, synthesis text streams to a preview, and a reaper clears stale jobs. Per-bundle drafts are kept in the version snapshot for audit, the compliance scrubber runs before storage, and edits and regeneration pass through the same gate.

Outcome

Measured before and after on the same kind of plan: 3 minutes 13 seconds to 63.5 seconds, with all three briefs served from cache and every stage passing first try. First visible text at about 35 seconds, where before there was nothing until the end. Three changes did it: a faster model tier, the brief cache, and the background job.

Who this fits

  • Products that generate long documents from a shared knowledge base plus a per-customer input
  • Teams paying for the same model call twice because their cache expires by clock
  • Anyone whose users stare at a spinner for three minutes

See it running first.

Proof first. Then we'll talk about your stack.

How we work →