The architecture diagram had nine agents on it. A planner, a researcher, three specialists, a critic, a synthesiser, a validator and a supervisor coordinating the lot. It was genuinely elegant. In production it was slower than a single well-prompted model call, cost roughly eleven times as much, and failed in ways nobody could reproduce — because by the time an error surfaced it had passed through four agents, each of which had paraphrased the previous one's output.
The rebuild had two agents and a deterministic state machine between them. It was faster, cheaper, and when it failed you could tell which step failed and why.
This is the most common mistake in agent architecture: treating agents as the unit of decomposition when they should be the exception. Every agent boundary you add is a place where structured state becomes natural language and back again, and each of those conversions is lossy, slow and expensive.
The cost of a boundary
It is worth being concrete about what an agent handoff actually costs, because the diagram makes it look free.
When agent A passes work to agent B, three things happen. The state is serialised into text. Agent B receives that text plus its own system prompt, its own tool definitions, and whatever context it needs to be useful — typically several thousand tokens before it has done anything. Then B re-derives an understanding of the situation that A already had.
Add latency: another model round trip, often several seconds. Add cost: the context is re-sent, so a chain of five agents can easily send the same background information five times. Add the failure surface: B may misread A's summary, and there is no type system between them to catch it.
A function call costs microseconds and cannot misunderstand its arguments. An agent handoff costs seconds, dollars and a paraphrase. Use the second one only when you need judgement that the first cannot provide.
The test we apply before adding an agent: does this step require open-ended reasoning over unstructured input, or does it require a decision that a competent engineer could express as code? If it is the latter — and it usually is — it belongs in the orchestration layer, not in a model.
Start with a deterministic pipeline
Most workloads that get described as multi-agent are actually a fixed sequence of steps with one or two genuinely uncertain decisions in the middle. Extract fields from a document, validate them against a schema, look up the counterparty, decide whether it needs review, write the result. Only the extraction and possibly the decision need a model. Everything else is code.
The pattern that works is a deterministic pipeline with model calls at specific stages, not a conversation between autonomous participants. The control flow lives in your language, where you can test it, log it, retry it and reason about it:
type Stage<I, O> = {
name: string;
run: (input: I, ctx: RunContext) => Promise<O>;
/** Deterministic gate: does this output satisfy the contract? */
validate: (output: O) => Result<O, ValidationError>;
retries: number;
};
async function runPipeline<T>(stages: Stage<unknown, unknown>[], input: T, ctx: RunContext) {
let current: unknown = input;
for (const stage of stages) {
let attempt = 0;
for (;;) {
const output = await stage.run(current, ctx);
const checked = stage.validate(output);
if (checked.ok) {
ctx.audit(stage.name, { attempt, output: checked.value });
current = checked.value;
break;
}
// Feed the validation failure back as context, do not just retry blind.
ctx.audit(stage.name, { attempt, error: checked.error });
if (++attempt > stage.retries) throw new StageFailed(stage.name, checked.error);
current = withRepairHint(current, checked.error);
}
}
return current;
}
Two things in there matter more than the structure. Every stage has a deterministic validator, so a model's output is checked by code rather than by another model. And a failed validation is fed back as a repair hint rather than triggering a blind retry — retrying an identical prompt against a non-deterministic model is a lottery, but telling it precisely what was wrong with its last attempt usually succeeds on the second try.
Four patterns that earn their keep
Router
One cheap, fast classification call decides which specialised path handles the request. The router does not do the work; it picks the handler. This is the highest-value multi-agent pattern because the routing decision is small, the specialised handlers can have tight focused prompts instead of one enormous prompt covering every case, and you can use a small model for the routing and reserve the expensive model for the work.
The failure mode is router misclassification, which is silent and cascades — the wrong specialist answers confidently. Mitigate by making the router return a confidence and a second choice, and escalating ambiguous cases rather than guessing.
Parallel fan-out with deterministic merge
Where subtasks are genuinely independent — analyse twelve documents, check a change against six policy categories — run them concurrently and merge the results in code. Latency becomes the slowest branch instead of the sum, which is often a five- or ten-fold improvement in wall-clock time.
The critical detail is that the merge should be deterministic. The instinct is to add a synthesiser agent to combine the outputs. That reintroduces a serial model call over a large context, and the synthesiser frequently drops findings from the middle of its input. If the merge is "collect all findings, deduplicate, sort by severity", write that in code.
Generate then verify
Two roles with genuinely different objectives: one produces a candidate, the other checks it against criteria. This works because the verifier's job is narrower and more objective than the generator's, and because a fresh context is better at spotting a flaw than the context that produced it.
It works considerably better when the verifier has tools that produce ground truth — running the test suite, executing the query, calling the schema validator — rather than forming an opinion. A verifier that only reasons agrees with the generator far more often than it should. This is the pattern behind the SDE and QA agents in our agent platform sharing one backlog: generation without independent, executable verification just produces confident output faster.
Supervisor with bounded delegation
The pattern people reach for first and should reach for last. A coordinating agent decides which specialist to invoke, reads the result and decides what to do next. It is the right choice when the sequence of steps genuinely cannot be known in advance — open-ended investigation, debugging, research where each finding determines the next question.
It needs hard bounds or it will not terminate. A maximum step count, a token budget, a wall-clock deadline, and a rule that the same subtask cannot be delegated twice. Without those, the characteristic failure is two agents politely handing a task back and forth while the meter runs.
State is the hard part, not coordination
Agent frameworks spend their documentation on how agents talk to each other. In production the difficulty is almost entirely about state: what is the source of truth, who may write to it, and what happens when a run dies at step four of seven.
Passing state as conversation history — the default in most frameworks — is the root of several problems. Context grows with every turn until you are paying to re-send the entire history on each call, and eventually truncating it, which means the system silently forgets its earliest and often most important instructions. It is also unqueryable: you cannot ask "what did the extraction stage decide" without parsing prose.
Keep a typed state object as the source of truth. Agents receive a projection of it — only the fields their step needs — and return structured output that is validated and merged back by the orchestrator. Conversation history becomes a debugging artefact rather than the data model.
This buys you the operational properties that matter:
- Resumability. Persist state after each stage and a failed run restarts from the last good checkpoint instead of from the beginning. On a workflow with six model calls, this is the difference between a retry costing one call and costing six.
- Idempotency. Side-effecting steps need a stable key derived from the run and the stage, so a retry after a timeout cannot send the same message twice. Assume every step will execute more than once, because under retries it will.
- Inspectability. When something goes wrong, you need to see the state at each boundary, not reconstruct it from a transcript.
- Partial failure handling. In a fan-out, decide explicitly what a run means when two of twelve branches fail. Returning ten results as if they were twelve is the quiet failure that damages trust.
Cost and latency compound
The arithmetic here is unforgiving and worth doing before you build, not after.
A single call with 4k tokens of context is one round trip. A five-agent chain where each agent carries its own 2k system prompt plus a growing shared context is five round trips and materially more than five times the tokens, because the shared context is re-sent each time. Sequential latency adds: five calls at three seconds each is fifteen seconds before any output reaches the user.
| Structure | Model calls | Latency | Relative token cost |
|---|---|---|---|
| Single call | 1 | 1 round trip | Baseline |
| Router plus one specialist | 2 | 2 round trips | Often below baseline — small router, tighter specialist prompt |
| Fan-out of 6, code merge | 6 | 1 round trip (slowest branch) | ~6x, bought with a large latency win |
| Sequential chain of 5 | 5 | 5 round trips | >5x from re-sent context |
| Supervisor, unbounded | Unbounded | Unbounded | Unbounded |
The router row is the interesting one: adding an agent can reduce total cost, because a focused specialist prompt is much shorter than a monolithic prompt that has to handle every case. That is the shape of a good boundary — one that reduces work downstream. The sequential chain is the shape of a bad one. Per-stage model selection and prompt-cache-friendly context ordering matter a great deal here, and we cover both in our note on controlling LLM cost and latency.
Debugging a system that is different every time
Non-determinism makes conventional debugging useless. You cannot reproduce the failure by rerunning it, and a stack trace tells you nothing about why a model chose what it chose.
What works is treating every run as a distributed trace. A run identifier threaded through every stage; for each stage the exact prompt sent, the raw response, the parsed output, the validation result, latency, token counts, and the model and prompt version used. Store it. When someone reports that the system did something strange last Tuesday, this is the only thing that will answer them.
Then add replay: the ability to take a recorded run and re-execute it against a modified pipeline, diffing the decisions. This is how you ship a prompt change with any confidence, and it is the piece teams most often skip and most often regret. Track stage-level metrics too — validation failure rate and retry rate per stage will point at your weak link long before users complain.
What this means in practice
Start with one model call and a lot of code around it. Add a second agent only when you can articulate what judgement it contributes that code cannot, and what it costs in latency and tokens. The burden of proof sits with the new boundary, not against it. Most systems that end up working well have two or three model calls in them, not nine.
Put the control flow in your programming language and the judgement in the model. Validate every model output with deterministic code and feed failures back as repair hints. Keep typed state rather than conversation history, checkpoint it between stages, and make every side effect idempotent. Bound anything that loops with step, token and time limits, and decide up front what a partially successful run returns.
Orchestration is a distributed systems problem with a non-deterministic component in it, and the discipline that makes it survivable is the same discipline that makes any distributed system survivable — clear contracts, checkpointed state, idempotent effects and real observability. The choices that matter are about where you draw the boundaries; for how to select the workloads worth orchestrating in the first place, see our note on finding the processes worth automating. If you are designing one of these and the diagram has more than three agents on it, talk it through with us before you build — that conversation is a large part of our AI engineering work.