Decision guide
Should this workflow use more agents, or one agent with better tools?
The short answer: start with one agent
For most workflows, the first thing to try is one agent with better tools, clearer instructions and a tighter output contract, not more agents. Adding agents is justified when the work splits into sub-tasks that are genuinely separable, when different parts need different tool permissions, or when independent pieces can run in parallel and wall-clock time matters. Even then, the split earns its place only if cost and latency per accepted result improve on your own tasks. Skelf has not yet measured this trade-off, so treat everything below as a framework for deciding, not a published result.
Deciding where to draw agent boundaries
The question is rarely “agents or no agents”. It is whether each boundary you draw between agents removes more difficulty than it adds in coordination. Work through the signals below for the workflow you actually have.
| Signal in your workflow | Points towards one agent | Points towards several |
|---|---|---|
| Shape of the work | Steps share most of their context; later steps need the reasoning behind earlier ones | Sub-tasks have narrow, well-defined inputs and outputs that can be written down as a contract |
| Tool access | One permission set is acceptable for every step | Some steps must not see certain data or must not hold write access (separation of privileges) |
| Parallelism | Steps are sequential and each depends on the last | Independent sub-tasks can run at the same time and latency is a hard constraint |
| Context size | The whole task fits comfortably in one context window | One context would be dominated by material irrelevant to most steps |
| Failure cost | A wrong result is cheap to detect downstream | A wrong result is expensive, and an independent check with different information is available |
| Where it fails today | Failures are tool errors, missing information or vague instructions | Failures are clearly one sub-skill done badly while the others are fine |
Two patterns deserve particular suspicion. The first is the “planner plus executor” split applied to a task whose plan is obvious: the planner adds a call and a handoff without adding information. The second is a verifier agent running the same model, with the same context, as the worker it checks. It tends to agree with the worker’s mistakes because it shares the reasoning that produced them.
What the evidence says, and does not
Measured by Skelf: nothing yet. Skelf has not published a comparison of single-agent and multi-agent architectures, and no Skelf number about agent accuracy, cost or latency exists. Any figure you see attributed to us on this question is wrong.
Well-established and arithmetic. Each additional agent adds at least one model call, its input tokens and its latency to every run that passes through it; sequential agents add their latencies. If a pipeline needs every handoff to succeed and the handoffs fail independently, the chance of a clean run is the product of the per-handoff success rates, so a chain of individually reliable steps can be unreliable as a whole. Parallel branches reduce wall-clock time only for work that is truly independent.
Inferred, not measured. The following are reasons to expect a split to help or hurt, drawn from how these systems are built:
- Lost state. When one agent hands work to another, it passes a summary, not its full reasoning. Assumptions, rejected options and caveats tend to disappear at the boundary, and the next agent acts confidently on an incomplete picture.
- Compounding error. A specialist downstream cannot fix a decomposition that was wrong upstream; it can only execute it well.
- Verifier error has two sides. A false accept lets a bad result through with extra confidence attached; a false reject triggers retries that cost money and time and can push a correct answer towards a worse one.
- Specialisation helps when the boundary is real. Narrow instructions, a smaller tool set and a smaller context can make a sub-task easier, provided the interface between agents is explicit.
Unknown. How large any of these effects is on a given task family, and whether the gains from specialisation outweigh coordination overhead at today’s model capabilities. This is the experiment that would settle it, and the one Skelf plans to run in public: three arms — a single agent, a planner with one executor, and a planner with specialists and a verifier — on a public tool-use benchmark, with the same model, the same tools and the same budget limits in every arm; at least three runs per task to expose variance; reported as cost and latency per accepted result, alongside a failure taxonomy that says where each architecture lost.
Where contracts between agents matter, Skelf’s MPL validates payloads against versioned semantic-type contracts and writes BLAKE3-hashed provenance records for each hop. Those records are evidence primitives, not a quality guarantee (claim), but they make “what did agent B actually receive?” an answerable question.
When this framework stops applying
This framework assumes a workflow that one capable model could plausibly complete with the right tools. It weakens when:
- The task exceeds one context window even after retrieval and summarisation, so decomposition is forced rather than chosen.
- Organisational boundaries are real. If different teams own different tools, data or approval steps, separate agents may simply mirror the organisation, and the right question becomes how to make the handoffs auditable.
- Regulation or security mandates separation. If a component must be provably unable to read some data, a split is required regardless of accuracy.
- Models change. A conclusion reached with one model generation may reverse with the next, in either direction, because stronger models reduce the need for decomposition while cheaper models make extra calls affordable. Re-run the comparison when the model changes.
- Your acceptance criteria are vague. Without a clear definition of an accepted result, none of the comparisons above can be made, and the architecture question is premature.
How to compare architectures on your own tasks
You can run a fair comparison in-house without new infrastructure. The discipline matters more than the tooling.
- Build a task set from real work. Take 50 to 200 tasks from recent production or from the backlog the workflow is meant to serve. Include the awkward ones. Write the acceptance criterion for each task before any arm runs, and have someone other than the system’s builder confirm it.
- Freeze everything but the architecture. Same model and version, same tools, same temperature, same maximum number of steps and tokens per run. Version the prompts for each arm — Blogus pins prompt hashes in a lock file and Promptel keeps them declarative — so the arms can be reproduced later and you know exactly what changed.
- Run each arm at least three times per task. Agent runs are noisy. Report how often each task passes across runs, not just the mean, because a task that passes one time in three is not a solved task.
- Log every call. Model calls, tool calls, retries, tokens in and out, wall-clock per step. Compute cost per accepted result as total cost across all runs divided by the number of accepted results, and report p50 and p95 latency per accepted result.
- Classify every failure. Use a fixed taxonomy: wrong decomposition, information lost at a handoff, tool misuse, hallucinated tool output, loop or timeout, verifier false accept, verifier false reject. Count by arm. This table usually tells you more than the headline accuracy.
- Audit the verifier separately. Take a sample of verifier decisions, label them by hand, and estimate its false-accept and false-reject rates. If the verifier is barely better than chance on the cases that matter, it is adding cost, not safety.
- Slice the results. Break results down by task type and difficulty. A multi-agent design that wins on long, separable tasks and loses on short ones suggests routing by task type rather than choosing one architecture for everything.
Decide in advance what difference would change your mind, for example “the multi-agent arm must lower cost per accepted result without raising the rate of verifier false accepts”. Writing that down before the runs stops the conclusion from being fitted to the results.
If the architecture decision is consequential
If the answer will set an architecture your team builds on for a year, or decides whether a product line ships, Skelf can run this comparison on your tasks rather than a public benchmark. We would agree the acceptance criteria and the decision threshold with you first, build the arms on your model and tools, run them repeatedly, and hand over the harness, the failure analysis and a report that says which design meets the threshold, including when the answer is “keep the single agent”. See Agent and RAG evaluation for scope and what you receive.
Related service: Agent & RAG evaluation · How we run an investigation · Evidence register
Questions buyers ask
- Are multi-agent systems more accurate than a single agent?
- Not as a rule. Splitting work across agents can help when the sub-tasks are genuinely separable, but every handoff is a place where context is lost and errors compound. Whether the split wins on your workflow is an empirical question, and Skelf has not published a measurement of it.
- What should we measure when comparing agent architectures?
- Cost and latency per accepted result, not per call or per run. Count every model call, tool call and retry in each arm, judge acceptance against criteria written before the runs, and classify every failure so you can see where each architecture loses.
- Does adding a verifier agent make the output safer?
- Only if its false-accept rate is low enough to matter and its false-reject rate does not destroy throughput. A verifier built on the same model as the worker often shares its blind spots, so measure the verifier against human labels before relying on it.
- When is a multi-agent design clearly justified?
- When parts of the work need different tool permissions, for example one component may read customer data and another may only write drafts, or when independent sub-tasks can run in parallel and wall-clock time matters. Separation of privileges is a valid reason even if accuracy does not improve.
- Has Skelf benchmarked single versus multi-agent workflows?
- No. A public experiment is planned: a single agent, a planner with an executor, and a planner with specialists and a verifier, run on a public tool-use benchmark with the same model and tools, at least three runs per task, reporting cost per accepted result and a failure taxonomy.
Is this decision on your desk?
If this decision is consequential, we can measure it on your workload: agent & rag evaluation, scoped to your deadline and the evidence that would settle it.