Skip to content
    AI Orchestration

    Design Multi-Agent Workflows With Six-Field Handoffs

    JK
    7 min read

    TL;DR

    1

    2

    3

    4

    5

    Start small. Type every handoff. Put checks where mistakes cost the most. That is most of the job.

    Teams tend to over-build first and pay for it later. So begin with the core parts: one orchestrator, a registry, a state store and basic logging. That set comes straight from the Microsoft reference architecture. You can add more once the simple version works.

    Four rules to design multi-agent workflows that hold up

    Good multi-agent design comes down to four rules. Break one and the system gets slow, costly or flaky.

    • Simple first. Build the smallest system that solves the problem. Add an agent only when one agent with tools can't cope.
    • Feedback where it pays. Add a review loop only where it buys real accuracy.
    • One main path. Keep a stable sequence. Add retries or checks at chosen points, not everywhere.
    • Pass summaries. Agents hand each other short, structured notes, not full chat logs.

    That last rule matters more than it looks. A study on scaling LLM multi-agent systems lists these four as its design principles. It pairs summary-based handoffs with a fixed map of allowed paths. Short notes keep context small. A fixed map stops agents wandering off.

    Pro tip: treat every extra agent as a cost decision, not a feature.

    The parts every system needs

    Think of this as the kitchen your agents work in. Miss one part and something breaks later.

    • Orchestrator and registry. The orchestrator routes the work. The registry lists which agents exist, what they do and which tools they can touch.
    • Knowledge layer. Where agents get facts and documents. Keep it apart from the tools they act with.
    • Tool layer. The actions agents can take, often through MCP (Model Context Protocol, a standard way for agents to call tools). Adapters should only pass calls through. They should not make decisions.
    • State store. One shared record per task: ID, status, owner agent and output.
    • Logs and governance. Health checks and audit trails, so you can see what ran, what failed and who approved it.

    Skip one of these and it tends to show up later as a failure nobody can explain. For the bigger picture, see how multi-agent systems work. Devwiz also has a plain-English primer on multi-agent systems.

    Which pattern to pick

    Pick the shape that fits the task. Not the one that looks clever on a diagram.

    PatternUse it whenTrade-off
    PipelineEach step needs the last step's outputCheap and easy to debug, but rigid
    Supervisor-workerOne agent must control several specialistsSlower, but consistent
    Fan-out, fan-inIndependent checks can run side by sideFast, but merging conflicts is hard
    Mesh or group chatAgents truly need to talk freelyFlexible, but hard to test and costly
    ManagerA lead agent assigns each new itemGood for mixed work, harder to predict

    Azure's agent design pattern guide says to use the lowest-complexity pattern that meets the need. Mesh and manager patterns look powerful. They often cost more in tokens and debugging than they save. See stopping runaway agent costs for how that cost climbs.

    How to create agent workflows, step by step

    Build in this order. Skip a step and you end up debugging in production.

    1. Define the outcome. Write down what success looks like and how you will measure it. Do this before any code.
    2. Map the process. Break the job into clear steps. Each step is a possible agent or tool call.
    3. Write the handoff contracts. Every pass between agents gets a typed contract. More on that below.
    4. Wire it up. Connect the orchestrator, tools and any agent-to-agent links. Add retries and approval points.
    5. Test in tiers. Test each agent alone. Then test the handoffs. Then run staging and shadow tests on real traffic.

    This is the same Explore, Map, Transform order we run in the 90-day build, just laid out as engineering steps.

    Pro tip: never skip staging. A workflow can pass every unit test and still fail the first time two agents disagree on a live handoff.

    The six-field handoff contract

    A handoff with no contract is a guess. Give every pass between agents six fields.

    FieldWhat it says
    Input schemaWhat the agent receives, and in what format
    Output schemaWhat it must return, as structured data, not free text
    Allowed toolsThe exact tools it may call, and nothing more
    Failure statesWhat counts as a failure and how it gets flagged
    Stop rulesThe condition that ends the agent's turn
    Escalation pathWho or what takes over when the agent is stuck

    Give each agent one job. Prefer tools that act like pure functions: the same input always gives the same output. Keep prompts in their own files, not buried in code. And make each agent say which agent runs next. Our guide to agent handoffs shows how this stops two confused agents passing work back and forth forever.

    Build it as an AI Operating System, not a pile of bots

    Here's the thing. The contract is the easy bit to write down. The hard bit is knowing what goes in it.

    The stop rules, the escalation paths and the "good enough" test all come from how the founder makes calls today. That is the IP. When you write it into contracts, your agents stop being generic bots. They become AI employees that make the calls you would make.

    We build these with Claude Code. Each agent is a folder of plain files: its role, its rules, its tools and its handoffs. You can read every one. You can change every one. That is how a non-technical founder ends up owning the system rather than renting it. See custom AI delivery systems built with Claude Code for how that works in practice.

    What to measure and test

    Measure two layers: each agent, and each full run.

    • Per agent. Success rate, accuracy on rule or fact checks, tool failures, delay and token use.
    • Per run. Handoff errors, retries, times a human stepped in, and end-to-end success.
    • Tests. Unit test each agent. Test the handoffs. Then run load tests and fault injection, which means feeding in bad data on purpose to see what breaks.

    Checks make mistakes cheap to catch. A Stanford researcher's write-up of a correctness-gated multi-agent workflow makes the point well. The better question is not which model is best. It is which process makes a model's mistakes cheap to spot. His answer was cross-agent review plus structural checks. It is one person's setup, not a guarantee. The lesson still holds: catch errors early. For the monitoring side, see stopping hidden token burn.

    Where to put the gates

    Gates belong where a mistake is expensive. Not at every step.

    • Schema checks. Reject bad input or output before the next agent sees it.
    • Automated tests. Run rule and fact checks against known good answers.
    • Cross-agent review. One agent checks another's work before it moves on.
    • Outside checks. Use an external tool or formal check for high-stakes calls.

    Add human approval at set pause points, not all the time. People stay in the loop without becoming the new bottleneck. At the system level, keep a registry, clear metadata and audit logs, so every action can be traced later.

    Common mistakes

    The same mistakes turn up again and again.

    • Adding agents before trying one agent with more tools.
    • Writing handoff contracts after the bugs appear, not before.
    • Adding logs late, once something has already failed quietly.

    A fixed build order stops all three, because the contracts and tests come before the wiring. For a worked example from another field, see this orchestration-first content planning workflow.

    When multi-agent is worth it

    Most tasks don't need many agents. One agent with good tools handles most jobs fine. Go multi-agent when the work splits into separate roles with different tools or knowledge. If you can't say why one agent isn't enough, don't build five.

    Next step

    Designing a workflow on paper is one thing. Wiring it, testing it and handing it to your team to run without you is another. That is the gap the 90-day program closes.

    • A working prototype built on your real processes, not a template.
    • A fully wired workflow, with handoffs, gates and logs in place.
    • Full ownership of the system when we finish. No lock-in.
    • Platform Access to run, watch and adjust your agents afterwards, from $90 a month or $900 a year.

    Want to know if your business is ready for this? Take the assessment. For more on how the build runs, see the services page or AI consulting for coaches and consultants.

    Sources

    Frequently Asked Questions

    JK

    James Killick

    Founder

    The AI Orchestrator. 10+ years building digital products and 200+ apps shipped, now helping $1M+ educators and consultants turn their IP into AI-powered delivery systems.

    James Killick founded and runs The AI Orchestrators.

    Ready to find out where your biggest AI opportunity is?

    Take the assessment. It takes about 5 minutes. You'll get a clear picture of how ready your business is.