Agents · Orchestration

Multi-agent Handoff

an orchestrator, workers with their own contexts

model
Qwen3-1.7B every agent
workers
5
tokens processed
4,189 vs 887 for one agent

All steps

  1. One agent, one context

    First, one agent with a lookup tool answers the request alone. It looks up both models at once, keeps both fact sheets in one context (609 tokens by the end) and answers.

    1 agent · 2 calls · 2 lookups

  2. Handing out the work

    An orchestrator gets the same request and one tool, ask_worker. It hands off 5 pieces, each a model and a question. A handoff is just a tool call whose result is another agent’s work.

    5 handoffs, one tool call each

  3. Each worker starts empty

    A worker starts from an empty context: its own short instructions and one task. It never sees the user’s request, the other workers or the plan. Its first reply must be a tool call (tool_choice “required”), so it looks the model up before it answers.

    a fresh context per worker

  4. Only the reports come back

    The fact sheets stay in the workers’ contexts. The orchestrator reads only the 5 short reports, as tool results, and writes the answer from them.

    5 reports in · 1 answer out

  5. Who runs when

    The workers do not depend on each other, so a server can run them at the same time; the orchestrator waits for all of them. On a task this small the extra calls outweigh the parallelism: the one agent still finishes first.

    fan out, run in parallel, gather

  6. What it costs

    Every worker re-reads its own instructions and fact sheet, so the split processes 4,189 tokens against 887, and here the single agent’s answer was also more complete. Splitting pays when the pieces are big, independent and would crowd one context.

    4,189 vs 887 tokens processed

Code

def ask_worker(model, question):                       # the orchestrator's one tool
    worker = Agent(instructions=WORKER, tools=[lookup])     # a fresh, empty context
    return worker.run(f'Model: {model}\nQuestion: {question}')   # only the report comes back

orchestrator = Agent(instructions=ORCHESTRATOR, tools=[ask_worker])
answer = orchestrator.run(request)       # calls ask_worker once per piece, then answers from the reports

Go deeper