AI system design · Orchestration, loops, cost per step
How to design a multi-agent AI system
An agent system completes a task by running a language model in a loop: decide the next step, call a tool, read the result, repeat. A multi-agent system splits the work between an orchestrator and specialised agents - for example research, coding and review. The shape is familiar from classic system design - a queue, a coordinator, workers and shared state - but the failure modes are new: loops that never end, context that grows every step, and cost that scales with every token.
Updated · 5 min read
Requirements
- Functional: accept a task (for example "research this company and draft a summary"), complete it using tools such as web search, code execution and internal APIs, and return the result with a record of the steps taken.
- Non-functional: bounded cost and time per task; tasks survive crashes and restarts; side-effecting actions (sending email, writing data) are safe to retry and can require human approval.
- Non-functional: every step is observable, so failures and costs can be traced.
Capacity estimates
| Quantity | Assumption | Result |
|---|---|---|
| Steps per task | 10 model calls | - |
| Tokens per step | 3,000 input + 300 output | ≈ 33,000 tokens per task |
| Task load | 1,000 tasks per minute | ≈ 17 tasks/s ≈ 550,000 tokens/s |
| Latency | ≈ 2 s per step, run sequentially | ≈ 20 s per task |
Two things stand out. Latency is a sum of sequential steps, so parallelism matters more than raw speed. And if every step re-sends the whole conversation so far, input tokens grow with each step, making total tokens per task grow roughly with the square of the step count.
High-level design
- API → task queue: tasks are accepted immediately and processed asynchronously; the client polls or receives events.
- Orchestrator: plans the task, splits it into sub-tasks, dispatches them to agents and combines the results.
- Agent workers: each runs a model loop for one sub-task with its own instructions and allowed tools.
- Model gateway: one place for provider routing, retries with backoff, rate limits, quotas, caching and cost accounting.
- Tool gateway: timeouts, sandboxing for code execution, credentials, and idempotency keys for side effects.
- State store: task state, step history and intermediate results, written after every step so a task can resume after a crash.
Durable execution
A twenty-second task with ten external calls will regularly hit a timeout, a deploy or a crashed worker. Treat each task as a workflow: persist the state after every step and make every step safe to re-run. Workflow engines such as Temporal, or a simple state machine over a database and a queue, let a task resume where it stopped instead of starting again and paying for every token twice.
Where it breaks
- Loops: an agent keeps calling the same tool or re-planning without making progress, burning tokens indefinitely.
- Context growth: step history accumulates until it exceeds the context window or makes each call slow and expensive.
- Retry storms: a failing tool or provider triggers retries at every layer, multiplying load on the thing that is already failing.
- Fan-out explosions: an orchestrator that spawns agents that spawn agents can multiply cost far beyond what the task is worth.
- Unsafe side effects: a retried step sends the same email twice or writes the same record twice.
Designing for control
- Budgets on every task: a maximum number of steps, a token budget and a wall-clock deadline, enforced by the orchestrator, not by the model.
- Summarise or trim history between steps instead of re-sending everything, and keep large tool outputs in the state store with only references in the prompt.
- Run independent sub-tasks in parallel to cut end-to-end latency.
- Retry in one place (the gateway) with backoff and a cap, and use circuit breakers for failing tools.
- Use idempotency keys for side-effecting tools, and require human approval for irreversible actions.
- Cache deterministic tool results and repeated model calls, and route simple steps to smaller, cheaper models.
Observability and evaluation
Trace every step with its inputs, outputs, tokens, latency and cost, grouped by task. Track task success rate, cost per successful task and the distribution of steps per task - a long tail of step counts usually means loops. Keep a set of reference tasks and re-run it on every change to prompts, tools or models.
What interviewers look for
- You estimated tokens, latency and cost per task, and noticed that tokens grow with step count.
- You made execution asynchronous and durable.
- You put gateways in front of models and tools for retries, limits and accounting.
- You bounded loops and cost with budgets the orchestrator enforces.
- You made side effects idempotent and observable.
Frequently asked questions
What is an AI agent system design interview?
+
It asks you to architect a system where language-model agents complete multi-step tasks with tools - for example a research assistant or coding agent. You design the orchestrator, workers, tool and model access, state and observability, and reason about tokens, latency, cost per task and runaway loops.
How is AI system design different from classic system design?
+
The components are familiar - queues, coordinators, workers, caches and storage - but the units change. Throughput is measured in tokens, memory is the context window, cost is per token, and new failure modes such as agent loops and context growth sit alongside the classic ones.
How do you control cost in a multi-agent system?
+
Give every task a step limit, a token budget and a deadline that the orchestrator enforces, trim history between steps, cache repeated calls, route simple steps to cheaper models and run independent steps in parallel.
What does the orchestrator do?
+
It receives a task, plans it, splits it into sub-tasks, dispatches them to specialised agents, enforces budgets and combines the results. It plays the role a coordinator plays in classic distributed systems, plus managing context and control flow for the agents.
How do you make agent tasks reliable?
+
Persist task state after every step so a task can resume after a crash, make every step safe to retry, use idempotency keys for side effects, retry with backoff in one gateway, and set timeouts on every tool call.