MIRKO
AI10 Min Read

Orchestrating AI Agents in Production: Graphs, Queues and the Human in the Loop

One agent is a demo. Ten agents working on a thousand tasks a day is an operations problem: routing, retries, budgets, approvals and visibility. Here is how we structure orchestration so it scales and stays controllable.

David Kukharchuk

By David Kukharchuk

Tech Lead at Mirko

Orchestrating AI Agents in Production: Graphs, Queues and the Human in the Loop

The question we hear most from engineering leads is "LangGraph or a queue?". The honest answer is both, for different jobs. A graph describes how one task moves through reasoning steps, tools and approval points. A queue describes how thousands of such tasks move through a system with limited budget, flaky providers and people who go home at six. Confusing the two is how pilots fail to become products.

Mirko Solutions Media9:18

AI Agent Orchestration in Production: Graphs, Queues and Human-in-the-Loop

Open on YouTube
Kanban or LangGraph? Scale, resilience and human-in-the-loop with practical models. More on our channel

Inside a task: the graph

For a single task we model the agent as a state graph. Nodes are reasoning steps and tool calls, edges are decisions, and state is explicit and serialisable. That last point matters: a serialisable state means a task can pause at an approval gate for two days, resume after a provider outage or be replayed in the evaluation suite with the exact same inputs. LangGraph gives us this structure with checkpoints out of the box.

Manager-worker and the ledger

Long tasks need delegation. In the Procurement Copilot a manager agent breaks an RFQ into extraction, matching and analysis work and hands each to a worker with a narrow tool set. The manager tracks progress in a ledger: what was asked, what came back, what is still open. The ledger is what prevents the classic failure of a single agent forgetting what it already did on step thirty, and it is what makes a long task resumable and auditable.

Across tasks: the queue

Between tasks we use plain, boring infrastructure: a queue with priorities, worker pools per model provider, rate limiting, retries with backoff and a dead-letter queue a person actually watches. Budgets live here too: a feature gets a token budget per day, and when it is spent, tasks degrade to a cheaper model or wait. This is where resilience comes from, not from the model.

Operational rules that saved us more than once

  • Every task has an idempotency key; a retried task never sends a second email or creates a second ticket.
  • Provider outages are expected; routing falls back to a second model with the same evaluation threshold.
  • Every state-changing action is logged with the evidence the agent used, before it is executed.
  • Approvals have deadlines; an unanswered gate escalates instead of silently blocking the queue.

Where the human sits

Human-in-the-loop is not a checkbox. It is a design choice about which decisions are worth a person's attention. We place gates at irreversible actions, at low-confidence results and at anything with regulatory weight. The gate shows the person the agent's evidence and proposed action, with approve, edit and reject options, and every decision flows back into the evaluation dataset. Over months the data shows which gates can be relaxed, which is how autonomy grows safely.

You May Also Like

Let's BuildSomething ThatMatters

Have a project in mind or looking for the right technology partner? Tell us what you're working on, and our team will get back to you to explore how we can help bring it to life.