Anatomy of an AI Agent Harness: What Turns a Model Into a Reliable Worker
A language model on its own is a very smart intern with no desk, no tools and no memory. The harness around it is what makes it an employee you can trust with real work. Here is what a production harness contains, piece by piece, and where teams usually underinvest.
By David Kukharchuk
Tech Lead at Mirko

When a client asks us to "add an AI agent", the model is rarely the hard part. Claude, GPT and the open-source models are all capable of reading a contract or drafting a support reply. What decides whether the agent survives contact with a real company is everything around the model: how it gets context, which tools it may call, what it remembers, who approves its actions and how its work is measured. We call that layer the harness.

Anatomy of the Agent Harness: From a Model to a Real Employee
1. Context assembly
The model sees only what the harness puts in front of it. A good harness assembles context deliberately: the task, the relevant documents retrieved for it, the user's permissions, the state of the conversation and the format of the expected output. A bad harness dumps everything into the prompt and hopes the model finds the relevant part. The difference shows up as cost, latency and wrong answers. We cover the mechanics in our article on context engineering below.
2. Tools with typed contracts
An agent without tools can only talk. Tools are how it reads your ERP, updates a ticket or queries a database. Each tool needs a typed contract: what it accepts, what it returns and what happens on failure. We expose tools through Model Context Protocol servers or function-calling schemas, and we design them as narrow verbs, such as `get_invoice` or `resend_invoice`, rather than a single `run_sql`. Narrow tools are easier to permission, easier to test and far harder to misuse.
Rules we apply to every tool
- Read-only tools are separated from tools that change state.
- Every state-changing tool declares whether it is reversible.
- Tools return structured results with an explicit error field; the model never has to parse stack traces.
- Tool access is scoped per user, not per agent, so the agent can never do more than the person who asked.
3. Memory that is scoped on purpose
Agents need short-term memory for the current task, working memory for a multi-step job and long-term memory for preferences and prior outcomes. The trap is treating all three as one growing transcript. We keep task memory in a ledger the manager agent owns, working memory in structured state and long-term memory in a store the user can inspect and clear. That separation is what lets an agent run a 40-step RFQ evaluation without forgetting step three.
4. Permissions and approval gates
No company will let an agent act without control, and it should not. The harness defines permission tiers: free actions, actions that notify, and actions that wait for a human. In our Support Resolution Agent project, resending an invoice is reversible and notifies the account owner; cancelling a subscription waits for a support engineer. The agent proposes, the person decides, and the decision becomes training data for the evaluation set.

Human-in-the-Loop & Security Sandboxing: Guardrails for Autonomous AI
5. Evaluation and observability
The last component is the one most prototypes skip. Every production harness we ship has an evaluation dataset of real tasks with expected outcomes, scoring that runs on every change, tracing of every production run with cost and latency, and a weekly loop that turns production failures into new evaluation cases. Without it the team cannot tell whether a prompt change helped or hurt, and autonomy never grows beyond the pilot.
The four classes of harness
Not every agent needs all of this at full strength. We distinguish four classes: IDE assistants that act inside a developer's editor, personal harnesses that managers set up in Claude Projects or ChatGPT, embedded copilots inside a product, and autonomous server-side agents that run without a person watching. Each class needs more of the components above than the previous one. Our video on the four classes explains how to pick the right level for a use case.

Four Classes of Agent Harness: From Cursor to Autonomous Servers





