MIRKO
AI9 Min Read

Context Engineering: Deciding What Your Model Sees, and When

Prompt engineering is about wording. Context engineering is about everything the model receives: retrieved documents, tool results, memory, instructions and their order. In production it decides cost, latency and accuracy more than the choice of model does.

David Kukharchuk

By David Kukharchuk

Tech Lead at Mirko

Context Engineering: Deciding What Your Model Sees, and When

Large context windows created a comfortable illusion: if the model can read 200,000 tokens, just give it everything. In practice, attention is not free. Models answer worse when the relevant passage is buried in noise, every token costs money on every call, and latency grows with input size. Context engineering is the discipline of deciding what goes into the window, in what order, and what stays out.

Mirko Solutions Media8:45

Context Engineering: Managing AI Attention in Production

Open on YouTube
Our practical guide to managing AI attention in production systems. More on our channel

The budget mindset

We treat the context window as a budget with line items: system instructions, retrieved knowledge, conversation history, tool results and the output reserve. For a customer assistant that might be 1,500 tokens of instructions, up to 6,000 of retrieved passages, the last six turns of conversation and whatever tools returned. Everything else is cut or summarised. Writing the budget down makes trade-offs explicit and stops the slow creep that doubles cost over a year.

Retrieval beats recall

Fine-tuning a model on company documents is rarely the answer to "the model does not know our data". Retrieval is: index the documents, fetch the three passages that matter for this question and put them in the context with their sources. Retrieval keeps answers current, lets you show citations and respects access rights, because you only retrieve what the user may see. In the Pharma QA Knowledge Copilot we built, version-aware retrieval was the entire difference between a toy and a tool auditors accept.

Retrieval details that matter in production

  • Chunk by document structure, not by a fixed token count, so a passage keeps its heading and table.
  • Store metadata (product, site, version, effective date) and filter on it before semantic search.
  • Return citations with every passage; the model must be able to say where a fact came from.
  • Measure retrieval separately from generation. Most wrong answers are retrieval failures.

Caching and ordering

Hosted models now cache the stable prefix of a prompt. If instructions and reference material come first and the volatile part last, repeated calls cost a fraction and respond faster. We order context as: system instructions, tool definitions, stable reference material, retrieved passages, conversation, user message. The LLM Evaluation Platform we built for a fintech cut token cost by 34% mostly through routing and ordering, with no change in answer quality.

Memory is context too

Conversation history is the most abused part of the window. Keeping the full transcript seems safe and is actually harmful: old turns compete with the current question for attention. We summarise history into structured state after a few turns, keep the last exchanges verbatim and store durable facts in a memory the user can see. The agent-harness article on this blog describes how that memory is scoped.

When to fine-tune after all

Fine-tuning earns its place for format and style, for narrow classification at very high volume and for small models that must run on-device. It does not replace retrieval for knowledge, and it does not replace scaffolding for multi-step tasks. Our decision matrix video goes through the choice case by case.

Mirko Solutions Media8:13

Fine-Tuning vs. RAG vs. Scaffolding: The AI Architectural Decision Matrix

Open on YouTube
A technical comparison of fine-tuning, Retrieval-Augmented Generation (RAG) and AI scaffolding. Explains how to choose the right architecture based on total cost of ownership, latency, adaptability and performance. More on our channel

You May Also Like

Let's BuildSomething ThatMatters

Have a project in mind or looking for the right technology partner? Tell us what you're working on, and our team will get back to you to explore how we can help bring it to life.