MIRKO
AI10 Min Read

Eval-Driven Development: How We Ship LLM Features Every Week Without Breaking Them

You cannot unit-test a language model, but you can measure it. Eval-driven development puts a dataset of real tasks, automated scoring and release gates around every LLM feature, so changes become safe and improvement becomes continuous.

David Kukharchuk

By David Kukharchuk

Tech Lead at Mirko

Eval-Driven Development: How We Ship LLM Features Every Week Without Breaking Them

A fintech client came to us with an assistant that worked well in demos and badly in the wild. Every prompt change took a week of manual review, so improvements shipped quarterly and regressions were found by customers. Six months later the same team ships weekly with zero compliance regressions. Nothing about the model changed. The evaluation loop did.

Mirko Solutions Media8:40

Eval-Driven Development (EDD): Testing & Evaluating LLMs in Production

Open on YouTube
Testing and evaluating LLMs in production: datasets, judges and release gates. More on our channel

Start with the dataset, not the prompt

Eval-driven development inverts the usual order. Before tuning a prompt, collect the cases it must handle: real conversations, real documents, real edge cases, each with the expected outcome. For the fintech assistant that meant 2,300 conversations grouped by intent: balances, disputes, limits, fraud, refusals and small talk. For the Procurement Copilot it was 400 historical RFQs with the buyer's final decision. The dataset is the specification.

What a good evaluation case contains

  • The input exactly as production would send it, including the retrieved context.
  • The expected outcome, which may be an answer, a refusal, a tool call or an escalation.
  • A rubric: what makes an answer acceptable, not just a reference string.
  • Metadata: intent, difficulty, source (synthetic, production, incident) and date.

Scoring that scales

Exact-match scoring fails for free text, and human review does not scale. We use three layers: deterministic checks where possible (did the agent call the right tool with the right arguments, did it cite a source), LLM-as-judge scoring against the rubric for free text, and a weekly human sample that calibrates the judges. When judges and humans disagree, the rubric is rewritten, not the judge.

Release gates in CI

Every pull request that touches a prompt, a retrieval setting or a model version runs the full suite. Each intent has a threshold; if any drops below it the release is blocked with a diff of the failing cases. This turns "I think the new prompt is better" into a table. It also makes model upgrades boring: when a new model version lands, the suite tells you within an hour whether to switch.

Production is part of the loop

Evaluation does not stop at deployment. Every production conversation is traced with cost, latency, retrieved passages and the final answer. A weekly job samples conversations, flags low-confidence or negatively rated ones and routes them to review, where they become new evaluation cases. The dataset grows with the product instead of rotting.

What this means for your AI roadmap

Teams that adopt eval-driven development stop arguing about prompts and start shipping. It is also the only honest answer to a compliance officer asking how you know the assistant is safe. If you are integrating an LLM into a product, build the evaluation set in the first sprint. We do it on every AI Integration project, and the LLM Evaluation Platform case study shows what the platform looks like when it is the product.

You May Also Like

Let's BuildSomething ThatMatters

Have a project in mind or looking for the right technology partner? Tell us what you're working on, and our team will get back to you to explore how we can help bring it to life.