MIRKO

LLM Evaluation Platform — Eval-Driven Development for a FinTech Assistant

An internal evaluation and observability platform for a fintech's customer assistant: regression suites, LLM-as-judge scoring, production tracing and release gates that let the team ship model changes weekly instead of quarterly.

Industry
FinTech
Country
United Kingdom
Team Size
4
Services
AI Integration, Machine Learning, QA & Testing
Technologies
Python, LangGraph, OpenAI, Claude, PostgreSQL, ClickHouse, Grafana, GCP
LLM Evaluation Platform — Eval-Driven Development for a FinTech Assistant

Overview

About The Project

A fintech company ran a customer assistant that answered questions about accounts, cards and payments. Every change to a prompt or model was followed by a week of manual review, so improvements shipped quarterly and regressions were found by customers.

We built an evaluation platform around the assistant: datasets of real conversations with expected outcomes, automated scoring with rubric-based LLM judges, tracing of every production conversation and release gates in CI.

The Challenge

Regulated Answers,No Safety Net

In financial services a wrong answer about fees or limits is a compliance issue. The team needed to prove, for every release, that the assistant still refused what it should refuse and answered what it should answer.

Manual review did not scale. Reviewers disagreed with each other, and the backlog meant most conversations were never checked at all.

Regulated Answers,
No Safety Net

Our Solution

Datasets, JudgesAnd Release Gates

We curated 2,300 conversations into evaluation datasets by intent: balances, disputes, limits, fraud, refusals and small talk. Each has expected behaviour and a rubric. LLM judges score accuracy, policy compliance and tone; a sample is double-checked by humans to calibrate the judges.

Every pull request runs the suite and blocks the release if any intent drops below its threshold. Production conversations are traced with token cost and latency, and weekly sampling adds new cases to the datasets, the loop we describe in our Eval-Driven Development video.

Datasets, Judges
And Release Gates

Key Features

What WeDelivered

The Impact

Results ThatMatter

Changes became safe to ship, so the assistant improved every week.

Weekly

Release cadence (from quarterly)

0

Compliance regressions in production

-34%

Token cost after routing changes

2,300

Evaluation conversations