Intent Datasets
2,300 curated conversations with expected outcomes and rubrics.

An internal evaluation and observability platform for a fintech's customer assistant: regression suites, LLM-as-judge scoring, production tracing and release gates that let the team ship model changes weekly instead of quarterly.

Overview
A fintech company ran a customer assistant that answered questions about accounts, cards and payments. Every change to a prompt or model was followed by a week of manual review, so improvements shipped quarterly and regressions were found by customers.
We built an evaluation platform around the assistant: datasets of real conversations with expected outcomes, automated scoring with rubric-based LLM judges, tracing of every production conversation and release gates in CI.
The Challenge
In financial services a wrong answer about fees or limits is a compliance issue. The team needed to prove, for every release, that the assistant still refused what it should refuse and answered what it should answer.
Manual review did not scale. Reviewers disagreed with each other, and the backlog meant most conversations were never checked at all.

Our Solution
We curated 2,300 conversations into evaluation datasets by intent: balances, disputes, limits, fraud, refusals and small talk. Each has expected behaviour and a rubric. LLM judges score accuracy, policy compliance and tone; a sample is double-checked by humans to calibrate the judges.
Every pull request runs the suite and blocks the release if any intent drops below its threshold. Production conversations are traced with token cost and latency, and weekly sampling adds new cases to the datasets, the loop we describe in our Eval-Driven Development video.

Key Features
2,300 curated conversations with expected outcomes and rubrics.
Calibrated judges for accuracy, compliance and tone.
Thresholds per intent block regressions before deployment.
Every conversation traced with cost, latency and retrieval hits.
Token spend per intent and per model with alerts.
Weekly review sample keeps judges aligned with compliance.
The Impact
Changes became safe to ship, so the assistant improved every week.
Weekly
Release cadence (from quarterly)
0
Compliance regressions in production
-34%
Token cost after routing changes
2,300
Evaluation conversations