Khachatur Pepanyan ← All notes

Notes · LLM evaluation · 6 min read

Evals before prompts

On an LLM-powered credit-decisioning platform, every prompt change was effectively a production deploy. This is the evaluation setup that let us change prompts and models quickly — without silently breaking the decision engine behind them.

1,700labeled golden cases
96%field-level accuracy
−45%AI cost per case
3.8 sp95 latency kept

Why a prompt needs a release gate

The LLM features sat on top of credit-decisioning services handling 50–70k requests a day: RAG over credit policies and historical cases, structured extraction from bank statements, payslips and ID documents, and a tool-calling agent that drafted recommendations for manual reviews. Their outputs fed a decision engine, so a wrong field was not a typo — it was a wrong decision.

A prompt tweak that fixes one case can quietly degrade fifty others. Without measurement you learn about it from analysts, or worse, from the decisions themselves. So the rule became simple: no prompt or model change ships unless the evals say it is at least as good as what is running.

The dataset comes first

The foundation is a golden set of 1,700 labeled cases, topped up with sampled production traces so the set does not drift away from what users actually send. Labels live at the field level, not the document level: “income correct, employer wrong” tells you what to fix; “document 80% fine” does not.

One metric per way of failing

The gate

Every prompt or model change runs the suite in CI, and a drop below the baseline blocks the merge. Production drift monitoring covers what the golden set does not yet contain. Prompts and models are versioned, and per-call tracing with Langfuse and OpenTelemetry ties any bad answer back to the exact prompt and model that produced it — together with latency, token cost and error-rate dashboards.

The point of the gate is speed, not caution: once regressions are caught automatically, you can try a new prompt or model the same afternoon instead of debating it for a week.

Evals pay for themselves: model routing

Benchmarking candidate models on the same suite showed that simpler cases did not need the largest model. Routing them to a smaller one cut AI cost per case by 45% while keeping p95 latency at 3.8 seconds. That is a change I would never have shipped on intuition — the eval numbers are what made it safe.

Humans stay in the loop

Evals tell you how good the model is; the human loop decides what happens when it is not. Low-confidence fields go to review, and the agent only drafts recommendations — an analyst approves every one of them.

What I would tell my past self

  1. Build the eval set before polishing prompts — otherwise you are tuning against vibes.
  2. Measure at the level you act on: fields, retrieved chunks, single claims.
  3. Calibrate the judge against humans, and re-check it when the domain shifts.
  4. Make the gate block merges. A dashboard nobody has to pass is a suggestion.
  5. Let eval data pick models and routes — that is where the cost savings hide.