Topic

How a trained AI evaluator drifts until it confidently certifies work its owner finds unusable — and why the confident number is what makes it dangerous.

Target Reader

A knowledge entrepreneur running an AI content or analysis pipeline with a quality gate in it. They took the good advice — train a rubric instead of tweaking prompts — and it worked. They now have an evaluator that scores their output, and they have started trusting the score more than their own read.

The Fear / Frustration / Want / Aspiration

The quiet unease of looking at something your system says is excellent and knowing it isn’t — and not being sure which of you to believe. Underneath: the fear that in automating your standards you handed them over, and can no longer tell whether the thing you built is still working for you.

Before State

Treats the evaluator’s score as evidence. When gut and rubric disagree, defers to the rubric, because a number looks more objective than a reaction. Has never audited the evaluator itself — auditing the outputs is the whole point of having one.

After State

Understands that a self-improving rubric is a drifting rubric, that drift produces no error and no warning, and that a confident score actively suppresses the doubt that would otherwise catch it. Runs a periodic calibration pass. Holds the correct hierarchy: you are ground truth, the rubric is a lossy compression of you taken at an earlier date.

Narrative Arc

Lou publishes a batch of long-form articles he’s disappointed by, then runs them through his evaluator agent out of habit. It returns 4.9 out of 5. Both are right — against the criteria that agent accumulated over many runs, the work genuinely scores 4.9. The turn is realising there was no malfunction: every criterion was reasonable when added, and the compound result is an instrument measuring something that quietly stopped being what he cared about. The resolution is a distinction he named in the article he wrote about it — did you do it right, or did you do the right thing? — and the practice that follows: the measurement instrument needs its own audit.

Core Argument

An evaluator that learns from use will drift away from what you actually value, and because it fails by producing a confident number rather than an error, it will keep certifying work you’d reject — so when your reaction and your rubric disagree, the rubric is the suspect.

Key Evidence / Examples

  • The 4.9 itself, and Lou’s reaction: “I’m looking at the quality of the writing, I’m going, this is not 4.9 out of 5.” The gap between reaction and score is the entire diagnostic.
  • The mechanism: criteria accumulated over time, each reasonable individually, compounding into an instrument nobody designed.
  • Lou’s framing — did you do it right, or did you do the right thing? — and the fact that a rubric can only ever answer the first.
  • The compounding risk: a drifted evaluator becomes the fitness function for a self-improving pipeline, which then gets measurably better at producing work you like less.
  • The contrast with a broken pipeline: a crash throws an error, a drifted rubric ships — with a number attached, which is worse.

Proposed Structure (5–7 beats)

  1. Open on the disagreement: the writer says bad, the machine says 4.9, and there is no error message anywhere.
  2. Establish that the rubric was good advice and did work — this is not a piece against quality gates. It’s a piece about their second-order failure.
  3. The mechanism of drift: reasonable increments compounding into an undesigned instrument. Nobody made a bad decision; the outcome is still bad.
  4. Why the number is the dangerous part — it looks like evidence, and evidence overrides intuition, which is exactly backwards here.
  5. The compounding case: drift as fitness function, and a pipeline optimizing toward work you don’t want.
  6. The calibration pass, concretely: five high-scoring artifacts, cold read, rank them yourself, diff against the machine, find the over-weighted criterion.
  7. Close on the hierarchy: you are the ground truth. Everything downstream is a compression of you, and compressions decay.

Editorial Notes

Direct sequel to Brief - The Output Quality Gate Train a Rubric Not a Prompt — do not re-argue the case for rubrics, assume it and go straight to the failure. That existing brief is the setup; this is the payoff, and the two should not be drafted by anyone who hasn’t read both. Resist any framing that makes the evaluator sound broken or the reader careless; the entire force of the piece is that everything worked as designed and the outcome was still wrong. The calibration exercise must be concrete enough to run in twenty minutes or it will not get run. Avoid statistics about AI evaluation accuracy — the argument is structural and doesn’t need borrowed authority.

Next Step

  • Approved for drafting
  • Needs revision
  • Deprioritised