Writing · 2025-09-18

The score is not the work

A high metric can still describe a system nobody can use. What claim extraction taught me about evaluation.

Automatic scores are useful fixtures. They are a terrible definition of success when the system is meant to help a person make a decision.

That is the thread I keep returning to: an evaluation has to preserve the shape of the work. A claim that wins a benchmark but cannot be checked is not a successful claim. An agent that completes a scripted trace but misses intent has not done the job.

Back to writing