Writing · 2025-09-18
The score is not the work
A high metric can still describe a system nobody can use. What claim extraction taught me about evaluation.
Automatic scores are useful fixtures. They are a terrible definition of success when the system is meant to help a person make a decision.
That is the thread I keep returning to: an evaluation has to preserve the shape of the work. A claim that wins a benchmark but cannot be checked is not a successful claim. An agent that completes a scripted trace but misses intent has not done the job.