What stops a description from producing a grader that measures the wrong thing?

Nothing, by itself. A grader written from a sentence like “flag responses that sound rude” encodes whatever the author pictured when they wrote it, and that picture can be wrong in ways that only show up once the grader is running: too broad, catching a tone that’s actually fine, or too narrow, missing the actual pattern in production.

The fix isn’t a better sentence. It’s grounding the description in a real trace, an example the grader is drafted against, checked, and adjusted until it agrees with a human reading the same trace. A grader that can point to the specific trace its rule came from is measuring something that happened; one that can only point to a rubric someone wrote is measuring what that person imagined.

Before trusting a new grader’s verdicts, ask what trace it was checked against. If the answer is none, the verdict is a guess wearing a pass rate.

keep reading

More on this.

Send us the traces you already emit.