What is a rubric in LLM-as-judge grading?

A rubric is a written list of specific, checkable criteria a judge scores one at a time, in place of one open-ended question like “is this good.” Instead of asking a model to weigh tone, correctness, and policy compliance all at once and hand back a single verdict, a rubric asks it to answer each separately: did the answer cite its source, did it follow the stated policy, did it address the actual question.

The split works because each sub-question is narrower than the one it replaces, and a narrower judgment leaves less room to average away a real problem inside one vague score. On agent tool-calling traces, a structured rubric prompt raised one judge’s agreement with human labels by up to 6.5 points. In a follow-up test on four models picked for the highest baseline self-preference bias, splitting a verdict into separate rubric dimensions cut that bias by 31.5% on average; a separate test across 20 judges found bias didn’t track with how capable the judge model was. A grader is the more general version of the same idea: one check, one specific thing it’s allowed to say yes or no to.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y