# What is a rubric in LLM-as-judge grading?

A rubric is a written list of specific, checkable criteria a judge scores one at a time, in place of one open-ended question like "is this good." Instead of asking a model to weigh tone, correctness, and policy compliance all at once and hand back a single verdict, a rubric asks it to answer each separately: did the answer cite its source, did it follow the stated policy, did it address the actual question.

The split works because each sub-question is narrower than the one it replaces, and a narrower judgment leaves less room to average away a real problem inside one vague score. On agent tool-calling traces, a structured rubric prompt raised one judge's agreement with human labels by up to 6.5 points. In a follow-up test on four models picked for the highest baseline self-preference bias, splitting a verdict into separate rubric dimensions cut that bias by 31.5% on average; a separate test across 20 judges found bias didn't track with how capable the judge model was. [A grader is the more general version of the same idea](/answers/graders/what-is-a-grader): one check, one specific thing it's allowed to say yes or no to.

---

Sources:
- AgentJudgeBench (arXiv:2608.26623): https://arxiv.org/abs/2608.26623 (fetched 2026-09-17)
- "Quantifying and Mitigating Self-Preference Bias of LLM Judges" (arXiv:2604.22891): https://arxiv.org/abs/2604.22891 (fetched 2026-09-17)

Source: https://tessary.ai/answers/llm-as-judge/what-is-a-rubric-in-llm-as-judge-grading
More on LLM as judge: https://tessary.ai/answers/llm-as-judge
From Tessary, agent reliability for AI agents in production: https://tessary.ai
