all answers

Agent reliability

LLM as judge

LLM-as-judge is a prompt that grades an agent's output by reasoning over it against one or more specific intents, the things the agent was supposed to do. The judge reads the output, works through whether each intent was met, and returns a verdict. It exists because many of the qualities that matter, like tone, helpfulness, or whether an answer actually addresses the question, need judgment to assess, and human judgment is too slow and expensive to apply to every output.

A judge's verdicts come from the model its prompt runs on, so they carry that model's properties. They're stochastic: run the judge twice on the same input and it can disagree with itself. And they depend on the model version: change the model and the verdicts change meaning, even with the prompt held fixed.

LLM-as-judge fits checks that need judgment: whether an answer is grounded in the retrieved context, whether the agent followed a policy, whether the tone fits. Because it costs the most per verdict of any grader, it usually runs on a subset rather than every output: against a fixed dataset in evals before a change ships, and in production on traces that were sampled or flagged by cheaper checks.

17 questions

Answered, plainly.

Did my agent get worse, or did my judge?You can't tell from a falling score alone; run a fixed, human-labelled anchor set through the current judge on a steady interleave and watch it separately.answer →Does an LLM judge know when it's unsure?Somewhat: a judge's self-reported confidence predicts its own errors better than sampling it five times and voting, at a fifth of the inference cost.answer →Does pinning the judge model version stop it from drifting?Not fully. A pinned version string is a record, not a guarantee, since a hosted provider can serve updated weights under the same stable alias.answer →Is 90 percent agreement good enough for an LLM judge?Not on its own. Raw agreement doesn't correct for chance, so on a skewed label set 90 percent can sit on a kappa near zero.answer →When should I use an LLM judge instead of a deterministic check?Use an LLM judge for checks that need judgment, like whether an answer addresses the question; skip it wherever a rule or a classifier can decide instead.answer →Why do judges flip their verdict when you swap the order of two answers?Position bias: an LLM judge weighs position alongside content. MT-Bench found GPT-4 kept its verdict only 65 percent of the time when the two answers swapped places.answer →Does an LLM judge score its own model's outputs more favorably?Often, yes, but not from recognizing its own writing: research ties the bias to judges rating familiar-sounding text higher regardless of its source.answer →Should an LLM judge return a pass/fail verdict or a score on a scale?Pass/fail, when the distinction is coarse enough for one. A finer scale needs quadratically more judge calls to trust at the same confidence.answer →What is an LLM judge?An LLM judge is a prompt that grades an agent's output by reasoning over it against a stated intent, for judgments a fixed rule can't make.answer →Can the content an LLM judge is grading manipulate its verdict?Yes. A 2025 study found text appended to a graded response swayed an LLM judge's verdict over 30% of the time, a documented prompt-injection attack.answer →If my LLM judge makes mistakes, is my measured pass rate biased?Yes, unless you correct for it. A judge's sensitivity and specificity are never 100 percent, and that error rate biases every pass rate it reports.answer →Does running more LLM judges on the same trace reduce error?Not by much: a 2026 study found nine LLM judges carried only about two judges' worth of independent signal, since they fail on the same items.answer →Are most LLM judges reliable?No. In a 2025 benchmark of 54 LLM judges scored against human raters, only half, 27, reached reliable agreement; the rest fell short.answer →What is a rubric in LLM-as-judge grading?A rubric is a written list of specific, checkable criteria a judge scores one at a time, in place of one open-ended question like 'is this good.'answer →What's the difference between an LLM judge and a reward model?An LLM judge reasons in text against criteria written at call time; a reward model is trained once to output a single learned score, with no prompt to edit.answer →How does an LLM judge compare to human evaluation?About as often as two human graders agree with each other: 85 percent versus 81 percent on MT-Bench, the benchmark that established LLM-as-judge as a working method.answer →Can you fix a correlated LLM judge panel?Yes: a 2026 study got 7 to 8 points higher accuracy by modeling which judges share failure patterns, instead of averaging votes as if each judge failed independently.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y