What is an LLM judge?

An LLM judge is a prompt that grades an agent’s output by reasoning over it against one or more stated intents, the things the agent was supposed to do. Instead of matching a pattern, it reads the output, works through whether each intent was met, and returns a verdict, which is why it fits properties that need judgment: whether an answer is grounded in what was retrieved, whether it follows a policy, whether its tone fits the situation.

Its verdicts inherit the properties of the model running it. Run the same judge twice on the same input and it can disagree with itself, and swapping the underlying model changes what its verdicts mean even with the prompt held fixed. It’s also the most expensive grader per verdict, since it pays for an inference call every time it runs, which is why most teams point it at a fixed eval set before a change ships and a sampled or flagged slice of production traffic, rather than every output.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y