What's the difference between an LLM judge and a reward model?

An LLM judge is a general model, prompted at grading time with the criteria you want checked; a reward model is a separate model trained specifically to output one preference score, with the criteria baked into its training data rather than a prompt.

Change what an LLM judge should check and you edit the prompt; change what a reward model should check and you retrain it on new preference data, which is slower and needs a labeled dataset to start. A reward model also returns a single number with no explanation, while an LLM judge can be asked for a verdict plus the reasoning behind it, so its calls can be read and audited one at a time rather than trusted as a black box. Reward models are the common choice inside training, like RLHF, where the same scoring policy has to run millions of times cheaply; an LLM judge is the more common choice for evaluation, where criteria change often. A judge’s per-event cost is why it usually runs on a sample rather than every trace, a tradeoff a reward model’s cheap, fixed inference doesn’t face the same way.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y