Is Jev more consistent than an LLM judge on the same trace?

Yes, by a wide margin, in one independent test. LangChain scored Jev and three LLM judges, two GPT-5.6 configurations and Claude Sonnet 4.6, on the same five agent traces, repeating each judgment 100 times per trace and measuring how much each judge’s quality score moved across those repeats. Jev’s variance was the lowest of the four, and the three LLM judges ranged from 92 to 913 times higher: Claude at 92x, one GPT-5.6 configuration at 433x, the other at 913x.

An LLM judge is stochastic by construction; run it twice on an unchanged input and it can still disagree with itself, because it generates text and a score gets parsed back out of that text. Jev never generates text: it returns a calibrated probability for a typed question in one pass, which LangChain’s own post credits for the gap, while being careful to call that a hypothesis rather than a proven cause. On the same traces, Jev also matched a human reviewer’s pass or fail label on all 500 repeated decisions; the three LLM judges matched on 80% to 99.8%. It’s one test on five examples, early and narrow by the authors’ own account, not a settled number.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y