Can Jev replace an LLM judge for grading agent traces?

Partly: Jev can take over the narrow, yes-or-no parts of an LLM judge’s rubric, but not a judgment that needs reasoning across a whole trace or a written reason for the verdict. Jev needs the rubric split into atomic questions, such as whether a reply broke a stated policy or whether a quote supports its claim. It answers each independently, and you combine them in code, as TypeSafe’s composite scoring guide describes.

TypeSafe’s citation-check cookbook runs that shape on eight citations with jev-1.12: the four accurate ones came back verified at 0.93 confidence or higher, and all four planted failures were caught, the fabricated quote by a string match in code rather than by Jev. That’s eight examples, run by the vendor.

Agent traces run into two of the failure modes TypeSafe publishes for jev-1.13. Accuracy falls as the input fills with detail unrelated to the question, so send the spans a question is about rather than the whole trace. And it doesn’t count reliably, so “how many tool calls failed” belongs in code. A failing verdict also comes back without a reason, which is the part you need when you go to fix the agent.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y