# Can Jev replace an LLM judge for grading agent traces?

Partly: Jev can take over the narrow, yes-or-no parts of an [LLM judge's](/answers/llm-as-judge/what-is-an-llm-judge) rubric, but not a judgment that needs reasoning across a whole trace or a written reason for the verdict. Jev needs the rubric split into atomic questions, such as whether a reply broke a stated policy or whether a quote supports its claim. It answers each independently, and you combine them in code, as TypeSafe's [composite scoring guide](https://docs.typesafe.ai/patterns/composite-scoring.md) describes.

TypeSafe's [citation-check cookbook](https://docs.typesafe.ai/cookbooks/citation_check.md) runs that shape on eight citations with jev-1.12: the four accurate ones came back verified at 0.93 confidence or higher, and all four planted failures were caught, the fabricated quote by a string match in code rather than by Jev. That's eight examples, run by the vendor.

Agent traces run into two of the failure modes TypeSafe [publishes for jev-1.13](https://docs.typesafe.ai/model-jaggedness/jev-1.13.md). Accuracy falls as the input fills with detail unrelated to the question, so send the spans a question is about rather than the whole trace. And it doesn't count reliably, so "how many tool calls failed" belongs in code. A failing verdict also comes back without a reason, which is the part you need when you go to fix the agent.

---

Sources:
- TypeSafe docs, Composite scoring: https://docs.typesafe.ai/patterns/composite-scoring.md (fetched 2026-09-21)
- TypeSafe docs, Citation check cookbook: https://docs.typesafe.ai/cookbooks/citation_check.md (fetched 2026-09-21)
- TypeSafe docs, Jev 1.13 jaggedness: https://docs.typesafe.ai/model-jaggedness/jev-1.13.md (fetched 2026-09-21)

Source: https://tessary.ai/answers/jev/can-jev-replace-an-llm-judge
More on Jev: https://tessary.ai/answers/jev
From Tessary, agent reliability for AI agents in production: https://tessary.ai
