Can Jev replace an LLM judge for grading agent traces?
Partly: Jev can take over the narrow, yes-or-no parts of an LLM judge’s rubric, but not a judgment that needs reasoning across a whole trace or a written reason for the verdict. Jev needs the rubric split into atomic questions, such as whether a reply broke a stated policy or whether a quote supports its claim. It answers each independently, and you combine them in code, as TypeSafe’s composite scoring guide describes.
TypeSafe’s citation-check cookbook runs that shape on eight citations with jev-1.12: the four accurate ones came back verified at 0.93 confidence or higher, and all four planted failures were caught, the fabricated quote by a string match in code rather than by Jev. That’s eight examples, run by the vendor.
Agent traces run into two of the failure modes TypeSafe publishes for jev-1.13. Accuracy falls as the input fills with detail unrelated to the question, so send the spans a question is about rather than the whole trace. And it doesn’t count reliably, so “how many tool calls failed” belongs in code. A failing verdict also comes back without a reason, which is the part you need when you go to fix the agent.
sources
- TypeSafe docs, Composite scoring fetched
- TypeSafe docs, Citation check cookbook fetched
- TypeSafe docs, Jev 1.13 jaggedness fetched