# How do I test Jev against the LLM call it would replace?

Run both against the same held-out sample with real ground truth, not the LLM's own past answers, and compare accuracy and cost per correct decision rather than cost per token. One independent benchmark put a single yes-or-no question to both jev-1.13 and Claude Haiku 4.5 on 2,000 labeled emails, half phishing and half legitimate: Jev scored 62.6% against Haiku's 81.3%, a clear loss. Splitting the same judgment into five narrow questions, fitting a regression on half the emails, and scoring the other half flipped the result: 95.0% for Jev against 93.2% for Haiku, a gap too small to call significant, at a twelfth to a twenty-seventh of Haiku's cost.

That second run isn't the whole test: the same benchmark checked a cheap non-AI baseline first, a two-line regex on the link's domain and hosting type, which reached 91.8% accuracy on the same held-out half, beating Jev's best signal (89.4%) and within 3.2 points of both fitted classifiers. Decompose the judgment, fit weights on half your labels, and score a rule-based baseline on the same split before trusting a fitted model's number. [That labeled sample is the same golden set an eval suite still needs once production traces pile up](/answers/eval-datasets/when-do-i-still-need-a-golden-set).

---

Sources:
- THE D*AI*LY BRIEF, "TypeSafe's Jev Scores 62.6% Asked Once and 95% Split Five Ways" (September 20, 2026): https://www.beri.net/article/typesafe-jev-typed-decision-model-calibration-decomposition-shadow-eval (fetched 2026-09-23)

Source: https://tessary.ai/answers/jev/how-do-i-test-jev-against-the-llm-call-it-would-replace
More on Jev: https://tessary.ai/answers/jev
From Tessary, agent reliability for AI agents in production: https://tessary.ai
