How do I test Jev against the LLM call it would replace?
Run both against the same held-out sample with real ground truth, not the LLM’s own past answers, and compare accuracy and cost per correct decision rather than cost per token. One independent benchmark put a single yes-or-no question to both jev-1.13 and Claude Haiku 4.5 on 2,000 labeled emails, half phishing and half legitimate: Jev scored 62.6% against Haiku’s 81.3%, a clear loss. Splitting the same judgment into five narrow questions, fitting a regression on half the emails, and scoring the other half flipped the result: 95.0% for Jev against 93.2% for Haiku, a gap too small to call significant, at a twelfth to a twenty-seventh of Haiku’s cost.
That second run isn’t the whole test: the same benchmark checked a cheap non-AI baseline first, a two-line regex on the link’s domain and hosting type, which reached 91.8% accuracy on the same held-out half, beating Jev’s best signal (89.4%) and within 3.2 points of both fitted classifiers. Decompose the judgment, fit weights on half your labels, and score a rule-based baseline on the same split before trusting a fitted model’s number. That labeled sample is the same golden set an eval suite still needs once production traces pile up.