Is Jev safe to run on untrusted user input?
Not unguarded: TypeSafe’s own notes on jev-1.13 say Jev doesn’t treat the input it evaluates as hostile by default, so an injected instruction, a misleading framing, or text that argues for its own classification can move the answer. It’s the same weakness LLM judges have with the content they grade, because the untrusted text and your question arrive in one request.
What differs is how far an attack can reach. Jev only returns answers you defined, so injected text can shift a probability or flip a choice, but it can’t make Jev write anything or call a tool. The damage is whatever your code does with that answer, so gate a destructive action behind a stricter threshold than a read-only one, as TypeSafe’s confidence guide does, or behind a person.
TypeSafe does document Jev as a jailbreak screen, and in its guardrails cookbook, run with jev-1.12 on ten sample prompts, the DAN prompt scored 0.98 as a jailbreak. That’s text aimed at another model being classified, not text written to steer Jev itself.
sources
- TypeSafe docs, Jev 1.13 jaggedness fetched
- TypeSafe docs, LLM guardrails cookbook fetched
- TypeSafe docs, Confidence fetched