Is Jev safe to run on untrusted user input?

Not unguarded: TypeSafe’s own notes on jev-1.13 say Jev doesn’t treat the input it evaluates as hostile by default, so an injected instruction, a misleading framing, or text that argues for its own classification can move the answer. It’s the same weakness LLM judges have with the content they grade, because the untrusted text and your question arrive in one request.

What differs is how far an attack can reach. Jev only returns answers you defined, so injected text can shift a probability or flip a choice, but it can’t make Jev write anything or call a tool. The damage is whatever your code does with that answer, so gate a destructive action behind a stricter threshold than a read-only one, as TypeSafe’s confidence guide does, or behind a person.

TypeSafe does document Jev as a jailbreak screen, and in its guardrails cookbook, run with jev-1.12 on ten sample prompts, the DAN prompt scored 0.98 as a jailbreak. That’s text aimed at another model being classified, not text written to steer Jev itself.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y