all answers

Frameworks and tooling

Jev

Jev is TypeSafe AI's decision model, launched September 15, 2026, and the first sold as a System One model. It doesn't generate text. You send it state, text or JSON, and named questions: a choice between options you define, a score on ordered levels, or a yes-or-no probability. Each comes back as a typed answer with a probability and a confidence score.

Questions run independently and in parallel, so you split a judgment into narrow questions and combine the answers in code. Jev isn't fine-tuned per customer; everything you control is in the request. It's trained for calibrated probabilities rather than preferred text. An early independent test found its choices and scores overconfident, most of all on questions the input can't answer.

TypeSafe's own numbers put it at 70 to 500 milliseconds per request and $0.042 per million input tokens. Each version ships with a published list of how it fails, including literal reading, counting, and injected instructions.

9 questions

Answered, plainly.

Can Jev choose which tool an agent calls?Yes. TypeSafe's function-calling cookbook has Jev answer a single choice question listing every tool, returning a confidence with the pick, from 1.00 down to 0.53.answer →Can Jev replace an LLM judge for grading agent traces?Partly. Jev can take the narrow yes-or-no parts of a judge's rubric, but not a judgment that needs reasoning across a whole trace or a written reason for the verdict.answer →Is Jev safe to run on untrusted user input?Not unguarded. TypeSafe says jev-1.13 doesn't treat input as hostile, so injected text can move its answer, though only among the answers you defined.answer →What is Jev?Jev is TypeSafe AI's decision model: send it text or JSON plus questions with answers you define, and it returns typed answers with probabilities instead of text.answer →Which agent tasks is Jev good for, and which should stay on an LLM?Jev fits narrow decisions with answers you can name in advance: classification, scoring, routing, verification. Anything needing generated text stays on an LLM.answer →How do I test Jev against the LLM call it would replace?Run both against the same held-out labeled sample and compare accuracy and cost per correct decision; one test found splitting the question flips the result.answer →Does Jev beat a frontier LLM's accuracy on the same task?On TypeSafe's own benchmark, Jev averaged 67.8% accuracy at $0.0004 a case, beating Claude Opus 5 run as a single prompt (64.8% at $0.34).answer →Can I self-host Jev?No. TypeSafe only serves Jev through its own hosted API or OpenRouter; there's no self-hosted build, downloadable weights, or on-premises option.answer →Is Jev more consistent than an LLM judge on the same trace?Yes: in an independent LangChain test, Jev's score on identical traces varied 92 to 913 times less across repeats than GPT-5.6 or Claude Sonnet 4.6 judges did.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y