Why do Tessary's classifiers produce false positives?

Because each classifier scores one narrow property of a trace and fires on a threshold. A threshold set anywhere flags traces that turn out to be fine, so a flag on its own is never the verdict.

The recall-precision trade differs per classifier, and some lean toward precision. Groundedness is one: on the RAGTruth benchmark, 83% of the answers it flags are unsupported, and it catches 39% of the unsupported ones, per its model card.

For most classifiers a firing writes a finding, and triage rules on it before a person sees anything, so a false positive costs compute rather than somebody’s afternoon. Frustration and groundedness guard earlier: a single flag never reaches anyone, and a finding opens only when a call site’s rate of flags rises above its own normal. A frustration finding then opens its case directly, while a groundedness finding still goes to triage.

One limit worth saying out loud: on a corpus nobody has labelled, a fire rate is a fire rate, not a false-positive rate.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y