Does Tessary truncate a long trace before grading it?

Partly: two classifiers read a bounded window, while triage and RCA, the LLM steps behind them, include fewer items rather than clip any one of them.

In the LLM lanes the bound is by count, not by content. A tool result or a model completion either goes in whole or doesn’t go in at all. Grading a clipped payload is worse than grading fewer items, because the one line that was wrong is usually the line that got cut, and the verdict comes back confident anyway.

The two classifiers that read text through a model clip on purpose. Frustration reads one user message and the four before it, swaps pasted blocks for a marker, and cuts a long message to its head and tail. Groundedness reads an answer and its retrieved documents together in one pass of up to 8,192 tokens, per its model card. A very long answer is cut at a fixed length, the documents are shortened to fit beside it, and an answer that still leaves no room for them isn’t scored at all.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y