all answers

Agent reliability

Eval datasets

An eval dataset is the set of cases an agent gets judged against: inputs paired with the behavior expected on them. Every score a suite produces is relative to this set, so the dataset defines what passing means. Cases come from two sources with different properties.

Golden sets are hand-built: curated inputs with expected behavior, labeled by people who understand the agent. They're precise and reviewable, and they do double duty. They test the agent, and they supply the ground truth for calibrating automated graders, since a grader's precision and recall are measured on labeled cases. Their weakness is staleness. The agent changes, its prompts change, and its users change, while a hand-built set stays fixed, so over time the set describes a past distribution of traffic. Pass rates on a stale set measure agreement with that past distribution.

Production traces are the other source: real sessions, including the messy inputs and rare paths nobody thought to write down. A failure observed in a production trace is a real failure by definition, which grounds the case in traffic the agent actually receives. A trace records one exact interaction, though, while the failure it exposes is usually a class of inputs. A case built from a trace stays meaningful across wording changes to the extent it captures that class.

Mature datasets draw on both sources: a curated core covering the behaviors that must hold, which also serves grader calibration, refreshed with cases promoted from production that keep the set aligned with current traffic. Provenance, a record on each case of where it came from and when it was last verified, is what makes staleness visible. Each human label added along the way is fresh ground truth for calibration, so labeled cases accumulate value beyond the single behavior they check.

10 questions

Answered, plainly.

Can I start running agent evals before I have a labeled dataset?Yes. An LLM judge can score production traces against a rubric with no labeled set required; the labeled cases build themselves from what gets flagged and reviewed.answer →How do I tell whether an eval case has gone stale?Check its capture date against what changed since: a case written before a prompt, tool, or model swap is testing a version of the agent that no longer exists.answer →How many eval cases do I need before I can start?Around 25, pulled from your docs or from production, is enough to begin, per Confident AI's guide to LLM evals for startups. Coverage of distinct behaviors matters more than count.answer →What should each eval case record about where it came from?At minimum: the trace it was pulled from, who or what confirmed the expected behavior, and the date it was captured. That record is what makes staleness visible later.answer →When do I still need a golden set?You still need a golden set the moment you want to trust the grader itself: precision and recall are measured against hand-labeled cases, and nothing else supplies that ground truth.answer →Where do the cases in an agent eval dataset come from?From two places: hand-built golden cases with a labeled expected behavior, and cases pulled from real production traces, especially ones where the agent already failed.answer →What's the difference between an eval dataset and a training dataset?A training dataset teaches a model behavior; an eval dataset measures that behavior afterward. Passing evals never means the same data trained it.answer →Can synthetic data replace production traces in an eval dataset?Not as a permanent substitute. Synthetic cases get an eval set running fast, but only production traces show how often a failure actually occurs.answer →Can a brand-new eval suite already have a coverage gap?Yes. A coverage gap is the mismatch between what an eval set contains and what production sends, and it can exist the day the set is written, before anything goes stale.answer →Does a higher pass rate mean an eval suite has better coverage?No. Pass rate is computed only over the suite's existing cases, so adding easy ones raises the score without closing any coverage gap.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y