Can I start running agent evals before I have a labeled dataset?

Yes. You don’t need a hand-labeled dataset to start grading an agent: write a rubric for what a good and a bad answer look like, and an LLM judge can score real production traces against it today. Nothing has to exist before you can start.

The labeled set builds itself as a side effect of doing that. Traces the judge flags, or that a human overrides, go into review, and once someone confirms what the agent should have done, that trace becomes a labeled case grounded in a real failure rather than one invented at a desk.

What this doesn’t give you is confidence in the judge itself. A rubric-based grader can be wrong in ways nobody notices until its verdicts get checked against real labels, which is a separate problem from grading the agent, and the reason a golden set still matters even once you’ve started.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y