Does running more tasks fix an unreliable agent benchmark?

Not on its own. A Bayesian reliability study of 22 agent benchmarks drawn from the Holistic Agent Leaderboard and the Harbor Index found that when a benchmark’s uncertainty comes mainly from thin scaffold coverage rather than too few tasks, running infinitely many more tasks of the same kind raises model-ranking reliability by at most 0.097 on its 0-to-1 scale. The same study found a fixed model-and-harness pairing already ranks reliably (0.935 to 0.994), while ranking the underlying model alone is far less reliable (0.148 to 0.841): system reliability counts a scaffold’s own differences as signal, but model reliability needs a difference that survives averaging over scaffolds. More tasks sharpen the same evaluation conditions; they can’t substitute for testing a model across different scaffolds, which is the variation a model-level ranking actually needs. The paper’s working fix is pooling several diverse benchmarks instead of piling tasks onto one: that raised projected reliability from 0.44 to 0.75 at the same task budget, and cut projected cost by up to 83%. A benchmark score describes the model-and-harness pair that produced it, so when a ranking looks shaky, check scaffold coverage before adding tasks.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y