# Does running more tasks fix an unreliable agent benchmark?

Not on its own. A Bayesian reliability study of 22 agent benchmarks drawn from the Holistic Agent Leaderboard and the Harbor Index found that when a benchmark's uncertainty comes mainly from thin scaffold coverage rather than too few tasks, running infinitely many more tasks of the same kind raises model-ranking reliability by at most 0.097 on its 0-to-1 scale. The same study found a fixed model-and-harness pairing already ranks reliably (0.935 to 0.994), while ranking the underlying model alone is far less reliable (0.148 to 0.841): system reliability counts a scaffold's own differences as signal, but model reliability needs a difference that survives averaging over scaffolds. More tasks sharpen the same evaluation conditions; they can't substitute for testing a model across different scaffolds, which is the variation a model-level ranking actually needs. The paper's working fix is pooling several diverse benchmarks instead of piling tasks onto one: that raised projected reliability from 0.44 to 0.75 at the same task budget, and cut projected cost by up to 83%. [A benchmark score describes the model-and-harness pair that produced it](/answers/agent-benchmarks/what-is-an-agent-benchmark), so when a ranking looks shaky, check scaffold coverage before adding tasks.

---

Sources:
- "Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard" (arXiv:2610.00651): https://arxiv.org/abs/2610.00651 (fetched 2026-10-09)

Source: https://tessary.ai/answers/agent-harness/does-running-more-tasks-fix-an-unreliable-benchmark
More on Agent harness: https://tessary.ai/answers/agent-harness
From Tessary, agent reliability for AI agents in production: https://tessary.ai
