Does running more tasks fix an unreliable agent benchmark?
Not on its own. A Bayesian reliability study of 22 agent benchmarks drawn from the Holistic Agent Leaderboard and the Harbor Index found that when a benchmark’s uncertainty comes mainly from thin scaffold coverage rather than too few tasks, running infinitely many more tasks of the same kind raises model-ranking reliability by at most 0.097 on its 0-to-1 scale. The same study found a fixed model-and-harness pairing already ranks reliably (0.935 to 0.994), while ranking the underlying model alone is far less reliable (0.148 to 0.841): system reliability counts a scaffold’s own differences as signal, but model reliability needs a difference that survives averaging over scaffolds. More tasks sharpen the same evaluation conditions; they can’t substitute for testing a model across different scaffolds, which is the variation a model-level ranking actually needs. The paper’s working fix is pooling several diverse benchmarks instead of piling tasks onto one: that raised projected reliability from 0.44 to 0.75 at the same task budget, and cut projected cost by up to 83%. A benchmark score describes the model-and-harness pair that produced it, so when a ranking looks shaky, check scaffold coverage before adding tasks.