How do I test a harness change before it reaches production?
Run it against a baseline on a held-out eval set before it reaches production, change one thing at a time, and check it didn’t regress a task that already passed, not just the one it was meant to fix. LangChain’s own harness-tuning recipe splits evals into an optimization set, used while iterating, and a holdout set never touched until a candidate is scored, since an autonomous tuning loop otherwise overfits to whatever it’s measured against; a human still reviews what passed before anything ships.
A full benchmark run is often too slow to repeat on every candidate. DeltaSelect, a 2026 method for testing coding-agent changes cheaply, found only 19.5% of a benchmark’s own tasks, 22 of 113 in a resampling analysis of DeepSWE’s trials, tracked the full run’s result closely enough to stand in for it, and selects just that subset to compare a baseline against a candidate within a fixed budget. One case study revising an agent’s custom skills and instructions this way cut cost 58% while raising the score, for about $28 across 13 runs.
Neither replaces a production readout; both make a harness regression cheap enough to catch before it ships.