What's the difference between an eval dataset and a training dataset?
A training dataset teaches a model behavior; an eval dataset measures that behavior afterward. The training, or fine-tuning, set is the examples the model’s weights get adjusted against, so what comes out the other side has learned something it didn’t do before. The eval set is a fixed collection of inputs paired with the outcome expected on each one, run against the trained model or agent afterward to produce a score.
The two have to stay separate for the score to mean anything. If a case a grader checks was also in the data used to train or fine-tune the agent, a high score partly reflects memorization of that exact case rather than the general skill it’s supposed to represent, the same contamination problem that undermines any benchmark. It’s also why eval cases keep value even when nothing is being trained: a golden set built to test an agent’s behavior does double duty as labeled ground truth for calibrating a grader, a role a training set was never built for.