Can synthetic data replace production traces in an eval dataset?

Not as a permanent substitute, though it’s a fine way to start. Synthetic cases, generated from your own docs or a scenario prompt, get an eval set running before real traffic exists to pull from, and they’re good at covering an edge case you know matters but rarely see. What they can’t do is tell you how often a failure actually happens: a synthetic case only tests what you thought to generate, while production carries the malformed inputs, unusual phrasing, and edge cases nobody wrote down. Hamel Husain and Shreya Shankar’s evals FAQ, built from questions asked by the 700+ engineers and PMs they’ve taught, puts it directly: synthetic data can’t tell you how common a failure is in production, and it can miss details that matter in specialized domains, legal or medical content especially. Start synthetic if that’s what gets you running. Compare synthetic cases against real traces the moment traffic exists, and let production keep replacing what you invented.

sources

keep reading

More on this.

Self-host Tessary.

Free and open source. Point it at the traces your agent already emits.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y