What's the difference between observability and evals?
Observability records what an agent did; evals judge whether what it did was right. Instrumentation is what makes observability possible: a span for each LLM call and tool call, a trace for the turn, so you can see the exact inputs, outputs, and steps after the fact. None of that carries a verdict on its own. An eval runs a grader, a deterministic check, a trained classifier, or an LLM judge, against that recorded behavior, or against a separate test case, and produces a pass, fail, or score. You can have deep observability and no evals at all: a complete trace of an agent nobody is checking for correctness. You can also run evals on weak observability, grading outputs without enough recorded context to tell why one failed. A grader needs something to read: observability supplies the trace, and the eval supplies the judgment the trace alone never makes.