What is Tessary's behavior_drift classifier?
It detects traces that depart from how an agent normally moves: the steps it takes and the order it takes them in, learned from that call site’s own history. It doesn’t judge whether any single answer was right.
That covers the changes a correctness check misses. A step the agent always ran and quietly stopped running. A tool it starts calling that it never called before. A path through a conversation it has never taken. None of that has to produce a wrong-looking answer to be worth knowing about.
The baseline is fit per call site from its own traffic, so a new project stays quiet until Tessary has seen enough sessions there to know what normal looks like. A pattern that keeps recurring graduates into the new normal instead of firing forever, and everything that does fire goes to triage first.