What are the different kinds of agent behavior drift?

One 2026 taxonomy splits agent behavior drift into three kinds: semantic drift, a gradual departure from the original task intent; behavioral drift, the agent settling into strategies or tool use nobody specified; and coordination drift, multi-agent agreement breaking down over a long interaction. Semantic drift shows up as the output wandering from what was asked, while coordination drift only shows once you compare what agents agreed on earlier against what they agree on now.

The paper is a single-author, simulation-based study, not a measurement on production agents, and its own simulations found semantic drift appearing in close to half of workflows by 600 interactions.

Tessary’s own behavior_drift classifier watches one of the three: it learns a call site’s normal sequence of steps and flags a trace that departs from it, which is the shape behavioral drift takes. It has nothing to say about whether a single agent’s answers have wandered from the original ask without changing which steps it took, or whether two agents in a handoff still agree; an LLM judge is the check built for that.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y