the log
What we measured.
latest ·
Most of them returned a normal-looking response and passed sampled evals.
AI agent failures we found across 22 companies
AI agent failures often return normal responses and escape sampled evals. We tested 22 customer-facing agents to see what those production errors look like.
read the post →newest first
0 posts
nothing in the log matches that. the latest post above is pinned, so it always shows.
Self-host Tessary.
Free and open source. Point it at the traces your agent already emits.
Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md
docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y