Would stratified sampling fix a silent failure?

Only partly, and only if you already have a guess about where the failure lives. Stratified sampling raises your review rate inside a slice you choose, a specific tool, a specific customer segment, a specific step in the flow, so a failure concentrated there gets caught sooner than uniform random sampling would catch it. That’s a real improvement when you have a hypothesis to test.

It doesn’t help with the failure you don’t have a hypothesis about yet. One that doesn’t correlate with any dimension you thought to stratify on hides in whatever slice you left thin, same as it would under plain random sampling. Stratification changes where your review budget concentrates; it doesn’t raise the total share of production you’re reading. The only way around that ceiling is reading more of production, or reading all of it.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y