all answers

General agent concepts

Agent harness

An agent harness is everything in an agent that isn't the model: the system prompt and instruction files, the tools and their descriptions, how context is assembled and trimmed, hooks and middleware, any sandbox or subagents, and the loop that calls the model, runs what it asks for, and decides when to stop. Claude Code and Codex CLI are harnesses, and so is the code around the model in a team's own support agent; the same model can sit inside any of them.

It exists because a model on its own only predicts the next message. It can't run a command, keep state between sessions, or notice it's repeating the same failed edit. The harness turns those predictions into actions, with memory and limits around them.

On every turn the harness decides what the model sees, which tools it can reach, and what happens to its output, so it shapes behavior alongside the weights. It's also the part of an agent a team keeps editing: a prompt line, a new tool, a changed default.

A harness change can make an agent worse with the model untouched. In EvoHarnessBench, a 2026 benchmark that grows an agent's harness stage by stage, adding tools, skills, or specialist agents, with nothing else changed, could lower results on tasks the agent had already solved. Harness engineering is a young discipline, so writers still draw a harness's boundary differently: some include observability and evaluation, others only the runtime around the loop.

8 questions

Answered, plainly.

Does a better harness matter more than a better model?Sometimes. Tuning the harness per model has closed the gap between two models on one benchmark, but the effect varies so much by model that there's no general rule.answer →Is my agent getting worse because of the model or the harness?It can be either. Rerun the failing inputs several times on the old model with today's harness, then today's model with the old harness; whichever swap restores results points at the cause.answer →What's the difference between an agent harness and an agent framework?A framework is a library you build an agent with; a harness is the running system around the model in one agent: its prompts, tools, context handling, hooks, and loop.answer →Can adding more tools make an agent worse?Yes: EvoHarnessBench found adding tools, skills, or specialist agents to a harness can drop performance by up to 46% on tasks it had already solved.answer →How do I test a harness change before it reaches production?Run a baseline on a held-out eval set, change one thing at a time, and check it didn't break a task that already passed before shipping it.answer →What is harness engineering?Harness engineering is the discipline of designing and evolving the harness around a model so an agent's mistakes get fixed at the system level.answer →Why does the same model score differently in different agent harnesses?Because the harness decides what the model sees; a 2026 study's combined treatment, trimming old tool results and responding to stalls, moved one model's score 21 points on a partial-credit metric.answer →Does running more tasks fix an unreliable agent benchmark?Not on its own. When uncertainty comes from thin scaffold coverage, infinite extra tasks raise ranking reliability by at most 0.097 on a 0-to-1 scale.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y