# Agent harness

An agent harness is everything in an agent that isn't the model: the system prompt and
instruction files, the tools and their descriptions, how context is assembled and trimmed, hooks
and middleware, any sandbox or subagents, and the loop that calls the model, runs what it asks
for, and decides when to stop. Claude Code and Codex CLI are harnesses, and so is the code
around the model in a team's own support agent; the same model can sit inside any of them.

It exists because a model on its own only predicts the next message. It can't run a command,
keep state between sessions, or notice it's repeating the same failed edit. The harness turns
those predictions into actions, with memory and limits around them.

On every turn the harness decides what the model sees, which tools it can reach, and what
happens to its output, so it shapes behavior alongside the weights. It's also the part of an
agent a team keeps editing: a prompt line, a new tool, a changed default.

A harness change can make an agent worse with the model untouched. In EvoHarnessBench, a 2026
benchmark that grows an agent's harness stage by stage, adding tools, skills, or specialist
agents, with nothing else changed, could lower results on tasks the agent had already solved.
Harness engineering is a young discipline, so writers still draw a harness's boundary
differently: some include observability and evaluation, others only the runtime around the loop.

## Questions answered under this concept

- [Can adding more tools make an agent worse?](https://tessary.ai/answers/agent-harness/can-adding-more-tools-make-an-agent-worse)
- [Does a better harness matter more than a better model?](https://tessary.ai/answers/agent-harness/does-a-better-harness-matter-more-than-a-better-model)
- [Does running more tasks fix an unreliable agent benchmark?](https://tessary.ai/answers/agent-harness/does-running-more-tasks-fix-an-unreliable-benchmark)
- [How do I test a harness change before it reaches production?](https://tessary.ai/answers/agent-harness/how-do-i-test-a-harness-change-before-it-reaches-production)
- [Is my agent getting worse because of the model or the harness?](https://tessary.ai/answers/agent-harness/is-my-agent-getting-worse-because-of-the-model-or-the-harness)
- [What is harness engineering?](https://tessary.ai/answers/agent-harness/what-is-harness-engineering)
- [What's the difference between an agent harness and an agent framework?](https://tessary.ai/answers/agent-harness/whats-the-difference-between-an-agent-harness-and-an-agent-framework)
- [Why does the same model score differently in different agent harnesses?](https://tessary.ai/answers/agent-harness/why-does-the-same-model-score-differently-in-different-harnesses)

---

Source: https://tessary.ai/answers/agent-harness
All concepts: https://tessary.ai/answers
From Tessary, agent reliability for AI agents in production: https://tessary.ai
