What Is an Agent Harness? Why the Scaffolding Matters More Than the Model
Published: 2026-09-28 · Author: Muhammad Tayyab Ilyas
Quick answer
An agent harness is everything in an AI agent except the model itself: the tools, memory, context handling, checks and recovery logic wrapped around it. For a coding agent, it is the system of controls that raises your confidence in generated code and lets the agent correct itself before a human looks. It matters because, in practice, the scaffolding often decides the outcome more than the model does.
What is an agent harness?
Birgitta Böckeler, writing on martinfowler.com, defines a harness as "everything in an AI agent except the model itself." Addy Osmani puts the same idea another way in Agent Harness Engineering: a coding agent is the model plus all the scaffolding around it.
Think of the model as an engine and the harness as the rest of the car. It decides what the agent can touch, what it remembers, how its work is checked and what happens when something breaks.
If you have used a coding agent for more than an afternoon, you know the same model can feel brilliant in one setup and useless in another. That difference is the harness.
Why it matters now
Harness engineering is described as the third phase of AI engineering maturity, after prompt engineering and context engineering (see the Faros write up). Prompt engineering asked how to phrase a request. Context engineering asked what the model should see. Harness engineering asks what system keeps an agent reliable over hours of autonomous work.
Osmani's line captures the shift: "A decent model with a great harness beats a great model with a bad harness." Gartner has published an "Innovation Insight: Coding Agent Harness Engineering" arguing that models alone are not enough to scale an AI native software lifecycle, and that leaders should invest in the harness.
Mitchell Hashimoto's rule is the working principle behind all of it: "Anytime you find an agent makes a mistake, you take the time to engineer a solution so that the agent never makes that mistake again." We go deeper on that habit in turning every agent mistake into a rule.
The components of a coding agent harness
Osmani lists the pieces that show up again and again:
- Filesystem and git. Durable state that survives a crashed session.
- Bash and code execution. The agent needs to run things, not just write them.
- Sandboxes. A safe place to run those things.
- Memory. AGENTS.md files and context injection, so lessons persist between runs.
- Context management. Compaction, offloading bulky tool output, and skills that load in stages.
- Long horizon execution. Ralph loops, planning, verification, and splitting planner from evaluator.
- Hooks. Checks that fire before a tool call, after an edit or before a commit.
- A short rulebook. AGENTS.md kept small, roughly under 60 lines in his suggestion.
His practices are worth copying. Work backwards from the behavior you want. Every constraint should trace to a past failure. Ten well designed tools beat fifty overlapping ones. Success should be silent, while failures inject detailed error context back into the loop. And in his words: "If you can't name the behaviour a component exists to deliver, it probably shouldn't be there."
Guides and sensors, computational and inferential
Böckeler gives us two useful axes for sorting controls.
The first is timing. Guides are feedforward: they steer the agent before it acts. Linters, architectural rules, documentation and codemods are guides. Sensors are feedback: they observe after the agent acts and trigger self correction. Tests, static analysis and code review agents are sensors. You want both, because guides prevent errors and sensors catch the ones that slip through.
The second is how the check works. Computational controls are deterministic and fast: tests, linters, type checkers. Inferential controls are semantic and AI based: an LLM code review or a custom judge. Computational checks are cheap and trustworthy but narrow. Inferential checks can judge intent but can be wrong or fooled. A good harness layers them.
She also names three things a harness regulates: maintainability (the most developed), architecture fitness, and behavior, meaning functional correctness, which is the least developed today. Her closing idea is a good test of any design: a harness should not aim to eliminate human input, but to direct it to where it matters most.
What the September 2026 study found
A new paper, An Empirical Study of Harness Design for Coding Agents by Fan et al. (arXiv 2609.20804, 17 September 2026), measured the effect of harness choices directly. It compared 176 matched configurations across four language models on SWE-Bench Verified and Terminal-Bench 2.1, looking at context management, planning and action space.
- Context management mattered most. It mainly helped by preventing context overflow failures and extending how long an agent can keep working. Staging rule based elision before LLM based summarization gave the best overall efficiency.
- Planning changes role with model strength. It "shifts from an accuracy scaffold for weaker models to a cost saver for stronger models."
- Action space depends on the model. Stronger models got substantially lower cost with bash only interfaces on command line tasks, while weaker models benefited from predefined tools.
The practical reading: there is no single best harness, because the right scaffolding depends on the model inside it.
Inside a real harness: how LoopCodeLab is built
LoopCodeLab is a good example because its whole product is a harness. You describe an idea. A planner turns it into user stories. Worker agents build stories in parallel, each in its own isolated git worktree on its own branch. A master agent reviews every story against its acceptance criteria before merge, and a finalize pass delivers a live preview, a GitHub repo you own and, depending on the output, things like an Android APK or Windows installer. Here is how that maps onto the components above.
Filesystem and git. Worktrees and branches are the durable state. A worker that dies leaves its branch behind, and nothing touches main until review passes.
Long horizon execution. The orchestrator is literally called Ralph, a nod to the Ralph loop. It runs each agent in a loop per story with a capped number of attempts.
An inferential sensor. The master review is a semantic check of each story against its acceptance criteria. The master decides by writing a verdict file, and the parse is echo proof: a CLI that prints back its own prompt, which contains the words ACCEPT and REJECT, cannot fake a verdict. That is a sensor designed to survive a confused model.
Escalation instead of endless loops. After two rejections of the same story, the master does one intervention pass itself. Only if that fails does the build ask a human, which is Böckeler's "direct human input to where it matters" in practice.
Fallback chains and a pre start check. If an agent runs out of quota, work fails over to the next agent you configured. Before a build begins, a probe checks whether the chosen agents are actually usable. LoopCodeLab drives many coding agent CLIs as interchangeable workers, so the harness cannot assume any one of them.
A failure taxonomy. Auth, quota, stall and review provider failures are classified so the right recovery runs. Partial work is preserved, and a retry can re review an existing branch instead of rebuilding. We wrote about the state model in why AI coding agents fail and how to recover safely.
Memory and skills. Builds suggest reusable lessons and a human approves them. A per build logbook keeps the master's rulings consistent. The codebase also keeps a short rule file per directory documenting traps that bit before, so the humans and agents who work on LoopCodeLab get the same treatment. More on that loop in how LoopCodeLab gets smarter every build.
A no spend test harness. A stub mode makes every worker, review and finalize step deterministic, with auto complete and auto accept, so the entire orchestrator can be tested end to end at no API cost. You cannot improve a harness you cannot test.
Fail soft by contract. Planning works without research keys, a delivery failure never fails a finished build, and token metering never blocks work.
How to evaluate a harness
Whether you are choosing a tool or building your own, ask these questions:
- Is state durable? Can a crashed agent be resumed without losing finished work?
- Are there both guides and sensors, and both computational and inferential checks?
- Can a model fake success? Look for verdicts that cannot be spoofed by echoed text.
- Does it escalate sensibly, with a capped retry count, then a stronger intervention, then a human?
- Does it classify failures instead of treating every error as generic?
- Can you swap the model without rewriting everything?
- Does every rule trace back to a real failure?
- Can you test the harness itself without spending money?
Frequently asked questions
What is an agent harness?
An agent harness is everything in an AI agent except the model itself. For coding agents it means the tools, memory, context management, checks and recovery logic that keep the agent reliable and let it correct itself before human review.
What is harness engineering?
Harness engineering is the practice of designing and improving that scaffolding. It is often described as the third phase after prompt engineering and context engineering, and it focuses on turning observed agent failures into permanent controls.
Is the harness really more important than the model?
Often it is. Addy Osmani argues that a decent model with a great harness beats a great model with a bad harness, and a September 2026 study of 176 configurations found that harness choices such as context management had a large effect on results.
What is the difference between guides and sensors?
Guides steer an agent before it acts, for example linters, architectural rules and documentation. Sensors observe after it acts and trigger correction, for example tests, static analysis and code review agents.
Do I need to build my own harness?
Not necessarily. You can adopt a platform whose harness already handles isolation, review, retries and recovery, then add your own rules on top as you learn where your projects fail.
See a harness at work
You can watch these pieces run on a real project: describe an idea to LoopCodeLab and follow the plan, build and review loop. For the mechanics, read how a build works, then see what the wider industry is doing in agentic coding trends for 2026 and the agent harness.