Harness Engineering in Practice: Turn Every Agent Mistake Into a Rule
Published: 2026-09-28 · Author: Muhammad Tayyab Ilyas
Quick answer
Harness engineering means treating every agent mistake as a missing piece of the system around the model, then adding that piece: a rule, a guide, a sensor, a hook or a test. You keep a short rule file where each line traces to a real failure, you prefer fast deterministic checks over AI judgment, and you let humans review only where their judgment matters most. This is how you make coding agents reliable without waiting for a smarter model.
This is the hands on playbook in a three part series. For the definition, start with what an agent harness is. For the market view, see the 2026 agentic coding trends. Here we get practical.
The principle: fix the system, not the prompt
Mitchell Hashimoto's rule is widely quoted in this debate: "Anytime you find an agent makes a mistake, you take the time to engineer a solution so that the agent never makes that mistake again."
Birgitta Böckeler calls the habit of acting on it the steering loop. You watch the agent fail, then you change the harness, meaning everything in the agent except the model itself. Her vocabulary is worth borrowing:
- Guides are feedforward. They steer before the agent acts: rule files, documentation, linters, codemods.
- Sensors are feedback. They observe after the agent acts and trigger self correction: tests, static analysis, review agents.
- Either can be computational (deterministic and fast, like a test or type checker) or inferential (semantic, like an LLM reviewer).
Addy Osmani puts the payoff bluntly in Agent Harness Engineering: "A decent model with a great harness beats a great model with a bad harness."
Six real failures, and what fixed each
LoopCodeLab runs a planner, parallel workers and a master reviewer across many coding agent CLIs, so we have plenty of failures to learn from. Each case has the same shape.
A. The silent reject trap
What failed. Each agent's launch line lives in the worker script and in three separate master scripts (review, finalize and research). When one agent was missing from a master script, the script fell through to "unknown tool" and wrote no verdict. Every story was silently rejected, even though the worker had built it fine.
Why a better model would not help. The model never ran.
The fix. A written rule in that directory's rule file: adding an agent touches all four scripts. When we added Muse (Meta's coding agent) in September 2026, the plan made all four mandatory, and a live smoke run proved a real review wrote an ACCEPT verdict. The written rule is a guide. The smoke run that proves a real ACCEPT is a computational sensor, and the two together are what closed the gap.
B. The logout that did not log out
What failed. A final review of the whole Muse branch found that its logout command does not delete the credentials file. It rewrites the file with an empty provider list. Our harness only checked whether the file existed, so a logged out user looked signed in and their saved API key was never used.
Why a better model would not help. No model could know that quirk without observing it. It is undocumented behavior of one tool.
The fix. Check the file's content, not its existence, plus a unit test for exactly that shape. This is a textbook sensor added because of one observed failure, and it is computational.
C. The verdict that ate its own reason
What failed. An earlier verdict parser mishandled backticks in the master's explanation and reported a false "master CLI broken".
Why a better model would not help. The reviewer wrote a perfectly good verdict. Our parser misread it.
The fix. A stricter parse and a test. Computational sensor.
D. Flags differ per CLI
What failed. One CLI rejects the auto approve flag in prompt mode and errors. Another needs a trust flag to run non interactively. Another takes the prompt on standard input while the rest take it as an argument.
Why a better model would not help. These are arbitrary conventions. A model cannot guess them, and they change between versions.
The fix. Write each quirk down as a rule. This is a guide, and it is inferential in the sense that a model reads and applies it, though the underlying facts are checked by real runs.
E. Prompts as files, not arguments
What failed. Review and finalize prompts carry large diffs. Passing them as a command line argument risks argument length limits and quoting failures.
Why a better model would not help. The prompt was fine. The delivery channel broke.
The fix. Our newest agent integration passes the prompt through a temporary file instead. That is a change in the tooling itself, and it is computational.
F. Reviews as sensors for the harness itself
What failed. Nothing failed yet, which is the point. LoopCodeLab's own features are built by subagents with a per task review and a final whole branch review on the most capable model.
Why it matters. That final review caught story B before release, after per task reviews had missed it.
The fix. Keep the whole branch review as a standing sensor. It is inferential, and it earns its cost because it catches what narrow checks cannot.
The playbook
Keep the rule file short and traceable
Osmani suggests keeping AGENTS.md under about 60 lines, and his test is strict: "Every line in a good AGENTS.md should be traceable back to a specific thing that went wrong." If you cannot name the failure behind a line, delete it.
We keep a short rule file per directory (server, orchestration, web, public, multi tenant) documenting the traps that bit before. Story A is one line in one of them.
Computational sensors first, inferential second
Böckeler's split gives you an order of operations. Tests, linters and type checkers are fast, cheap and repeatable, so run them first. Add an LLM reviewer for the questions they cannot answer, such as whether the code meets the acceptance criteria.
Stories B and C were fixed with deterministic checks. Story F needed an inferential one.
Make success silent and failure loud
Osmani's guidance: success should be silent, while failures inject detailed error context back into the loop. A passing suite that prints two thousand lines burns context. A failure that says only "error" gives the agent nothing to correct. Print the file, the expectation and the actual value.
Our master reviewer follows the same idea. It writes a verdict, and the parse is "echo proof": a CLI that prints back its own prompt, which contains both ACCEPT and REJECT, cannot fake a verdict.
Use a focused tool menu
Osmani's rule of thumb is that ten well designed tools beat fifty overlapping ones. Research points the same way. In An Empirical Study of Harness Design for Coding Agents, stronger models got substantially lower cost with bash only interfaces on command line tasks, while weaker models benefited from predefined tools. Match the menu to the model and prune duplicates.
Use hooks at lifecycle points
Hooks turn a wish into a guarantee. Osmani lists pre tool call, post edit and pre commit hooks. A rule in a document is advice. A pre commit hook that rejects a bad commit is enforcement. If a mistake recurs despite a written rule, promote it to a hook.
Test the harness itself with a stub mode
The harness is software, so it needs tests. Ours has a stub mode where every worker, review and finalize step is deterministic, auto completing and auto accepting. We can run the whole orchestrator end to end with no API cost. If you can only test your harness by spending real tokens, you will test it too rarely.
Direct human review to where it matters
Böckeler writes: "A good harness should not necessarily aim to fully eliminate human input, but to direct it to where our input is most important." Sensors handle the routine. After two rejections of the same story, our master does one intervention pass itself, and only if that fails does the build ask a person. Save human attention for judgment calls, and see why coding agents fail and how to recover safely for what those handoffs look like.
Start this week
- Collect the last five agent mistakes. Pull them from review comments, reverted commits and failed runs.
- Classify each one. Was it missing guidance, a missing check, or a tooling problem?
- Write the rule file lines. One line per failure, under about 60 lines in total.
- Turn one repeated mistake into a computational check. A test, linter rule or pre commit hook beats another paragraph of instructions.
- Add a stub mode and a weekly review. Run your loop for free, then look at what failed and repeat.
Frequently asked questions
What is harness engineering?
Harness engineering is the practice of improving everything around a coding model, such as rules, tools, tests, hooks and reviews, so the agent behaves reliably. Whenever the agent makes a mistake, you change the harness so the mistake cannot easily happen again.
What should go in an AGENTS.md file?
Only rules that trace back to a specific past failure, such as a tool quirk or a step the agent keeps skipping. Osmani suggests keeping the file under about 60 lines, so anything generic or unproven should be removed.
What is the difference between a computational and an inferential sensor?
A computational sensor is deterministic and fast, like a test, linter or type checker. An inferential sensor uses a model to judge meaning, like an LLM code reviewer, so it is more flexible but slower and less predictable.
How do I make coding agents reliable without a better model?
Log each failure, then add a guide or sensor that prevents it, starting with cheap deterministic checks. Keep the tool menu small, make errors detailed, and test the harness itself so changes do not quietly break it.
How do I test an agent harness without spending money?
Build a stub mode where workers, reviewers and finalizers return deterministic results. Then you can run the whole pipeline end to end and verify routing, merging and failure handling with no API cost.
Put the loop to work
Every failure you fix makes the next build steadier. You can read about how it gets smarter every build or how a build works in the docs. When you are ready, describe an idea to LoopCodeLab and watch the team plan, build and review it.