Harness Engineering for AI Agents: Building the Runtime That Makes Models Useful

AI coding agents can help with implementation, but their usefulness also depends on the environment: the context, tools, permissions and feedback available to them.
Think of the model as an engine and the harness as the surrounding vehicle: a useful analogy for the runtime’s role.
Harness engineering asks how to build systems that make models consistently useful alongside choosing the model. Harness engineering complements prompt design by supplying the surrounding runtime. This guide breaks down what harness engineering means in practice, the patterns that actually work, and how to build a runtime that scales beyond what any single prompt could achieve.
What Is Harness Engineering?
Harness engineering concerns the runtime around an agent: tool interfaces, context delivery, planning, verification, memory and control. A common shorthand is Agent = Model + Harness; Birgitta Böckeler’s discussion on Martin Fowler’s site uses that framing. It is not an equation credited here to a single inventor.
Agent = Model + Harness
If you're not the model, you're the harness.
The preprint What makes a harness a harness proposes necessary and sufficient conditions for its operational definition. Treat this as a research proposal for distinguishing harnesses from adjacent tools, rather than a universally adopted standard.
The Five Design Principles
1. Expose Everything as Files
The Azure SRE Agent team describes exposing source code, runbooks, query schemas and investigation notes through a filesystem workspace. Their case study supports this design option; it does not establish a universal requirement to replace every API with files.
The lesson: consider exposing reference material as files so the agent can read, search, and navigate it. Keep dedicated tools where they provide needed operations or access controls.
2. Encode Conventions in Skills
A new agent session needs the relevant context loaded or retrieved. Missing intent gets filled with confident guesses — what practitioners call intent debt. Project instructions (such as AGENTS.md or CLAUDE.md) and reusable skills help preserve that context: conventions, build steps, and reasons behind earlier decisions. Project instructions and skills have different loading rules; Claude Code skills are loaded when relevant or explicitly invoked.
Reusable skills are one way to preserve organizational procedures; retrieved documentation, project instructions, and application state can also supply context across runs. The key question isn't "what model should I use?" — it's "what conventions have I encoded so the agent doesn't have to guess?"
3. Separate Creation from Verification
An independent reviewer can catch mistakes an implementer misses. Self-verification can also help when it includes tests and comparison against requirements; neither approach guarantees correctness.
This is the adversarial code review pattern: different agents (or at least different model instances with different instructions) for creation and verification. The checker's default stance is REJECT. Separate review is a useful additional control; it should complement tests and self-verification rather than replace them.
4. Externalize Memory
A model call does not by itself persist application state across sessions. The loop must read from and write to something durable: STATE.md, PLAN.md, IMPLEMENT.md, DOCUMENTATION.md, or a Linear board.
One practical file convention is to separate planning, implementation notes and documentation. The filenames below are a suggested layout, not mandatory OpenAI artifacts:
Plan.md — task breakdown, dependencies, current status
Implement.md — implementation notes, decisions, trade-offs
Document.md — architecture decisions, design rationale, what was learned
These files are the agent's persistent brain. They're what gets read at the start of each run ("where did we leave off?") and written at the end ("here's what I did and what's next").
The agent forgets. The repo doesn't.
5. Design for Cost, Not Just Correctness
Token costs can explode in agent loops. The best practice: triage first — spawn sub-agents only when state says actionable. Empty watchlist → exit before unnecessary model work; choose a budget for your workload.
Three cost factors compound in loops:
Factor | Impact |
|---|---|
Cadence | Linear multiplier (5min vs 1d = 288× runs/day) |
Sub-agents per run | Each = full model + tool round-trips |
Context size | Large repos + full CI logs = expensive triage |
A loop with no actionable work can exit before expensive generation. Build explicit early-exit conditions — machine-checkable stop criteria that tell the loop when there's nothing to do.
The Architecture of a Production Harness
Here's what a production-grade harness looks like in 2026:
┌──────────────────────────────────────────────────┐
│ SCHEDULER: cron /loop or GitHub Action │
└────────────────┬─────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────┐
│ TRIAGE: │
│ Read STATE.md • Discover work • Prioritize │
│ Output: structured list of actionable items │
│ Early exit if empty watchlist │
└────────────────┬─────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────┐
│ IMPLEMENTER (sub-agent): │
│ Read relevant files • Follow skills • Code │
│ Write to PLAN.md • Create worktree │
└────────────────┬─────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────┐
│ VERIFIER (sub-agent): │
│ Different model • Run tests in isolation │
│ Default stance: REJECT │
└────────────────┬─────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────┐
│ MCP CONNECTORS: │
│ Read/write Linear/Jira • Post to Slack │
│ Update CI/CD • Query production data │
└────────────────┬─────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────┐
│ HUMAN GATE: │
│ Escalate on ambiguity • Review complex changes │
│ Kill switch: loop-pause-all label │
└──────────────────────────────────────────────────┘
Harness vs. Loop Engineering
The distinction matters:
- In this guide, harness engineering means configuring the environment, tools, permissions and context; a harness may also support multiple runs.
Loop engineering designs systems that prompt agents for you — scheduled, multi-run, with verification and state. The loop runs on a timer, spawns helpers, and feeds itself.
In 2026, the interesting engineering isn't in picking the model. It's in designing the scaffolding around it so it makes better decisions while you sleep.
Getting Started: Your First Harness
You don't need to build a 100-tool empire. Start with the minimum that works:
Write CLAUDE.md. Three files matter most: CLAUDE.md (project conventions), PLAN.md (current task), IMPLEMENT.md (notes and decisions). These are your harness's nervous system.
Expose context as files. Source code is already files. Add runbooks, checklists, past incident notes. Let the agent
grepandreadinstead of calling specialized tools.Add a verifier. A second agent (different model or instructions) that reviews the implementer's work with a default stance of REJECT. Run it in an isolated worktree.
Add a scheduler.
/loopevery 5 minutes, a GitHub Action on push, or a cron job. Start with report-only mode — have the loop discover work and report, don't auto-fix yet.Add a kill switch. A label (
loop-pause-all) or environment flag that the human can flip. No harness should run 24/7 with no pause criteria.
L1/L2/L3 Rollout Model
Don't skip levels:
Level | Description | Week 1 behavior | Risk |
|---|---|---|---|
L1 | Report-only. Triage skill returns structured output. | L1 report only | Low |
L2 | Assisted. Loop can suggest fixes, spawn sub-agents. | L2 report + suggestions | Medium |
L3 | Unattended. Auto-fix, auto-PR, self-healing. | Disabled until L1 quality proven | High |
Roll out from L1 to L2 and then L3 only against task-specific evaluation and oversight criteria; one error-free week is not a general safety threshold.
Anti-Patterns That Kill Harnesses
Failure modes to consider when designing a harness:
Same agent implements and verifies — confirmation bias; weak tests get rubber-stamped
No attempt cap — infinite fix loops, token burn, wrong fixes merged
Vague triage output — "paragraphs of narrative" instead of structured sections
L3 before L1 quality — auto-fix and auto-PR on day one
Shared state without schema — three loops appending to one unstructured STATE.md
MCP with write-everything scope — loop can merge PRs, post to Slack, edit tickets on day one
No kill switch — loop runs 24/7 with no pause criteria
Auto-merge without allowlist — verifier passed, merge it (security/business-logic bugs pass weak verifiers)
The Harness Engineering Resource Map
Resource | URL | Purpose |
|---|---|---|
OpenAI Harness Engineering | https://openai.com/index/harness-engineering/ | Engineering case study: repository context and verification loops |
Anthropic Building Effective Agents | https://www.anthropic.com/research/building-effective-agents | Three principles: simplicity, transparency, and a carefully designed agent-computer interface |
Martin Fowler Harness Engineering | https://martinfowler.com/articles/harness-engineering.html | Patterns and architecture for production harnesses |
Microsoft Azure SRE Agent | https://techcommunity.microsoft.com/blog/appsonazureblog/the-agent-that-investigates-itself/4500073 | Case study: filesystem-as-interface |
Awesome Harness Engineering | https://github.com/ai-boost/awesome-harness-engineering | Curated resource list; consult linked primary sources |
What This Means for You
The era of "write a good prompt and get a good result" is ending. What replaces it isn't more prompting — it's system building.
Three things matter now:
Externalize everything. Persist needed state in files or another durable store the agent can access on the next run.
Separate creation from verification. Use independent review alongside the implementer’s tests and self-checks.
Design for cost, not just correctness. Every loop needs a token budget and a kill switch.
Harness engineering is learning to stop prompting and start designing.
Coding tools expose different combinations of context, tool use, persistence, verification and control. Check the capabilities and limits of the version you use; the design choices around those facilities still matter.
Related reading: Loop Engineering patterns · A2A Protocol · Awesome Harness Engineering