Harness Engineering for AI Agents: Building the Runtime That Makes Models Useful

ProductivityUse cases •

Harness Engineering for AI Agents: Building the Runtime That Makes Models Useful

AI coding agents can help with implementation, but their usefulness also depends on the environment: the context, tools, permissions and feedback available to them.

Think of the model as an engine and the harness as the surrounding vehicle: a useful analogy for the runtime’s role.

Harness engineering asks how to build systems that make models consistently useful alongside choosing the model. Harness engineering complements prompt design by supplying the surrounding runtime. This guide breaks down what harness engineering means in practice, the patterns that actually work, and how to build a runtime that scales beyond what any single prompt could achieve.

What Is Harness Engineering?

Harness engineering concerns the runtime around an agent: tool interfaces, context delivery, planning, verification, memory and control. A common shorthand is Agent = Model + Harness; Birgitta Böckeler’s discussion on Martin Fowler’s site uses that framing. It is not an equation credited here to a single inventor.

Agent = Model + Harness

If you're not the model, you're the harness.

The preprint What makes a harness a harness proposes necessary and sufficient conditions for its operational definition. Treat this as a research proposal for distinguishing harnesses from adjacent tools, rather than a universally adopted standard.

The Five Design Principles

1. Expose Everything as Files

The Azure SRE Agent team describes exposing source code, runbooks, query schemas and investigation notes through a filesystem workspace. Their case study supports this design option; it does not establish a universal requirement to replace every API with files.

The lesson: consider exposing reference material as files so the agent can read, search, and navigate it. Keep dedicated tools where they provide needed operations or access controls.

2. Encode Conventions in Skills

A new agent session needs the relevant context loaded or retrieved. Missing intent gets filled with confident guesses — what practitioners call intent debt. Project instructions (such as AGENTS.md or CLAUDE.md) and reusable skills help preserve that context: conventions, build steps, and reasons behind earlier decisions. Project instructions and skills have different loading rules; Claude Code skills are loaded when relevant or explicitly invoked.

Reusable skills are one way to preserve organizational procedures; retrieved documentation, project instructions, and application state can also supply context across runs. The key question isn't "what model should I use?" — it's "what conventions have I encoded so the agent doesn't have to guess?"

3. Separate Creation from Verification

An independent reviewer can catch mistakes an implementer misses. Self-verification can also help when it includes tests and comparison against requirements; neither approach guarantees correctness.

This is the adversarial code review pattern: different agents (or at least different model instances with different instructions) for creation and verification. The checker's default stance is REJECT. Separate review is a useful additional control; it should complement tests and self-verification rather than replace them.

4. Externalize Memory

A model call does not by itself persist application state across sessions. The loop must read from and write to something durable: STATE.md, PLAN.md, IMPLEMENT.md, DOCUMENTATION.md, or a Linear board.

One practical file convention is to separate planning, implementation notes and documentation. The filenames below are a suggested layout, not mandatory OpenAI artifacts:

  • Plan.md — task breakdown, dependencies, current status

  • Implement.md — implementation notes, decisions, trade-offs

  • Document.md — architecture decisions, design rationale, what was learned

These files are the agent's persistent brain. They're what gets read at the start of each run ("where did we leave off?") and written at the end ("here's what I did and what's next").

The agent forgets. The repo doesn't.

5. Design for Cost, Not Just Correctness

Token costs can explode in agent loops. The best practice: triage first — spawn sub-agents only when state says actionable. Empty watchlist → exit before unnecessary model work; choose a budget for your workload.

Three cost factors compound in loops:

Factor

Impact

Cadence

Linear multiplier (5min vs 1d = 288× runs/day)

Sub-agents per run

Each = full model + tool round-trips

Context size

Large repos + full CI logs = expensive triage

A loop with no actionable work can exit before expensive generation. Build explicit early-exit conditions — machine-checkable stop criteria that tell the loop when there's nothing to do.

The Architecture of a Production Harness

Here's what a production-grade harness looks like in 2026:

┌──────────────────────────────────────────────────┐
│  SCHEDULER: cron /loop or GitHub Action          │
└────────────────┬─────────────────────────────────┘
                 │
                 ▼
┌──────────────────────────────────────────────────┐
│  TRIAGE:                                         │
│  Read STATE.md • Discover work • Prioritize      │
│  Output: structured list of actionable items     │
│  Early exit if empty watchlist                   │
└────────────────┬─────────────────────────────────┘
                 │
                 ▼
┌──────────────────────────────────────────────────┐
│  IMPLEMENTER (sub-agent):                        │
│  Read relevant files • Follow skills • Code      │
│  Write to PLAN.md • Create worktree              │
└────────────────┬─────────────────────────────────┘
                 │
                 ▼
┌──────────────────────────────────────────────────┐
│  VERIFIER (sub-agent):                           │
│  Different model • Run tests in isolation        │
│  Default stance: REJECT                          │
└────────────────┬─────────────────────────────────┘
                 │
                 ▼
┌──────────────────────────────────────────────────┐
│  MCP CONNECTORS:                                 │
│  Read/write Linear/Jira • Post to Slack          │
│  Update CI/CD • Query production data            │
└────────────────┬─────────────────────────────────┘
                 │
                 ▼
┌──────────────────────────────────────────────────┐
│  HUMAN GATE:                                     │
│  Escalate on ambiguity • Review complex changes  │
│  Kill switch: loop-pause-all label               │
└──────────────────────────────────────────────────┘

Harness vs. Loop Engineering

The distinction matters:

  • In this guide, harness engineering means configuring the environment, tools, permissions and context; a harness may also support multiple runs.
  • Loop engineering designs systems that prompt agents for you — scheduled, multi-run, with verification and state. The loop runs on a timer, spawns helpers, and feeds itself.

In 2026, the interesting engineering isn't in picking the model. It's in designing the scaffolding around it so it makes better decisions while you sleep.

Getting Started: Your First Harness

You don't need to build a 100-tool empire. Start with the minimum that works:

  1. Write CLAUDE.md. Three files matter most: CLAUDE.md (project conventions), PLAN.md (current task), IMPLEMENT.md (notes and decisions). These are your harness's nervous system.

  2. Expose context as files. Source code is already files. Add runbooks, checklists, past incident notes. Let the agent grep and read instead of calling specialized tools.

  3. Add a verifier. A second agent (different model or instructions) that reviews the implementer's work with a default stance of REJECT. Run it in an isolated worktree.

  4. Add a scheduler. /loop every 5 minutes, a GitHub Action on push, or a cron job. Start with report-only mode — have the loop discover work and report, don't auto-fix yet.

  5. Add a kill switch. A label (loop-pause-all) or environment flag that the human can flip. No harness should run 24/7 with no pause criteria.

L1/L2/L3 Rollout Model

Don't skip levels:

Level

Description

Week 1 behavior

Risk

L1

Report-only. Triage skill returns structured output.

L1 report only

Low

L2

Assisted. Loop can suggest fixes, spawn sub-agents.

L2 report + suggestions

Medium

L3

Unattended. Auto-fix, auto-PR, self-healing.

Disabled until L1 quality proven

High

Roll out from L1 to L2 and then L3 only against task-specific evaluation and oversight criteria; one error-free week is not a general safety threshold.

Anti-Patterns That Kill Harnesses

Failure modes to consider when designing a harness:

  1. Same agent implements and verifies — confirmation bias; weak tests get rubber-stamped

  2. No attempt cap — infinite fix loops, token burn, wrong fixes merged

  3. Vague triage output — "paragraphs of narrative" instead of structured sections

  4. L3 before L1 quality — auto-fix and auto-PR on day one

  5. Shared state without schema — three loops appending to one unstructured STATE.md

  6. MCP with write-everything scope — loop can merge PRs, post to Slack, edit tickets on day one

  7. No kill switch — loop runs 24/7 with no pause criteria

  8. Auto-merge without allowlist — verifier passed, merge it (security/business-logic bugs pass weak verifiers)

The Harness Engineering Resource Map

Resource

URL

Purpose

OpenAI Harness Engineering

https://openai.com/index/harness-engineering/

Engineering case study: repository context and verification loops

Anthropic Building Effective Agents

https://www.anthropic.com/research/building-effective-agents

Three principles: simplicity, transparency, and a carefully designed agent-computer interface

Martin Fowler Harness Engineering

https://martinfowler.com/articles/harness-engineering.html

Patterns and architecture for production harnesses

Microsoft Azure SRE Agent

https://techcommunity.microsoft.com/blog/appsonazureblog/the-agent-that-investigates-itself/4500073

Case study: filesystem-as-interface

Awesome Harness Engineering

https://github.com/ai-boost/awesome-harness-engineering

Curated resource list; consult linked primary sources

What This Means for You

The era of "write a good prompt and get a good result" is ending. What replaces it isn't more prompting — it's system building.

Three things matter now:

  1. Externalize everything. Persist needed state in files or another durable store the agent can access on the next run.

  2. Separate creation from verification. Use independent review alongside the implementer’s tests and self-checks.

  3. Design for cost, not just correctness. Every loop needs a token budget and a kill switch.

Harness engineering is learning to stop prompting and start designing.

Coding tools expose different combinations of context, tool use, persistence, verification and control. Check the capabilities and limits of the version you use; the design choices around those facilities still matter.

Related reading: Loop Engineering patterns · A2A Protocol · Awesome Harness Engineering


Related Posts

Site Logo Artifilog

Artifilog is a creative blog that explores the intersection of art, design, and technology. It serves as a hub for inspiration, featuring insights, tutorials, and resources to fuel creativity and innovation.

Categories