Loop Engineering: Building Autonomous Agent Systems That Work While You Sleep

Agent systems can discover work, write code, review it, test it, and open a PR on a schedule. Addy Osmani describes one such loop, while stressing that engineers still need to verify what it produces.
Design the system that prompts the agent, and keep verification part of that system. — a practical takeaway from Addy Osmani’s June 2026 essay
Welcome to loop engineering: the discipline of designing systems that prompt agents for you. Where harness engineering configures the environment for a single agent run, loop engineering builds the scheduler, the state tracker, the verifier, and the escalation path that runs long after you've closed your laptop.
This guide walks through the five primitives of loop engineering, a suggested L1/L2/L3 rollout model, and the anti-patterns that can make "autonomous automation" expensive and unreliable.
What Loop Engineering Actually Means
Loop engineering is the practice of building systems — schedulers, state managers, verifiers, and escalation paths — that run agentic workflows on a timer, with verification and memory, replacing the developer as the prompter.
One way to organize the progression, building on Addy Osmani’s June 2026 discussion, is:
Layer | Question | Tool |
|---|---|---|
Prompt Engineering | "How do I phrase this to get what I want?" | Chatbot, system messages |
Harness Engineering | "What do I need to give the model to succeed?" | Tool config, context, memory |
Loop Engineering | "How do I build the system that prompts agents for me?" | Schedulers, verifiers, state files, hooks |
Loop engineering sits one level above harness engineering. The harness runs inside a single session; the loop runs on a timer, spawns helpers, and feeds itself.
An exploratory arXiv study, revised in August 2026, confirmed autonomous agent loops operating in 217 of 256 repositories matched by heuristics across 36,710 repos mined. This was not a representative estimate of adoption across all software projects. Its review of practitioner literature identified recurring components: triggered agent runs, machine-checkable stop conditions, persistent state files, verifier sub-agents, token budgets, and defined escalation points.
The Five Primitives + Memory
Loop engineering decomposes into six components. Use them as a design checklist; a small loop may not need a worktree, connectors, or sub-agents.
Primitive 1: Automations / Scheduling (The Heartbeat)
Without a recurring trigger or a feedback cycle, it may be just a one-off agent run. The scheduler is the heartbeat that triggers the loop on a cadence:
/loopfor in-session intervals; external schedulers for persistent recurring runsGitHub Actions with repository dispatch triggers
CI workflows with scheduled triggers (
cron:in.github/workflows/)Webhook-triggered runs (push events, issue creation, etc.)
The critical decision: how often to run vs. how much each run costs. Measure cost per run using your model, context, tool usage, and billing arrangement. A five-minute cadence creates far more opportunities to spend than a daily run.
Primitive 2: Worktrees (Parallelism Without Chaos)
The Git worktree documentation describes separate working directories attached to one repository, with shared repository data. One worktree per attempt, tracked in a manifest, swept on reject or escalation.
Why not just branches in one checkout? Because isolation matters:
- Separate worktrees keep ordinary checkout edits apart when each loop stays within its assigned directory; they do not prevent cross-directory writes
- Each new worktree checks out the chosen commit; shared services and external state need separate reset or isolation
- Failed attempts can be removed with git worktree remove; external side effects still need cleanup
- In this workflow, the main branch changes only after the configured approval and merge steps
# One worktree per attempt
attempt="loop-attempt-$(date +%s)"
git worktree add -b "loop/$attempt" "../$attempt" HEAD
# After reviewing or saving any work you need:
git worktree remove "../$attempt"
Primitive 3: Skills (The Persistent Memory of Intent)
Skills encode project conventions, "we don't do it this way because of X incident," build/test/lint commands, and review standards. Without durable project instructions, loops may need to rediscover conventions; skill files are one way to preserve them.
A well-maintained CLAUDE.md, AGENTS.md, or skill file answers:
- What are the coding conventions on this project?
- What build/test/lint commands do I run?
- What review standards must changes meet?
- What incidents happened before that I should avoid?
Intent debt is what happens when the loop doesn't have skills. A fresh session may lack project context unless the tool restores history or loads durable instructions.
Primitive 4: Plugins & Connectors (MCP)
A loop that can only read the filesystem is limited. Some loops need to read or update issue trackers, query databases, or create branches and PRs; others can stay entirely within a repository. This is where MCP (Model Context Protocol) and other connectors bridge the gap.
The security rule: scope MCP permissions narrowly. A connector exposing write tools with broad credentials may let the loop merge PRs, post to Slack, and edit tickets — that's anti-pattern #6. Start read-only and escalate permissions only as quality proves itself.
Primitive 5: Sub-agents — The Maker/Checker Split
This is the single most important structural pattern. A separate review pass can challenge assumptions in generated code. Independence is useful, but neither the maker nor the checker is automatically reliable. The implementer knows what they intended; the checker evaluates what actually shipped.
The pattern:
- Session 1 (Implementer): writes code, runs tests, creates a PR
- Session 2 (Verifier): reviews the PR, runs expanded test suite, checks for edge cases — default stance: REJECT
- Session 3 (Gatekeeper): final human approval, or auto-merge with an allowlist
Each session can use different model instances, different instructions, and different tools. Different instructions and context can help the verifier challenge the implementation, but shared model biases can remain.
Primitive 6: Memory / State (What Outlives Sessions)
A fresh model invocation does not inherently retain previous session state. The surrounding tool may restore conversation history or load persistent memory; the loop needs durable state it can recover.
This takes the form of:
- STATE.md — the single source of truth for what's done, what's next, current assignees
- A Linear/Jira board — tasks persist across runs
- A database row — machine-checkable state
The agent forgets. The repo doesn't.
The arXiv study found that its sampled repositories committed loop configuration but almost none committed the state files described in practitioner discussions. Runtime state can live in a durable database or external service too. Version suitable project state when useful; keep secrets and sensitive runtime data out of Git.
The L1/L2/L3 Rollout Model
The following is a suggested rollout progression, not a production standard. Increase autonomy only when evaluation and recovery controls support it.
Level | Description | Week 1 behavior | Risk |
|---|---|---|---|
L1 | Report-only. Triage skill returns structured output. | L1 report only | Low |
L2 | Assisted. Loop can suggest fixes, spawn sub-agents. | L2 report + suggestions | Medium |
L3 | Unattended. Auto-fix, auto-PR, self-healing. | Disabled until L1 quality proven | High |
Roll out: L1 report → L2 assisted → L3 unattended, using representative evaluations and explicit risk criteria rather than elapsed time alone.
L1: Report-Only
Start here. The loop runs on a schedule, discovers work, and reports — no actions taken. In this example, the triage pass reads STATE.md, scans for actionable items, and outputs a structured list:
## Actionable
- [HIGH] Unassigned PR #1247 — fix looks straightforward
- [MED] CI failure in deploy-workflow — needs investigation
## Not Actionable
- Closed PRs, resolved tickets
Reporting avoids implementation work, but measure its actual usage. If the watchlist is empty, the loop should exit before launching unnecessary worker calls. That's the early-exit condition — the single most important design principle for cost control.
L2: Assisted Mode
Once L1 reports are reliable across representative cases, consider enabling L2. The loop can now suggest fixes and spawn sub-agents for investigation — but still requires human approval for any write operation.
L2 output looks like:
For example, PR #1247 might have a missing await. An illustrative suggested fix:
- const data = fetchData();
+ const data = await fetchData();
Approve with your workflow’s apply command, such as /apply --pr=1247, if you have implemented one.
L3: Unattended (After Proof)
Enable L3 only after representative evaluations and recovery drills meet your risk criteria. The loop now auto-fixes, auto-PRs, and self-heals. L3 is where the compound benefits kick in: the loop can run unattended on available infrastructure and cover repetitive work across repositories, while people still maintain and review it.
But L3 is also where things go wrong fastest. Anti-patterns #4 (L3 before L1 quality) and #9 (auto-merge without allowlist) are two risks to address before trusting unattended changes.
Anti-Patterns That Kill Loops
These are the critical failures that turn "autonomous automation" into a nightmare. Every production loop needs explicit guards against these.
1. Same Agent Implements and Verifies
The implementer knows what they intended; the checker evaluates what was actually built. Using the same context for implementation and review can leave assumptions unchallenged; separate review and independent tests help.
Guard: Always use a separate verifier session with different model instructions, defaulting to REJECT.
2. No Attempt Cap — Infinite Fix Loops
Without a limit on how many times the loop tries to fix something, a subtly broken test can trigger infinite fix-merge-fix cycles. Each cycle burns tokens, each fix may introduce a new regression.
Guard: Set a maximum attempt count (e.g., 3 fixes per issue). If the verifier rejects after 3 attempts, escalate to human with full context.
3. Vague Triage Output
Unstructured narrative can make triage output difficult to process reliably. The loop needs machine-parseable output: clear priorities, clear boundaries between actionable and not-actionable, clear stop conditions.
Guard: Enforce a strict output schema. If the triage can't produce structured output, the loop doesn't know when to stop.
4. L3 Before L1 Quality
Enabling auto-fix and auto-PR on day one, before the report-only mode has proven reliable, is how you get a loop that confidently merges broken code at 3 AM.
Guard: Choose a trial period long enough to cover representative tasks and failures; elapsed weeks alone are not an acceptance test.
5. Shared State Without Schema
When several loops update shared state without coordination, their writes can conflict. A schema helps with format, while locking or transactional updates help with concurrency. Field names drift, formats diverge, and the state file becomes unreadable.
Guard: Define a state schema. Each field has a type and a producer/consumer contract. The loop validates state before reading it.
6. MCP with Write-Everything Scope
The loop can merge PRs, post to Slack, edit tickets, and deploy — all from day one. One misconfiguration and the loop wipes a production database.
Guard: Start read-only MCP permissions. Escalate to write permissions only after L2 has proven reliable with human-in-the-loop.
7. No Kill Switch
A loop that runs 24/7 with no pause criteria will run forever — even if it's broken, even if the repo is gone, even if the world ends.
Guard: A loop-pause-all label on GitHub, an env var LOOP_ENABLED=false, or a STOP file in the repo root. An authorized operator can flip it. The loop checks for it every run.
8. Fixing Flakes with Code
When CI tests are flaky, a naive loop tries to "fix" them by changing application code. This corrupts the codebase — quarantine or retry may contain the symptom, but a fix can also require changes to the test, infrastructure, or application code.
Guard: Distinguish "flake" from "failure." Treat intermittently passing and failing tests as candidates for flake investigation. Determine the root cause before classifying or changing them. Investigate the cause before deciding whether to quarantine, retry, or change code.
9. Auto-Merge Without Allowlist
The verifier passed → merge it. Except security bugs, business-logic bugs, and infrastructure issues all pass weak verifiers. A correctly enforced empty allowlist permits no automatic merges.
Guard: Auto-merge only for changes that touch allowlisted files (e.g., docs/, README.md, tests/e2e/). Everything else requires human approval. The allowlist starts empty and grows as trust builds.
10. No Run Log
If the loop records only current state, reconstructing why a run went wrong becomes harder. "Why does STATE.md say PR #1247 is fixed but it's still open?"
Guard: Append to a run log on every execution: logs/loop-{date}-{time}.log. Include: triage result, actions taken, verification outcome, and any exceptions.
Token Cost Management
Token usage depends on cadence, worker calls, and context size. Estimate the total from observed runs and your billing arrangement.
| Factor | Impact |
|---|---|
| Cadence | Linear multiplier (5m vs 1d = 288× runs/day) |
| Sub-agents per run | Each spawns full model + tool round-trips |
| Context size | Large repos + full CI logs = expensive triage |
Illustrative Budget Arithmetic
For arithmetic only, suppose a triage run uses 50,000 tokens. This is a hypothetical budget assumption, not a measured benchmark; your runs may use much less or much more.
| Loop | Cadence | Runs/day | Rough daily tokens |
|---|---|---|---|
| Daily triage (report only) | 1d | 1 | ~50k |
| CI sweeper (light) | 15m | 96 | 4.8M under the 50,000-token assumption |
| PR babysitter | 5m | 288 | 14.4M under the same assumption; use early exit |
The Early-Exit Rule
The single most important cost optimization: spawn sub-agents only when state says actionable.
# Cheap triage — exit if nothing to do
watchlist = read_state("actionable_items")
if len(watchlist) == 0:
log("Empty watchlist — exiting")
sys.exit(0) # Avoid starting the implementer and verifier
Empty watchlist → exit before launching workers. If the watchlist isn't empty, spawn the implementer. If the implementer finds nothing actionable, the verifier doesn't run. Each layer gates the next.
Verifier Model Selection
For unattended L3 loops, run the verifier on a stronger model. Weigh the impact of a missed bug against the measured cost and effectiveness of the verifier. Select the implementation and verification models using task-specific evaluations and observed cost; a stronger verifier is useful only when it improves error detection.
From Theory to Practice: Building Your First Loop
You don't need to start with a $5,000/month infrastructure pipeline. Start with the minimum that works and grow.
Step 1: Define One Problem
Pick one specific problem that wastes your time:
- "Review my PRs for missing tests"
- "Check if my CI failures are flaky"
- "Close stale issues older than 90 days"
Not "build a general automation system." One problem, one loop.
Step 2: Write STATE.md
Your loop needs a memory. Start with the simplest possible state file:
# STATE.md
## Current Tasks
- [ ] Review PR #1247 for missing tests
## History
- 2026-09-22: Reviewed PR #1246 — all tests present ✓
The loop reads this, discovers PR #1247 needs review, does the work, updates STATE.md.
Step 3: Write the Triage Skill
Create CLAUDE.md or a skill file with:
- What to look for (e.g., "check for test files matching changed source files")
- What action to take (e.g., "comment on the PR with findings")
- What the stop condition is (e.g., "no unreviewed PRs in STATE.md")
- What to log (e.g., "append to logs/triage-YYYY-MM-DD.log")
Step 4: Schedule It
Start with L1 — report only, no actions:
# .github/workflows/loop-triage.yml
name: Triage
on:
schedule:
- cron: '0 9 * * 1-5' # Weekdays at 09:00 UTC
workflow_dispatch:
permissions:
contents: read
id-token: write
jobs:
triage:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
claude_args: '--allowedTools "Read,Glob,Grep"'
prompt: |
Read STATE.md and report any unreviewed PRs with missing tests.
Report only; do not modify files or external systems.
This example reads local repository files. Configure the API secret and the action’s authentication before running it; add narrowly scoped GitHub read tools if your triage also needs live PR data. See the official GitHub Actions setup guide.
Step 5: Evaluate L1 Before Expanding It
Watch the reports. Are they accurate? Are false positives caught? Are real issues surfaced? Only move to L2 when you trust the triage output.
Step 6: Add the Verifier
Once triage is reliable, add a second claude invocation as the verifier:
- name: Verifier
uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
claude_args: '--allowedTools "Read,Glob,Grep"'
prompt: |
Review the implementer’s saved diff and report against the task criteria.
Report findings only. Do not edit files or merge changes.
Save the implementation diff and report in the checked-out workspace before this step. Add specific read or test tools only when the review needs them.
The verifier starts a separate invocation with its own instructions, but this workflow uses the same runner and workspace. Use a separate sandbox if the review needs filesystem or process isolation.
Step 7: Graduated Escalation
After L2 meets your evaluation and recovery criteria:
- Enable L3 for low-risk files only (docs, tests)
- Require human approval for everything else
- Consider expanding the allowlist only when representative results support it
What This Means for You
Loop engineering isn't about replacing developers with bots. It's about building systems that handle the repetitive, high-signal work so humans can focus on the creative and strategic.
Three principles matter now:
Start small, graduate slowly. One problem, one loop, with evaluation before increasing autonomy. L1 → L2 → L3.
Separate creation from verification. The implementer cannot grade its own homework.
Always design for cost — and failure. Every loop needs a kill switch, an attempt cap, and structured state.
Loop engineering is learning to stop doing the work yourself and start designing the system that does it for you.
Osmani describes variations of these primitives in Claude Code and Codex; inspect the actual scheduling, isolation, and review capabilities of the tool you choose. The interesting engineering isn't in which tool you use. It's in how much judgment you encode into the system so it makes the right call while you sleep.
