Skip to main content

Harness Engineering in Practice — The Environment You Build Around Your Coding Agent

· 12 min read
Bruno Carneiro
Fundador da @TautornTech
Harness Engineering in practice

In the post about prompt, context and harness engineering I split the three layers apart and talked about the harness from the point of view of someone building an agent: idempotency, circuit breakers, approval gates, audit trails.

But most of us aren't building agents. We're using one: Claude Code, OpenCode, Cursor, Codex. Which raises the obvious question: if the tool already ships with a harness, what's left for me to do?

A lot. And it's exactly the part that changes results the most day to day.

Agent = model + harness​

The definition that stuck in 2026 is simple: an agent is a model plus a harness. The harness is everything that isn't the model: the loop, the tools, context management, permissions, subagents, extension points.

Birgitta Böckeler, in her article on Martin Fowler's site, proposes a mental model I find very useful: concentric circles.

┌──────────────────────────────────────────────┐
│ User harness (you) │
│ AGENTS.md, skills, hooks, tests, lint, │
│ subagents, architecture rules │
│ ┌──────────────────────────────────────┐ │
│ │ Builder harness (the tool) │ │
│ │ system prompt, loop, tools, search │ │
│ │ ┌──────────────────────────────┐ │ │
│ │ │ Model │ │ │
│ │ └──────────────────────────────┘ │ │
│ └──────────────────────────────────────┘ │
└──────────────────────────────────────────────┘

You don't control the middle circle (or barely do). The outer one is yours. And it decides whether the agent works like someone who knows the project or like an intern on their first day.

Harness engineering, for agent users, is designing that outer circle.

Guides and sensors​

The core idea in Böckeler's article is that a harness has two kinds of controls:

  • Guides (feedforward): act before the agent does something. They raise the odds it gets it right the first time. AGENTS.md, conventions, skills, examples, templates.
  • Sensors (feedback): act after. They observe the result and send a signal back so the agent can correct itself. Tests, type-checking, lint, static analysis, review.

And the point many people miss: one without the other doesn't work.

Sensors only: the agent repeats the same mistakes every time and fixes them afterwards. Guides only: the agent follows rules but never finds out whether they worked.

She also splits controls into computational (deterministic, fast, cheap: linter, compiler, tests) and inferential (semantic, done by another model: agent review, LLM-as-judge). Computational controls are reliable and cost milliseconds. Inferential ones catch things rules can't, but they're more expensive and probabilistic.

Rule of thumb: anything you can check deterministically, check deterministically. Leave the inferential layer for what's left.

Guides: give it a map, not a manual​

A short AGENTS.md that points elsewhere​

The most common mistake I see is the giant AGENTS.md (or CLAUDE.md). Everything anyone remembered to write ended up there. The result is a file that eats context in every session and that the model starts ignoring.

The OpenAI team that built an entire product with Codex alone reached the same conclusion: they replaced the monolithic file with an AGENTS.md of roughly 100 lines that works as a table of contents. It says where to look, and the real knowledge lives in docs/.

# AGENTS.md

## Commands
- Install: `pnpm i`
- Tests: `pnpm test` (run before saying you're done)
- Types: `pnpm typecheck`
- Lint: `pnpm lint`

## Where things are
- Architecture and dependency rules: `docs/architecture.md`
- API and validation patterns: `docs/api.md`
- Plans in progress: `docs/plans/`
- Past decisions (ADRs): `docs/adr/`

## Rules that don't change
- Never edit existing migrations, only create new ones
- UI never imports from `src/repositories/` directly
- Don't use `any`; use `unknown` and validate

Short, verifiable, with pointers. The agent reads docs/api.md when it needs to touch the API, not in every session. That's progressive disclosure applied to your repository.

What the agent can't see doesn't exist​

That line from the OpenAI post became my motto: from the agent's point of view, anything not accessible in the repository doesn't exist. The architecture decision buried in a Slack thread, the convention "everybody knows", the reason behind that weird if: none of it exists for the agent.

If the knowledge matters, it needs to be in the repo, as text, somewhere AGENTS.md points to.

Skills for repeated processes​

Skills (in Claude Code, OpenCode and others) are folders with instructions for a specific kind of task: creating an endpoint, writing a migration, cutting a release. The difference from AGENTS.md is that a skill only enters the context when it's relevant.

If you're explaining the same thing to the agent for the third time, it should become a skill.

Plans as artifacts​

For big changes, ask for a plan first and save the plan in the repo (docs/plans/auth-migration.md). The plan becomes a guide for future sessions, survives /clear, and makes it clear what was decided. OpenAI treats plans as first-class artifacts, versioned alongside the code.

Sensors: let the agent find out on its own that it got it wrong​

This is, for me, the part with the biggest payoff. In the Claude Code post I talked about giving a verification criterion in the prompt ("run npm test and fix the failures"). Sensors are the next step: the criterion no longer depends on you remembering to ask.

Hooks running sensors automatically​

In Claude Code, a PostToolUse hook runs after every edit. If the script exits with code 2, whatever it writes to stderr goes back to Claude as feedback, and it fixes the problem right away.

// .claude/settings.json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write|MultiEdit",
"hooks": [
{ "type": "command", "command": ".claude/hooks/sensors.sh" }
]
}
]
}
}
#!/bin/bash
# .claude/hooks/sensors.sh

if ! OUT=$(pnpm -s typecheck 2>&1); then
echo "Type errors after your edit. Fix them before continuing:" >&2
echo "$OUT" | head -40 >&2
exit 2
fi

if ! OUT=$(pnpm -s lint 2>&1); then
echo "Lint failed. Fix it before continuing:" >&2
echo "$OUT" | head -40 >&2
exit 2
fi

exit 0

The head -40 is deliberate: tool output is context too, and 2,000 lines of errors pollute more than they help.

tip

If type-checking the whole project is slow, run it only on changed files or leave the heavy sensor for pre-commit. A sensor that takes 40 seconds on every edit becomes a sensor someone turns off.

Error messages written for the agent​

This detail makes a huge difference. A sensor is much more useful when the error message already says how to fix it. OpenAI uses custom linters whose messages inject remediation instructions straight into the agent's context.

You can do this with what you already have. For example, with ESLint's no-restricted-imports:

// eslint.config.js
export default [
{
files: ['src/ui/**/*.{ts,tsx}'],
rules: {
'no-restricted-imports': ['error', {
patterns: [{
group: ['**/repositories/*'],
message:
'UI does not access repositories directly. Use the matching service in ' +
'src/services/ (see docs/architecture.md, "Layers" section).',
}],
}],
},
},
]

The agent makes the mistake, lint flags it, and the message itself explains the right path. You didn't have to be there.

Architecture rules as tests​

An architecture rule that only exists in a document is a suggestion. To become a rule, it needs a sensor. Some that work well:

  • Circular imports: madge --circular in CI (I wrote about it here).
  • Layer boundaries: no-restricted-imports, dependency-cruiser or structural tests.
  • Simple limits: max file size, no console.log, mandatory structured logging.

Böckeler's article calls this an architecture fitness harness: functions that define and check characteristics of the architecture.

Inferential sensors: one agent reviewing another​

For what rules can't catch (overengineering, bad naming, logic that doesn't make sense for the domain), use review by another agent, ideally with clean context that didn't watch the implementation happen. A review subagent or the GitHub Action I showed in the Claude Code post work well here.

Just remember it's probabilistic. It filters, it doesn't guarantee.

Make the application legible to the agent​

Another strong lesson from the OpenAI write-up: the bottleneck stopped being the model and became what the agent can observe. They made the application boot in isolation in each git worktree, gave access to local logs and metrics, and integrated Chrome DevTools so the agent could actually see the UI.

You don't need to go that far, but you can start small:

  • A single command to run the app locally (pnpm dev), documented in AGENTS.md.
  • Readable logs in the terminal, not only in an external service.
  • Browser access (Playwright MCP, Chrome DevTools or the tool's own browser) to validate UI changes.
  • A reproducible database seed to reproduce bugs.

The more the agent can reproduce and verify on its own, the less you become the only feedback loop.

Subagents and the right model for each task​

Not every task needs the most expensive model. Exploring the codebase, generating a session title, summarizing history, compacting context: all of that works fine with a smaller, faster model.

In OpenCode, for example, you can set the model per agent in opencode.json:

{
"$schema": "https://opencode.ai/config.json",
"model": "anthropic/claude-sonnet-5-5",
"small_model": "anthropic/claude-haiku-4-5",
"agent": {
"explore": { "model": "anthropic/claude-haiku-4-5" },
"title": { "model": "anthropic/claude-haiku-4-5" },
"summary": { "model": "anthropic/claude-haiku-4-5" },
"compaction": { "model": "anthropic/claude-haiku-4-5" }
}
}

explore is the read-and-search subagent. title, summary and compaction are hidden system agents that generate titles, summaries, and compact the context as the conversation grows. Check the available IDs with opencode models and the agents documentation for your version, since the structure changed in V2.

In Claude Code, the same idea becomes a file in .claude/agents/:

---
name: explorer
description: Investigates the codebase and returns a short summary. Use before large changes.
tools: Read, Grep, Glob
model: haiku
---

You investigate code and never edit anything.
Answer in at most 400 words: relevant files, main flow,
risk points. Cite file paths.

The gain is twofold: lower cost and clean context in the main session, because the subagent reads 30 files and returns only the summary.

The golden rule: an agent mistake becomes a harness improvement​

If you take only one thing from this article, take this one.

When the agent makes a mistake, the natural reaction is to fix the code and move on. That works, but the same mistake comes back tomorrow, in another session, with someone else on the team.

Böckeler calls this the steering loop: whenever a problem happens more than once, you improve a guide or a sensor so it becomes less likely, or impossible.

The agent...Fix the code and...
Imported the repository directly in the UIAdd no-restricted-imports with a remediation message
Forgot to run the testsPostToolUse hook or an explicit line in AGENTS.md
Edited an existing migrationA deny rule in settings + a line in AGENTS.md
Created an import cyclemadge --circular in pre-commit / CI
Got the same setup wrong for the third timeTurn it into a skill
Didn't know about an architecture decisionWrite the ADR in docs/adr/ and point to it in AGENTS.md

And the best part: the agent itself helps build the harness. Ask it to write the lint rule, the structural test, or the ADR draft. You review.

The harness rots too​

A harness is code. And code without maintenance rots.

An AGENTS.md with commands that no longer exist, rules that contradict each other, a skill describing an old flow, a sensor everyone ignores because it always fails. All of it makes the agent worse without anyone noticing why.

The OpenAI team calls the fix garbage collection: recurring tasks that look for deviations from the project's principles and open small corrective PRs, instead of letting debt pile up. For a regular team, a periodic review of AGENTS.md, skills and hooks already covers most of it.

Where to start​

If your project today has only a generic CLAUDE.md, this is the order I'd suggest:

  1. A short AGENTS.md with commands, where things are, and the few rules that don't change.
  2. A fast sensor in the loop: type-check + lint via a hook after every edit.
  3. Error messages with remediation instructions on the rules that break most often.
  4. One architecture rule as a test (start with circular imports or the UI → repository boundary).
  5. An exploration subagent on a smaller model.
  6. The habit: every repeated mistake becomes a guide or a sensor.
warning

None of this replaces reviewing what goes into the repository. A good harness doesn't remove the human from the process; it directs your attention to where it really matters.

Conclusion​

Prompt engineering is about what you ask. Context engineering is about what the model sees. Harness engineering, from the agent user's side, is about the environment it works in: what it knows before starting and what it finds out after acting.

  • Guides raise the odds of getting it right the first time.
  • Sensors let the agent discover on its own that it got it wrong.
  • Computational first, inferential for what's left.
  • The repository as the source of truth: what the agent can't see doesn't exist.
  • Every repeated mistake becomes a harness improvement, not just a one-off fix.

Models will keep getting better, and you don't control that. The harness is the part you do control. And it's what separates the agent that "sometimes gets it right" from the agent you trust to work while you do something else.

References​