Skill Issue: Harness Engineering for Coding Agents
We've spent the past year observing coding agents stumble in various ways: ignoring instructions, executing dangerous commands without prompts, and getting stuck on simple tasks. Teams have shipped sloppy code, and we've even contributed to the mess ourselves.
Every time, the instinct was the same:
- "Better models like GPT-6 will fix it."
- "Improved instruction-following is the key."
- "It'll work once my niche library is in the training data."
However, after numerous projects and countless agent sessions, we realized it's not a model problem. It's a configuration problem. Models will get smarter, but as they do, we'll give them bigger challenges, and they'll continue to fail in unexpected ways. This is a fundamental issue with non-deterministic systems.

Instead of waiting for gpt-6.4-codex-ultrahigh_extended to save us, we focus on maximizing the potential of today's models. There are many ways to enhance your coding agent's performance. If you use coding agents for moderately complex tasks, you've likely configured them a bit. Have you used skills, MCP servers, sub-agents, memory, or AGENTS.md files?
Text
These are all part of the coding agent's configuration, which we call the coding agent's harness. Think of it as the agent’s runtime or its peripherals: what does the model use to interact with its environment?
Harness Engineering
Harness engineering, a term coined by Viv, involves leveraging these configuration points to customize and improve your coding agent's output quality and reliability.

As Mitchell Hashimoto explains, harness engineering is about creating solutions so that agents don't repeat mistakes.
Harness Engineering as Context Engineering
We see harness engineering as a subset of context engineering, a concept introduced by my cofounder Dex in 12-factor agents. Context engineering encompasses “prompt engineering” and other techniques to systematically improve AI agents’ reliability.

It addresses:
- How do we give our coding agent new capabilities?
- How do we teach it things about our codebase not in the training data?
- How do we add determinism beyond
CRITICAL: always do XYZin the system message? - How do we adapt the agent’s behavior for our specific codebase?
- How do we increase task success rates beyond “magic prompts”?
- How do we prevent our context window from inflating too rapidly, or with too much bad context?
Skills, MCP servers, sub-agents, hooks, and back-pressure mechanisms are tactical solutions we've developed.
Views on Harness Engineering
Viv’s posts on harness engineering are worth reading alongside this one — the first frames the four customization levers (system prompt, tools/MCPs, context, sub-agents), and the second works backwards from what models can’t do natively to derive why each harness component exists.

We’d add two levers he doesn’t emphasize:
- Hooks for automated integration and deterministic control flow.
- Skills for progressive disclosure of knowledge. (Dex refers to them as "Instruction Modules" - more on this in another post.)
After months of solving complex problems in enterprise-scale codebases, we've found sub-agents to be particularly powerful. When tackling hard problems requiring many context windows, sub-agents maintain coherency across sessions. They act as a "context firewall," ensuring discrete tasks run in isolated context windowsundefinedundefined






