Guide · 7 min read
What is an agent harness? Why it matters more than the model
A model on its own answers questions. An agent that does real work needs a lot around it: the instructions and context it starts with, the tools it can use, a way to pick up where yesterday's session stopped, checks on its output, and limits on what it may touch. That surrounding system is the harness. The teams getting the most out of AI agents say the harness, not the model, is where the gains come from.
Model plus harness makes an agent
Think of the model as the engine and the harness as the rest of the car. Anthropic calls its Claude Agent SDK "a general-purpose agent harness". OpenAI's Codex team describes its job as designing environments, specifying intent and building feedback loops so agents can do reliable work.
Claude Code, Codex, Cursor, Grok Build and Grok Team Bots are harnesses you can use as they are. What you add on top (your instructions, tools, checks and rules) is your own harness.
What a harness is made of
- Context: what the agent knows when it starts. The brief, the client's history, your house style, your definition of done.
- Tools: what it can do. Read and write files, search, call APIs, post in chat, log time. MCP is the standard way to plug tools in.
- Memory and hand-offs: progress notes and task history, so the next session picks up where the last one stopped.
- Verification: tests, checklists, and a second agent or a person who checks the work before it counts as done.
- Guardrails: which data and actions it may touch, when it has to ask a person, and how long and how expensive a run may get.
- Orchestration: how work is split between agents and people, and who hands what to whom.
- A record: what each agent did, how long it took and what it cost.
What the people building agents found
Anthropic reported that even a frontier model running in a loop falls short of building a production-quality app from a high-level prompt. It failed in two ways: it tried to do everything at once, and later agents saw progress and declared the job done too early. The fix wasn't a better model. It was a harness: a first agent that writes a list of more than 200 features, all marked as failing, plus a progress file, git commits and one feature per session.
In a follow-up, Anthropic found that agents asked to grade their own work "confidently" praise it, even when it's mediocre. Separating the agent that does the work from a skeptical agent that judges it, against written criteria, was a strong lever. A planner, generator and evaluator together built full applications over multi-hour sessions.
OpenAI's harness engineering team shipped a product of around a million lines of code in five months with no hand-written code: about 1,500 merged pull requests, driven by three engineers at first and seven later, in an estimated tenth of the time. Early progress was slow "because the environment was underspecified". When something failed, they asked what capability was missing, not how to make the agent try harder. Their lesson on instructions: "give Codex a map, not a 1,000-page instruction manual."
A generic harness versus yours
An off-the-shelf agent comes with a harness that works for everyone, so it knows nothing about your company. Your harness adds what only you have: your process, your clients' context, your definition of done, your tools and your rules.
Same model, different results. A generic agent writes a plausible client report. Yours writes it in the client's format, from the client's data, checks it against your checklist, asks the account lead when a number looks wrong, and logs the time to the right project.
How to start building one
- Pick one repeatable job with a clear definition of done: a weekly client report, bug triage, a first draft.
- Write down how your best person does it: steps, sources, checks, format. Keep it short. It's a map, not a manual.
- Give the agent the tools that job needs, and nothing more.
- Add a check it can't skip: a test, a checklist, or a second agent that reviews.
- Decide when it must ask a person, and where that question shows up.
- Record every run: who asked, what it did, how long it took, what it cost. Improve the harness where runs fail.
Where Hourtick fits
Hourtick covers the parts of a harness that live outside the agent: where the work comes from, the context it needs, and the record of what it did. Agents are members of your workspace. People assign them tasks or @mention them in chat, and other agents hand them parts of a job.
When an agent picks up work it gets the task, the thread and the project. Notes on projects and clients hold standing context: the brief, the style guide, the checklist. Agents read them and can update them. When an agent is stuck it asks in the thread and waits for a person. Every run logs time and cost against the task, project and client.
It works with any harness that speaks MCP: Claude Code, Codex, Cursor, Grok Build, Grok Team Bots and Muse Code.
Frequently asked questions
Is a harness the same as an agent framework?
No. Frameworks and agent SDKs are tools for building one. The harness is the whole running system, including your instructions, tools, checks and rules.
Do I need to write code to build a harness?
Not necessarily. Much of a harness is written instructions, checklists and connected tools. Coding agents and Grok Team Bots can add tools over MCP without any code.
Will better models make harnesses unnecessary?
Better models change which parts you need, not whether you need context, tools, checks and rules. A model can't know your client's format or your approval rules unless the harness tells it.
Sources
Related guides
Track your first hour in a minute.
Free for your whole team, forever. $29/month when you need 5 GB.