92. Deep with Harnesses 1
Understand what a harness is, what it owns, how the agent loop runs, and how to tell whether a problem belongs to the model, the harness, or the context.
By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.
The lesson
A harness is everything around the model that turns raw text prediction into an agent: the tools, the system prompt, the permission layer, the context management, the session store. The model is only parameters. It takes tokens in and produces tokens out, once per request, and it cannot read a file, run a command, or remember the previous turn. The provider is whatever serves the model for inference, often a remote service and sometimes a local process on your own machine. The agent is the thing you actually talk to, the model in motion, configured for a purpose. Keeping these four words apart is the first professional habit, because the same complaint about poor output means four different fixes depending on which layer is at fault.
The agent loop is the mechanism that makes it look like the model is working. The harness assembles a request: the system prompt, the conversation so far, every tool result. The model proposes an action, usually a tool call written as structured text. The harness parses that call, checks it against the permission rules, executes it if allowed, and appends the result to the context. Then the whole context goes back to the model, and the cycle repeats until the model answers without calling a tool. One turn of your conversation can contain many of these round trips. The model never executes anything itself; it only writes the request, and the harness is what acts.
A harness owns the system prompt assembly. In a coding harness that brief is often tens of thousands of tokens of behavioural rules, tool descriptions, and edge cases, and it is re-sent on every request. A harness also owns the tool definitions: the names, descriptions, and parameter schemas the model can choose from, and the tool list is the real ceiling on what the agent can do. It owns the permission layer as well, the ladder from read-only through ask-on-write to full auto, and that ladder decides how often you are interrupted and how much damage a bad call can do. Two products running the same model with different prompts, different tools, and different permissions will behave like different products, because at that point they are.
The quieter half of a harness is state management. It owns context-window management: when the window fills, it decides whether to summarise the session, drop the oldest tool results, or stop and ask. It owns the session store, the transcript kept on disk that lets you resume or read back what happened. It owns hooks, the small scripts that fire around tool calls, for example formatting a file after an edit or blocking a dangerous command before it runs. It owns memory files, the rules documents loaded at session start, which is where your standing project instructions live. And it owns subagents, the child agents spawned to run a noisy search in their own context and report back a short result. None of this is the model, and all of it changes the outcome.
Diagnosis follows directly from that split. When something goes wrong, ask three questions in order. Is this the model, meaning would a stronger or different model plausibly do better on the same input? Is this the harness, meaning did the system prompt, the tool set, the permission default, or the context policy shape this behaviour? Or is this the context, meaning was the relevant fact ever in the window, or was it buried under stale tool output? Most disappointing output traces to context first and harness second, and to the model last. Blaming the model is the most expensive diagnosis, because it usually means a migration when the real fix was a rules file and a fresh session.
A worked example makes the separation concrete. Take one mid-tier model and two harnesses. In the first, a chat harness with no filesystem tools, you ask it to fix a failing test. It explains the likely cause, quotes an invented line number, and offers a corrected snippet you must paste in yourself. In the second, a coding harness with read, write, and shell tools, the same model on the same prompt reads the test, reads the source, runs the suite, sees the actual error, edits the file, and re-runs until green. The model did not improve between the two runs. The harness gave it eyes and hands, and the difference in result came entirely from the tools, the permission layer, and the loop.
This is why TPEE treats the harness as a separate world. TPEE is a strictly offline prompt-engineering editor: it never launches a harness, never contacts a provider, and never runs an agent on your behalf. What it does is help you prepare the structured prompt, the rules text, and the conventions that you then carry into whatever harness you have chosen. The transfer is manual. You copy the prepared prompt out of TPEE and paste it into the harness yourself, in a separate application that you installed and control. Documenting a harness inside a lesson is not the same as connecting to one, and the app makes no network or MCP call of any kind.
The practical takeaway is to stop thinking of the model as the product. When you evaluate a harness, look at what it owns: how it assembles the system prompt, how many and which tools it exposes, how it gates risky actions, how it manages a filling context window, where it keeps sessions and memory files, and whether it can spawn subagents. Those are the levers you actually configure. The model is one input to that machine, and a good harness with an ordinary model will often beat a poor harness with a frontier one.
Check yourself
Question 1: In the agent loop, what does the model do when it calls a tool?
- It executes the command itself and reads the output
- It writes a structured call that the harness parses and runs — correct
- It sends the call to the model provider to execute
- It writes the result into the session store on disk
Answer: It writes a structured call that the harness parses and runs
The model only produces text, and a tool call is structured output. The harness parses it, checks permissions, executes it, and feeds the result back in the next request.
Question 2: A teammate says the model is bad because it invents a field that is not in the type definition. What is the most useful first question?
- Should we switch straight away to a much larger and stronger model?
- Is the model provider having an outage at this moment?
- Should we raise the effort setting to its maximum for the session?
- Was the type definition actually loaded into the context window? — correct
Answer: Was the type definition actually loaded into the context window?
Most disappointing output traces to context first. If the type definition was never loaded, no model swap fixes it, because the model cannot read a fact it was never given.
Question 3: Which statement correctly describes TPEE's relationship to a coding harness?
- TPEE is strictly offline; you copy the prepared prompt out and paste it into the harness yourself — correct
- TPEE launches the chosen harness in the background and forwards your prompt to it automatically
- TPEE connects to the harness through MCP and streams each tool result back into the editor
- TPEE runs the harness only when you tick the allow-network box in the preferences panel
Answer: TPEE is strictly offline; you copy the prepared prompt out and paste it into the harness yourself
TPEE never launches, contacts, or drives a harness. It prepares the prompt and the conventions offline, and you move the text into the separate application by hand.
← Previous lesson · All 94 lessons · Next lesson →
The full course — 94 lessons and 282 quiz questions — ships inside the app. Get TPEE to study it offline.