94. Deep with Harnesses 3

Choose, configure, drive and audit a coding harness on a real project, and apply a harness-agnostic checklist that keeps the work verifiable.

By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.

The lesson

Choosing a harness is a decision you make once and then live with, so make it deliberately. Read the decision along a few axes: does the work need write access or only review, where does the code actually live, how closely must you watch it, how isolated must an unattended run be, and what is your realistic budget? Shortlist two candidates, then give each a throwaway task that touches three files, so you can judge the rules file, the permission prompts, and the diff review before you commit. Whatever you choose, the preparation happens in TPEE and the execution happens in the harness; the app itself never launches or contacts anything.

Configuration starts with the rules file, the document the harness loads at the start of every session. Claude Code calls it CLAUDE.md and most others accept AGENTS.md, and the content is the same either way: the build and test commands, the project conventions, the boundaries of the working directory, and a plain statement of what done means. Keep it short and specific, because every line is re-sent on every request and competes for attention. Version it alongside the code, so the rules travel with the project rather than living in one person's notes. TPEE is the natural place to draft and keep that text before you copy it across.

Next come the safety controls. Permission modes set how often the harness stops to ask: a read-only or plan mode for research, an ask-on-write default for everyday work, and auto-edit or full-auto only where the blast radius is contained. A sandbox is what makes a loose mode defensible, because it limits what a bad call can reach even when it runs without asking. Hooks are the third control, small scripts that fire around tool calls, for example formatting a file after every edit or refusing a command that would delete outside the project. Start tight, loosen only as the harness earns it, and never loosen on an unattended run without isolation.

Then tune the engine. Match the model tier and the effort dial to the task rather than to the session: a stronger model at high effort for planning and hard debugging, a smaller one at low effort for mechanical edits and lookups. Reach for subagents when a search is noisy, because a child agent can burn its own context on the hunt and hand back one short result instead of flooding your window. Resist the urge to set everything to maximum, which buys a slow, expensive session that produces nothing the sensible settings would not have.

Driving well is mostly about keeping the steps small. Ask for a plan before you allow edits, then approve one verified step at a time. Prefer a short plan you can argue with over a large change you have to accept or reject as a whole. Keep one task per session, because unrelated work leaves residue that colours the next answer. When a step lands, check it before the next one starts, so a wrong turn is caught while it is still cheap to reverse. Autonomy is something you grant in increments, not a setting you flip at the start.

Context hygiene is what keeps a long project sharp. Watch the context indicator and compact or clear at a natural boundary, such as after the plan is settled and before the implementation begins, rather than letting an automatic compaction fire mid-task and decide for itself what to keep. Write decisions into memory files and handover artefacts on disk, where no summary can lose them, and point the next session at those files first. A short plan document and a written handover turn a cold start into a cheap retrieval step, and both are things you can prepare in TPEE before the session opens.

Know the failure modes by name. Context rot is the gradual slide in quality as the window fills, the same model getting vaguer and less obedient without any error being raised. Permission fatigue is approving prompts without reading them, which gives you all the interruption and none of the protection. A hallucinated tool result is the agent claiming it ran the tests when the transcript shows no such call. Silent scope creep is a small request quietly widening into a refactor you never approved. And cost blowout comes from re-sent input tokens: every request carries the whole history, and anything that breaks the prefix cache, such as an injected timestamp in the system prompt, makes each request bill at full rate.

Keeping a harness honest is a discipline, not a setting. Review every diff before you accept it, and read the transcript rather than a summary of what was done, because the summary is a secondary source and the transcript is the primary one. Run the build and the tests yourself when the stakes are real, and treat a green claim from the agent as a hypothesis until your own gate confirms it. Keep the work inside a dedicated project folder, keep it under version control, and commit in small steps so any bad change is easy to find and easy to undo. A harness that is checked this way can be trusted further; one that is not will drift.

A harness-agnostic checklist you can apply to any tool: one, does it read a rules file, and is yours short, current, and versioned? Two, is the permission mode matched to your trust and, for unattended runs, backed by isolation? Three, are the tools it exposes the ones the task needs, and no more? Four, is the model tier and effort dial matched to the step? Five, is the context being compacted or cleared at boundaries rather than by accident? Six, are decisions written to disk as memory files and handovers? Seven, is every diff reviewed and every build gate run by you? Eight, is the whole task kept inside one working directory? And nine, is the preparation done offline in TPEE and the execution done in the harness, with the prompt copied across by hand? If you can answer all nine, the harness you picked matters far less than the discipline you brought to it.

Check yourself

Question 1: Which practice best prevents an agent from silently widening the scope of a change?
  1. Setting the effort dial to its maximum for the whole session
  2. Approving every permission prompt without reading it, to keep momentum
  3. Keeping one very long session so the agent never loses earlier decisions
  4. Asking for a short plan first, then approving one small verified step at a time — correct

Answer: Asking for a short plan first, then approving one small verified step at a time

A plan you can argue with, followed by small steps you verify as they land, keeps scope visible. Long sessions and blind approval both let the change grow out of sight.

Question 2: Why can a long coding session become surprisingly expensive even when the agent writes very little?
  1. Output tokens are billed again and again on every request throughout a long session
  2. Each request re-sends the whole history as input tokens, and a broken prefix cache removes the discount — correct
  3. The provider charges a fixed monthly fee for every tool that is defined in the harness
  4. Reading files from local disk is billed per file by the model provider on each request

Answer: Each request re-sends the whole history as input tokens, and a broken prefix cache removes the discount

The model is stateless, so every request carries the full history as input tokens. The prefix cache makes that affordable, and anything that breaks the shared prefix removes the discount.

Question 3: A colleague asks whether TPEE can run a harness for them. What is the correct answer?
  1. Yes, provided the harness is installed in the same folder as the editor itself
  2. Yes, but only for read-only tasks that make no changes to any files
  3. No: TPEE is strictly offline and never launches or contacts a harness; the user runs it separately — correct
  4. No, because TPEE has simply not yet added the harness integration menu

Answer: No: TPEE is strictly offline and never launches or contacts a harness; the user runs it separately

The boundary is permanent policy, not a missing feature. TPEE prepares the prompt and the conventions offline, and the user copies them into a harness they installed and run themselves.

← Previous lesson · All 94 lessons

The full course — 94 lessons and 282 quiz questions — ships inside the app. Get TPEE to study it offline.