Chapter 19The harness: architecting the agent loop
Part IV closes with the harness design: the loop every prior chapter has attached to but none has yet specified.
Every chapter so far has described something that attaches to the agent: bounded autonomy (Chapter 5) bounds it, governance (Chapter 6) gates it, memory (Chapter 7) feeds it, skills (Chapter 10) extend it, the trace (Chapter 12) records it. None of them has designed the thing they all attach to. The concept of the harness, the deterministic envelope around the model, was introduced in Chapter 4; this chapter designs it.
The loop and the code that runs it are the harness: the deterministic program that turns a model into an agent. The book’s thesis locates reliability in that shell: the harness. It is where the bounds are checked, the governance pipeline is called, the context is built, the tool results are dispatched, and the trace is written. A team can adopt every pattern in this book and still ship an unreliable system if the harness is an afterthought, because the harness is the one component that touches all the others on every turn. It is the last thing the book designs because it is the thing that holds the rest together.
Most teams adopt an agent framework rather than writing the loop by hand; below develops what production discipline requires around that default.
This chapter designs the deterministic envelope the model runs inside. Its central practical question is where a new capability belongs: in harness code, in a tool, in a skill, or in a new agent. Prompt craft, model selection, and techniques for improving model reasoning lie outside this architectural treatment (Preface).
The harness is the deterministic envelope
Industry usage often groups the system prompt, tools, filesystem, sandbox, and loop together under “harness.” That broad usage can label even a request interceptor around a model call a “harness engine,” although it has no loop and runs no agent.
This book uses harness for the deterministic loop and gateway. It assembles context, calls the model, parses intent, dispatches proposed actions through bounds and governance, observes results, and decides whether to loop again. The system prompt is content it carries. The filesystem, sandbox, and browser are tools it governs. This boundary lets a team test the harness’s guarantees separately from prompt edits and tool implementations; the broader definition would blur that responsibility.
A reasonable objection arises immediately. Chapter 1 argued that the reason-act-observe loop has been absorbed into the models: a modern reasoning model plans, acts, and self-critiques without external orchestration code, and Chapter 4 marked ReAct, Plan-Execute, and Reflection as largely model-internal by 2026. If the loop is inside the model, what is left to design?
The answer is the distinction between the inner loop and the envelope. What migrated into the model is the inner cognitive loop, the back-and-forth of reasoning, deciding which tool to reach for, and reasoning again about what to do next. What did not migrate is everything around it: assembling the context the model reasons over, executing the actions it proposes, enforcing the bounds, calling the governance pipeline, persisting state, and recording the trace. Chapter 1 named these in a single sentence, “bounding it, governing it, persisting its state, observing it, and recovering from its failures.” This chapter is the depth treatment of that sentence.
Tool selection can move into the model while tool execution stays in the harness. The model may decide mid-reasoning that it wants a tool. The harness carries out the call and receives its observation, giving bounds a chance to refuse, gates a chance to escalate, and the trace a chance to record the effect.
The useful test is whether the harness can refuse or modify a call before its effect occurs. An explicit tool loop in a graph or agent runtime exposes that point.
A computer-use action executed by your process exposes it too, even when the model selected the action. The boundary between deciding and doing is logical and can sit within a single turn.
Provider-hosted execution breaks that refusal point. A server-side code interpreter, hosted search, or other provider-resident tool may be selected and run inside one API call. By the time your code receives the result, the effect has happened on the provider’s infrastructure. The harness cannot refuse, rescope, or meter it beforehand.
The team must therefore keep effectful and irreversible capabilities on the harness side of the seam, either as client-executed tools or in a provider sandbox whose effects cannot escape. Provider-hosted execution is suitable for read-only or otherwise low-consequence operations whose completed effects are acceptable without preflight governance. Idempotence does not make a mutation safe to delegate: deleting a resource or freezing an account remains consequential even when repeated calls have the same effect.
Chapter 4 drew this line for patterns that had become mostly model-internal. Bounding, observability, and the tool surface remain architectural responsibilities. Chapter 5 accordingly makes the bounding layer the first surface an agent’s proposed action encounters. A provider-hosted effector concedes that first surface.
The harness is the deterministic envelope around the model’s inner cognitive loop. A single harness turn may wrap several model-internal reasoning steps. The harness assembles context and calls the model, which can reason and select tools internally before returning proposed actions. The gateway then executes those actions under bounds and governance.
The harness records what went in and which actions came out without needing to inspect the intervening deliberation. As models absorb more cognitive work, this envelope carries more responsibility for behavior decided inside an opaque call. When a provider also absorbs execution, the envelope loses its preflight control. That is why the execution seam must remain explicit.
The loop, structurally
Concretely, the harness is the code that runs one turn and decides whether to run another. Its responsibilities, in order:
-
Assemble context. Build the input the model will reason over this turn from the system instruction, relevant memory (Chapter 7), loaded skills (Chapter 10), tool declarations, and current task state. Assemble what the turn needs instead of appending the entire history, as developed below.
-
Call the model. The single probabilistic step. Everything before and after is deterministic code the team owns.
-
Parse intent. Interpret the response as a final answer or structured tool-call proposals with arguments. Free text alone does not authorize an executable command.
-
Dispatch through the gateway. Route every proposed action through the bounding layer (Chapter 5) and governance pipeline (Chapter 6) before it takes effect. The gateway may refuse, escalate, or execute the proposal.
-
Observe and update. Feed each result, success, refusal, or escalation, back as an observation, update memory and the bounds ledgers, and write the trace (Chapter 12).
-
Decide whether to continue. Check the iteration, cost, and deadline bounds, and loop again or terminate.

The per-action gateway internals, the action-surface check, the per-call cost and data-scope checks, and reversibility routing, are developed as pseudocode in the Concord worked example (Chapter 17). What matters here is the loop around the gateway, and the loop-level bounds the harness itself enforces: the iteration counter, the session deadline, and the running cost budget, checked before the expensive call rather than after (Chapter 5). In skeletal form:
def run_turn(session, task_state):
# The deterministic envelope. call_model() is the only DIRECTLY probabilistic
# line; a tool or sub-agent reached via the gateway may wrap its own model call.
if session.now() >= session.deadline or session.over_budget():
return Aborted("time or cost budget") # Ch 5: check before the costly call
context = assemble_context(session, task_state) # memory, skills, tools, state
response = call_model(context,
max_tokens=session.remaining_token_budget())
session.trace("model.responded", response.summary)
intent = parse_intent(response) # final answer or proposed actions
if intent.is_final:
return Done(intent.answer)
for action in intent.actions:
# the gateway enforces every bound axis AND the governance pipeline (Ch 5, 6, 17)
result = gateway(session, action) # may allow, refuse, or escalate
task_state.observe(action, result) # refusals are observations too
session.memory.record(action, result) # Ch 7
session.trace("action.result", action, result)
session.iter_count += 1
if session.exceeds_any_bound(): # iteration, cost, deadline (Ch 5)
return Aborted("bound exceeded")
return Continue(task_state)
Two properties of this skeleton matter. First, the only directly probabilistic line is call_model; the tools and sub-agents reached through the gateway may each wrap a model call of their own, but every one of those is itself a bounded step, and the harness code between them is ordinary software, testable, reviewable, and version-controlled (Chapter 12). Second, the harness never lets a model output reach the world directly. The model returns intent; the harness decides what becomes an effect. That separation is not an implementation convenience; it is the architectural seam on which the entire discipline depends, and it is the subject of the next section.
Reasoning versus hands and eyes
The most useful principle for designing the seam is a division of labor that runs through the whole book without ever being named: the model reasons, and it has no hands or eyes of its own. It cannot read a file, call an API, write a record, or perceive the state of the world. It can only emit text proposing that one of those things happen. Every effector (a tool that acts) and every sensor (a tool that perceives or retrieves) is dispatched by code the harness owns; the implementation behind a tool may itself be probabilistic (a search service, a reranker, a model-backed sub-agent), but its execution passes through the harness, where authorization, bounds, validation, and tracing apply. Chapter 4 drew the line precisely: the cognitive act of deciding to reach for a tool is distinct from the action of executing the call; the model decides which tool, and the architecture decides which tools exist, with what authorization, and with what logging.
The architectural rule that follows is simple to state and governing in practice: the model decides; the harness acts. No effector is wired directly to a model output, and where a platform would run one for you (the provider-hosted execution discussed above), effectful tools are kept on the harness side of the seam. The value of routing every effector and sensor through harness code is that each such routing is a place where a bound can be checked, a policy gate can fire, an argument can be validated against a schema, and a side effect can be recorded in the trace. A model wired directly to an effector is outside governance; routed through the harness, its every effect is bounded, observed, and recoverable.
The principle cuts in two directions, and both are design errors when ignored. Pushing into the model what should be deterministic, asking it to compute a total, enforce a policy, or format an identifier in prose, takes a guarantee the harness could have made for free and replaces it with a probability. Pushing into code what genuinely requires judgment, hard-coding a branch the task actually needs the model to reason about, throws away the one thing the model is for. Designing the harness well is largely the discipline of drawing this line in the right place, capability by capability, which is the decision the rest of the chapter is about.
Assembling context
Step one of the loop, assemble context, is where several prior commitments converge into a single harness responsibility: a bounded window (Chapter 3), memory gateway scoping (Chapter 7), and skills loading (Chapter 10). The context window is the model’s entire field of view for a turn, and it is a bounded, costly resource. The naive harness treats it as an accumulator: it appends every turn, every tool result, and every document to a growing transcript and passes the whole thing back to the model. That harness exhausts its window, pays to re-read stale material on every call, and buries the relevant signal in noise. That is the prompt stuffing failure mode (Glossary, Chapter 11).
The disciplined harness assembles context for each turn rather than accumulating it. It keeps a large, stable prefix for cache reuse and puts volatile material last (Chapter 18). It retrieves context rather than stuffing it into the prompt (Chapter 7), discloses skills as needed and evicts them when their task ends (Chapter 10), and budgets the context window for each turn. The harness is where these are exercised because it is the component that builds the input; no other layer can.
Context is constructed on every turn by the harness. The harness is the constructor. Treating the window as an append-only log is the most common way an otherwise sound architecture becomes slow, expensive, and unreliable in production.
Where capability lives
The recurring question a team faces once the harness exists is practical: where does a new capability go? The book’s earlier chapters answer fragments of this. This is the place to answer it whole, because the choice is the single most consequential design decision the harness forces, and it is what the “code way versus the emerging skills way” of building agents is about.
First separate capability from knowledge. A new fact, corpus, or retrieval source belongs in memory and the ingestion pipeline (Chapter 7, Chapter 8). A change to standing behavior belongs in the system instruction the harness assembles. Executable or procedural capabilities have four possible homes, distinguished by their build, change, and governance costs.
Harness code. When the behavior must be guaranteed and does not itself touch the world, internal logic, identical every time, not subject to the model’s judgment, it belongs in deterministic code: control flow, enforcement, parsing, retries, formatting, the bounds themselves. Code is the cheapest home at runtime and the most expensive to change, because changing it means a deployment. Anything essential for correctness or safety lives here. The bounds are code; the governance pipeline is code; the loop is code.
A tool. When the agent needs to act on or perceive the world, the capability is a tool: an effector or sensor admitted to the governed action surface (Chapter 5). A tool is the unit the bounding layer governs and the trace records; its implementation may be perfectly deterministic code, but “tool” denotes the governed surface and the fact of an external effect, not the presence of judgment. Add a tool when the agent needs to do something it structurally cannot do today, and accept that every tool widens the action surface and so must be justified against it.
A skill. When what is changing is not the tools but the know-how for using existing tools, a procedure, a project convention, a workflow over verbs the agent already has, the capability is a skill (Chapter 10): procedural knowledge, loaded at runtime, editable without a deployment. A skill is the right home for the large and fast-moving body of “how we do things here.” Reach for a skill instead of writing a new prompt or a new playbook, and instead of hard-coding a procedure that will change next month.
A new agent. When the work needs its own bounded context, action surface, or reasoning budget, and only when coordination genuinely benefits from the separation, the capability is a new agent: a sub-loop with its own envelope (Chapter 9). This is the most expensive home in every dimension, and the book’s standing caution holds: most production systems are a single agent with tools, and a new agent should be the last resort, not the first reach.
A short decision sequence resolves most cases:

The sequence is a heuristic, not a strict partition: a sub-agent acts on the world too, and a tool’s implementation is itself code, but it resolves the common case by asking the cheapest distinguishing questions first: is this knowledge rather than capability, and if capability, does it need a new effector or sensor the agent lacks, or only new know-how over the tools it already has?
Three trade-offs move along the sequence. Iteration speed improves as capability moves from code toward skills because a skill can change without a deployment. That also lowers deployment risk. Governance surface and testability become harder to control in the same direction. Code can be unit-tested and is tightly governed; a runtime-loaded skill requires different validation because it arrives as content.
Skills are therefore attractive when a procedure changes often, but the team must account for the weaker verification available at that boundary.
Skills may carry procedure but never authority (Chapter 10). A skill is a document the agent reads; it cannot expand the action surface, bounds, or enforcement. Adding a tool requires admitting that tool in code, even when its operating instructions live in a skill.
The emerging “skills way” lets teams stabilize the loop, action surface, and governance in code while iterating on procedure as content. It crosses the boundary when a skill purports to admit tools or data scopes absent from the declared action surface.
Tool design as harness design
When an agent seems to reason badly, audit its tool surface. Tools are the harness’s hands and eyes; a poor interface can make sound reasoning produce poor actions. Prefer coarse, intention-revealing operations over thin wrappers around raw power. A governed operation is safer and easier to interpret than a query interface exposing the whole database, as Chapter 14 explains.
Keep the surface small enough for the task. A registry of hundreds of tools expands the attack surface and dilutes the model’s choice (Chapter 5). Enforce call schemas in the harness. Treat tool names and descriptions as contracts because the model reasons over them; stale descriptions can degrade behavior without an explicit error (Chapter 11).
Split a tool that performs unrelated operations. If several tools are always called in lockstep, consider a coarser operation that can be governed as one action. Trace replay (Chapter 12) helps test these choices: during a cascade investigation, inspect what the tools exposed alongside what the model said.
Retrofitting bounds and governance into a framework
Most teams will not write the loop above by hand. They adopt an agent framework, LangGraph, a provider Agent SDK, or an in-house orchestrator, that ships a default harness: the loop, the state plumbing, the tool dispatch. The default is real but partial. It gives you the skeleton of step 1 through step 6, but it does not give you the bounds, the governance pipeline, the context discipline, or the trace, and it is built around its own assumptions about where tools execute. The production harness this chapter describes is mostly what you build around that default, and the practical question is not whether to adopt a framework but where, once it runs an agent loop, your code can still refuse a tool call before its effect occurs. That question has a small number of answers, and they map onto the seams a framework either exposes or does not.
Four interception points matter, in the order a request passes through them:
-
The context seam, where the framework assembles the prompt before the model call. Hooking it lets you enforce the assembled-not-accumulated discipline, scope retrieval through the memory gateway, and inject the loaded skills’ declarations (Chapter 7, Chapter 10).
-
The model-call seam, the primary model call. This is the loop’s only directly probabilistic step, as the skeleton above noted. A pluggable inference client lets you insert the model gateway for egress filtering, capability-tier routing, and cost attribution (Chapter 15), and set the per-call
max_tokensthat bounds a single generation against the remaining budget (Chapter 5). -
The tool-execution seam, the point at which the framework, having received the model’s proposed tool call, executes it. This is the load-bearing one: a dispatch hook here routes every proposed action through the bounding gateway and the governance pipeline before it touches anything, which is exactly the interception the loop above depends on.
-
The observation seam, where the tool’s result returns to the model. A transform hook here sanitizes tool responses, enforces output schemas, and writes the structured trace (Chapter 6, Chapter 12).
A framework is governable when tool execution passes through a seam your code can intercept before an effect occurs. A graph engine or in-house orchestrator that runs tools in your process exposes that seam. You can wrap its dispatcher so every effect passes through the gateway.
A provider Agent SDK that runs hosted tools inside one API call may hide the seam. By the time your code sees the result, the effect has already occurred on provider infrastructure. A wrapper around the SDK cannot restore the missed refusal point; the tool-execution choice must be made when the framework is adopted.
The retrofit discipline for a framework that does this is to read its tool list as a contract: keep every effectful or irreversible tool on your side of the seam (client-executed tools, or a provider tool sandboxed so its effects cannot escape), and restrict the provider-hosted tools to the read-only or low-consequence operations whose completed effect is acceptable without preflight governance. Where a framework offers both hosted and client-executed tools, the choice between them is the choice the seam makes for you.
Retrofit the framework in the same three stages as the brownfield case (Chapter 18). First, wrap model-call and tool-execution hooks to emit a structured trace. This reveals the distributions of cost, iteration, and action before setting limits.
Second, interpose the bounding gateway at the tool-execution seam. Set ceilings from the observed distributions with headroom so runaway cases abort. Third, route consequential actions through schema validation, policy gates, and approval, starting with the highest-risk action. The framework’s loop continues to run; your code governs the seams it exposes.
Two practical tests tell you whether the retrofit has landed. The first is the refusal test: can you write a test in which a bound or a policy gate refuses a proposed action, and assert that the action never reached its downstream effect? If the framework swallows your refusal and runs the tool anyway, the tool-execution seam is not on your side, and the framework cannot carry this book’s discipline. The second is the trace test: does the framework emit the events the governance-event matrix of Chapter 12 requires, or does it emit only its own logs? If it emits only its own logs, the framework’s observability is for debugging its loop, not for governing yours, and you must wrap the seams to emit the typed events the trace store expects. A framework that passes both tests is governable; one that fails either is a component to build around, with your own loop, rather than to build on.
Harness anti-patterns
The failure-mode catalog (Chapter 11) develops these failures at length. Here they mark harness-level design mistakes. Reasoning in prose for what should be code asks the model to enforce a limit or compute a value the harness can guarantee (Chapter 5, Chapter 6). The kitchen-sink tool surface exposes every tool to every task (Chapter 5). Prompt stuffing accumulates context across turns (Chapter 11).
Anthropomorphizing the loop assigns the system’s guarantees to the model even though the harness enforces them. Skill sprawl without eviction keeps procedural content in the context after its task ends (Chapter 10). Each mistake weakens a guarantee that belongs in the team’s deterministic architecture.
Summary
The harness is the deterministic envelope that turns a model into an agent: it assembles context, calls the model, parses intent, dispatches every action through the bounds and governance, observes the result, and decides whether to continue. The inner cognitive loop has largely moved into the model; the envelope has not, and as the model absorbs more reasoning the envelope carries more, so long as execution stays on the harness side of the seam, which is why provider-hosted execution of effectful tools is the case to avoid. The principle that organizes the design is that the model reasons and the harness acts; no effector runs outside the harness’s reach, so that every effector and sensor passes through a place where a bound, a gate, a schema check, and a trace entry can attach. Context is assembled on every turn, not accumulated. And the recurring question of where a new capability belongs has, for executable capability, four answers, code for guaranteed internal logic, a tool for a new effector or sensor, a skill for know-how over existing tools, a new agent only when the work needs its own envelope (new knowledge goes to memory instead), with the firm constraint that know-how may move into skills while the action surface and its enforcement stay in code. Design the harness as the system, with the model as the one bounded component inside it, and the rest of the book’s discipline has something solid to attach to.