Preface

The mandate from enterprise leadership is blunt: build autonomy, automate the workflows, and deploy the agents. It has a real economic basis. Software can now attempt work whose inputs and acceptable outcomes cannot be specified completely in advance: triaging a messy ticket, drafting a considered reply, or working with an unfamiliar API. Each task requires judgment about what matters and what to do next. That capability creates a legitimate reason to reconsider which work software can perform and how much human attention the work requires.

Urgency also encourages structural mistakes in problem selection, system design, and preparation for failure. A team may prescribe a multi-agent swarm where a simple router is sufficient, adding coordination costs without a corresponding need. It may treat conversational logs as databases even though a transcript does not provide the structure expected of durable state.

A system prompt may be assigned the role of a security boundary, leaving enforcement inside the component whose behavior requires control. Open-ended reasoning loops may then reach production without adequate controls for failure. These mistakes arise at different points in the design, but each assigns an architectural responsibility to a probabilistic component that cannot guarantee it.

Successful demonstrations can conceal the consequences. A stage-managed demonstration controls the inputs, the available tools, and the path through the task closely enough to make a behavior appear dependable. Production changes those conditions. Inputs become less orderly, state persists across interactions, and actions have consequences outside the demonstration. Confidence earned from the demonstration therefore exceeds what the production architecture can guarantee. The industry sometimes calls this AI theater. The phrase is rude, but it captures the substitution: a convincing performance standing in for an engineered guarantee.

I wrote this book because I could not find an integrated architectural account of how these parts hold together in production. Existing sources provide substantial treatments of individual patterns and frameworks. Gulli, Anthropic, CSIRO, and Andrew Ng supply canonical accounts that this book cites instead of reproducing. Governance as architecture is also an emerging consensus, not a position introduced here. The contribution of this book is integration: it connects bounded autonomy, governance, memory, control, skills, failure analysis, evaluation, operational discipline, and the harness into one account intended to remain useful as models and frameworks change.

Unnamed incidents are sourced composites of documented public cases, altered to remove identifying detail. Statements about the state of the field are dated to mid-2026. Those facts establish the conditions under which the book was written; the architectural arguments concern responsibilities and boundaries that are intended to outlast them.

The intended reader is an experienced software practitioner responsible for putting agentic systems into production: a senior engineer, architect, or technical leader who already understands standard software engineering. The book assumes familiarity with ordinary concerns such as services, persistence, interfaces, testing, and operations. Its practical concern is how those responsibilities change when a system contains a component whose outputs remain probabilistic, especially for the people who will carry the pager.

An agentic system embeds probabilistic reasoning components inside the deterministic infrastructure of a distributed system. Its volatile components can behave differently even when a task appears familiar. The surrounding infrastructure must bound what those components can do, govern consequential actions, observe their behavior, and recover when their reasoning or execution fails. Reliability depends on assigning those responsibilities to mechanisms capable of enforcing them.

This architectural reference explains how to embed an agent with probabilistic reasoning in a system whose behavior must remain operable and accountable. It assumes the reader already knows basic server construction and leaves prompt engineering and model training to their respective literatures.

Chapter 1 establishes the architectural shift, and Part I introduces the harness concept in Chapter 4. The full design of that harness follows in the capstone, Chapter 19, after the structures attached to it have been developed.

The book proceeds in four parts. Part I establishes the foundations: the definition of an agent, the forces that shape agentic design, and the cognitive patterns that shape reasoning. Part II develops the architecture around reasoning, including bounded autonomy, governance, structured memory, ingestion, control and coordination, and runtime-loaded skills. Part III examines production concerns through failure modes, evaluation and trace discipline, interaction, enterprise integration, and model routing. Part IV composes those elements through system vignettes and a worked example, then addresses deployment, cost, scaling, observability, lifecycle, and the operational discipline required to retrofit ungoverned agents in Chapter 18. The harness design completes that synthesis.

Infrastructure creates and constrains the operational state we call autonomy. Cost limits, iteration limits, risk controls, policy gates, and recovery paths remain outside the reasoning component so that they continue to apply when its judgment fails. The economic case for autonomous work depends on this arrangement. A model may attempt judgment-laden tasks at useful scale only if the deterministic shell catches its errors before those errors become uncontrolled system behavior. This book is about how to build that shell.