Harness Engineering: Building Reliable AI Agents That Actually Deliver (AI and Agentic Engineering) - Softcover

Book 2 of 5: AI and Agentic Engineering

Halloran, Wes

 
9798181738430: Harness Engineering: Building Reliable AI Agents That Actually Deliver (AI and Agentic Engineering)

Synopsis

Everyone in the room saw the agent work. That is the problem.

In production it fails one run in twelve, never the same way twice, and you carry the pager. You are not stuck because the agent cannot do the work. You watched it do the work. You are stuck because doing it once, on stage, with you watching is the easy part, and doing it every time, while you sleep, is the job nobody named for you.

Harness Engineering names that gap: the demo cliff, and the climb back up. This is not a book about building an agent. You already built one. This is the book about making it deliver: the evals, verification, guardrails, observability, and recovery that turn an impressive toy into a production LLM application you would put your name on the pager for.

The engineer's toolkit for reliable AI agents that actually deliver:

  • Reliability as a number: a success rate against a fixed task set, tracked run over run, so better is a measurement instead of a feeling.
  • The eval gate: agent evals run like CI, blocking a bad change before it ships.
  • The verification wall: cheap checks first, expensive ones behind them, catching wrong output before a user ever does.
  • The recovery path for the failures you cannot prevent: retry with judgment, fall back, escalate, roll back.
  • Run records: LLM observability and monitoring that capture what the agent actually did, so a bad run is a query instead of an archaeology dig.
  • The reliability budget for one real agent: a target you can defend and the guardrails that hold it.

Read it and you will measure your agent instead of demoing it, ship changes behind an eval gate, catch wrong answers before your users do, and answer the 3 AM page with a query instead of a guess. A reliable agent still fails. Its failures are rare, cheap, caught before the user, and survivable when they are not. That is software reliability engineering, the discipline SRE brought to servers, applied to agents. This book is the engineering.

For software, ML, and platform engineers who already run agentic loops and now own one in production. Part of the AI and Agentic Engineering series.

"synopsis" may belong to another edition of this title.