Runbook Structure for Agentic System Outages
Agent failures hide in plain sight, leaving traditional monitoring blind to the damage.

A page arrives at 03:14. No error rate moved, no latency graph spiked, and no exception showed up in any log the on-call engineer knows to check. This is the structural gap at the center of agentic system reliability: the runbooks built for deterministic software assume that failures throw exceptions, emit error codes, or cause visible degradation that trips an alert, and agentic systems break that assumption at the root. Reliability for an AI agent extends past availability and latency into decision quality, task completion fidelity, cost-per-operation bounds, and whether the system knows to stop and ask a human, as the discipline of SRE for AI agent systems describes it. None of those dimensions are things a CPU graph can show.
An agent's execution path is not fixed at deploy time. A traditional runbook assumes a fixed failure surface: known error codes, known retry paths, known thresholds. An agentic system's surface reshapes itself with every execution, so when you run the first move in a traditional runbook, checking the error rate, it often checks nothing relevant.
Two documented incidents make the gap concrete. There was no attacker, no exception, no alert: the agent was doing its job by the only measure traditional monitoring had been built to watch, as Cyera's research on agent-inflicted damage documents. In both cases, the systems watching for failure reported nothing wrong, because no one had built them to detect what had happened.
The conclusion is structural. A runbook that opens with "check the error rate" does not run slowly against agentic failures, it runs never: the failure modes that matter most, the ones that do the most damage, are precisely the ones that leave no trace in the telemetry a traditional runbook was built to read. Fixing that requires rebuilding the runbook's first steps, not tuning its thresholds.
The silent failure problem that makes detection the hardest step
The most dangerous failure class in agentic systems is the agent that keeps going: asserting task completion while the environment state says otherwise, producing plausible output on top of a tool error it never acknowledged, accumulating small deviations across a long trajectory without any single step crossing a threshold that would trip an alert.
That accumulation has a specific mechanism: no individual deviation is large enough to trigger a hard error. The agent's own self-regulation also tends to fail silently under this entropy: token overrun, runaway loops, budget exhaustion, and unbounded retries are symptoms of the same underlying accumulation, consuming resources while emitting no error signal.
Trajectory analysis across tau-bench and AppWorld shows agents often say they finished when the environment state disagrees, and in single-control domains this false-success pattern is the dominant failure type. The postmark-mcp package on npm behaved identically to the legitimate Postmark connector library it cloned across many releases, until a single added line began silently copying every outgoing email to the attacker's server. No error fired, no alert triggered, and no metric a traditional runbook watches ever deviated.
The engineering conclusion follows directly: detection cannot wait for a threshold to breach, because the failures that matter never breach one. Watching for noise will not catch a system that fails by going quiet. Detection has to be built around continuous instrumentation of what the agent intended versus what it actually did, checked at every span of the trace, rather than bolted onto the alerting thresholds built for a different kind of software.
The five failure categories a runbook must distinguish before prescribing any action
Conflating failure categories is what turns agent incident reviews into theater. The five-category taxonomy exists to stop that cycle, and identifying the category is the first mandatory step in any agentic runbook, not a retrospective exercise performed after the fire is out.
Planning errors occur when the agent picks the wrong tool sequence, loops, skips a required step, or keeps executing past the point where the plan should have terminated. The trace fingerprint is a plan span whose emitted tool sequence does not match what actually ran, or a goal evaluation that shows no progress across several steps in a row.
Tool errors happen when the plan was correct but a tool call returned malformed output, was invoked with the wrong arguments, or hit an error path the agent never handled, so it kept operating on fabricated data. The trace fingerprint is a tool call that returns a non-2xx or malformed payload, followed by an LLM response that never acknowledges the failure occurred.
Retrieval errors happen when the agent retrieves the wrong chunk, context goes missing, or a role-switch token leaks into the retrieved payload. The trace fingerprint is a grounding score that drops while the generator's output length and confidence stay flat.
Reasoning errors cover factual mistakes, math errors, ignored instructions, or hallucination layered on top of grounded context. The trace fingerprint is a generator span producing a confident, well-formed output that fails a grounding check against the retrieval payload it was actually given.
Safety and policy violations cover PII leaks, jailbreak compliance, out-of-policy commitments, harmful output, or refusal failure. The fix lives in the output scanner, a refusal eval, or an adversarial pushback set. The trace fingerprint is an output that passes every format check but fails a policy rubric.
The Microsoft AI Red Team's updated taxonomy, published June 2026, adds seven categories for agentic AI systems more broadly: Agentic Supply Chain Compromise, Goal Hijacking, Inter-Agent Trust Escalation, Computer Use Agent Visual Attack, Session Context Contamination, MCP/Plugin Abuse, and Capability/Architecture Disclosure. Each of these maps onto the planning, tool, or policy fix surfaces already described, so it doesn't open a separate remediation track, and the five-category model still holds as the operational backbone even as the vocabulary around it grows. A complementary framing, Greyling's harness-centric taxonomy, organizes failures by harness function instead, across Environment Contract, Operation Skills, Action Execution, and Trajectory Regulation. That gives engineers a second axis: useful when the five-category classification correctly identifies the fix surface but leaves ambiguous exactly which layer of the harness the fix belongs in.
Observability requirements the runbook depends on
A runbook built around failure categories is only as good as the trace data available to confirm which category applies, and most observability stacks built for traditional software do not capture that data. Agent observability has to record end-to-end reasoning sequences, tool calls, memory operations, and agent-to-agent handoffs, not just CPU, latency, and error rates, as the current generation of agent observability practice describes it. A runbook referencing trace attributes the stack does not emit is a flowchart with no inputs, and discovering that gap during an active incident is the worst possible time to find it.
A handful of span attributes carry most of the diagnostic weight during an incident. Per-user cost attribution closes a third gap: runaway users, malformed integrations, and abusive callers become visible per user within a day of cost attribution, and at the latest by the time the monthly invoice rolls up.
The requirement that follows is structural. Observability has to ship with the agent at launch, not get bolted on after the second incident makes clear what was missing the first time. A well-built runbook states, before prescribing any action, which of these signals the responding engineer should expect to find, so an engineer reading it at three in the morning knows immediately whether the data needed to classify the failure actually exists.
Walking a trace to confirm failure category during an active incident
A loop incident makes the clearest starting point, because it shows how two failures that look identical from the outside require opposite diagnostic moves. When a loop-count alert fires, or a task's duration crosses a meaningful multiple of its p95 baseline, the first action is to pull the in-progress trace for that task. The finish_reason field on the generation span settles the question quickly: if it cycles through tool-use responses without ever completing, the loop is a reasoning loop; if the tool call itself returns an error or malformed output on every pass, the loop is a tool error loop.
That same discipline, walking the span tree in a fixed order, applies past loops to any multi-agent failure. The sequence runs from input to plan to retrieval to tool arguments to tool output to LLM rewrite to final answer, and the responding engineer scores each span in that order to find the first one that produced bad output. A complementary lens, treating the trace as a sequence of behavioral segments rather than a list of raw span attributes, helps specifically with reasoning errors, where the output is syntactically well-formed and only wrong in a semantic sense that a flat log scan will not surface.
Confirming a category this way only holds up if the finding can be tested. That is where the replay requirement comes in. Given any trace URL from any point in the system's history, the responding engineer needs to be able to reconstruct what the agent saw, the rendered prompt, the tool outputs, and the retrieval payload, then re-run the agent in a sandbox and validate a candidate fix against that historical trace before the fix ships. If a runbook prescribes a fix without that replay step, you get a postmortem that says the team thinks the fix works, not one that shows the fix replayed cleanly against the failed traces from the day before.
None of this holds up under pressure if the runbook is written in prose that an engineer has to interpret at three in the morning. A runbook written as narrative is a runbook nobody actually reads when the page comes in, and the formatting discipline here is what makes the five-category classification something an engineer can execute under pressure rather than something they have to reconstruct from memory.
Routing remediation to the right surface once the category is confirmed
A fix only pays off if it lands on the surface that category actually points to, and that routing is neither obvious nor automatic. A planning error patched at the output scanner, a tool schema error patched in the prompt, or a retrieval error patched at the generator each produces a fix that looks like progress for a few weeks and then ages out, because it addresses a symptom at the wrong layer of the stack than the mechanism that produced it.
Planning errors route to the planner prompt, where explicit termination conditions can be added, to a cycle detector flagging the same tool called with similar arguments more than three times in a single trace, and to a plan-versus-execute evaluator gating the next request before it fires. Two of the agents, an Analyzer and a Verifier, locked into a clarification ping-pong with no budget cap, no termination condition, and no cycle detector in place, and per-call rate limits did nothing to stop it, because rate limits address volume, not the structural absence of a stopping condition.
Retrieval errors route to the retriever rubric, chunker re-evaluation, and ingestion controls that prevent role-switch tokens from entering retrieved payloads in the first place, not to the generator's system prompt. Reasoning errors route to the generator rubric, an LLM-judge eval run against sampled production traffic, and an updated calibration set, not to the retriever or the output scanner. Safety and policy violations route to the output scanner, a refusal eval, and an adversarial pushback set, while security violations, meaning failures originating in inter-agent trust rather than the final output, route to inter-agent controls, identity verification, and trust boundary enforcement. The CASE framework for enterprise agentic AI governance offers a useful reference point here, since its control layers map onto the same planner, tool, retrieval, generator, and output surfaces, which makes it a workable template for assigning escalation ownership inside a runbook.
Escalation itself follows the category, not the symptom. A tool error traced to a supply chain compromise, such as MCP or plugin abuse, escalates to security. A reasoning error surfacing on grounded context escalates to the model team rather than the data team, because the data was correct and the failure happened downstream of it.
The return on getting this routing right compounds with every incident handled this way. When a named failure category routes to a named fix surface, that fix becomes a labeled dataset entry and a regression test inside CI: the next pull request has to clear it, and the next incident review can determine in minutes whether a new failure is a regression of something already known or a genuinely new mode. If a runbook skips classification and jumps straight to a fix, it resets the team's knowledge with every incident, because nothing from the last one accumulates into defense against the next. A runbook built around classification turns every incident into infrastructure for the one after it.
Closing the incident loop with evals that prevent the same category from recurring
A runbook that ends the moment a fix ships has only managed the incident in front of it.
To close the loop on a specific incident, you add a small, specific set of checks. Format, length, grounding, and refusal checks run on every turn at low latency, so they catch the output-layer symptom common to any of the five categories. And when eval scores live on the same span as latency and token counts, the next on-call engineer sees quality and reliability together in one place instead of stitching the two together from separate tools at three in the morning.
Benchmarks alone will not substitute for this. The regression suite a team builds out of its own production incidents ends up more predictive of real failure than any public benchmark, because it is built from the failures that already happened in this system.
The same discipline extends backward from response into prevention. The goal is to classify and remediate the drift before it becomes an incident, not during one.
This closes the argument the piece opened with. The page that arrives at 03:14 with no error metric moved becomes less likely over time, because the categories behind it are now checked continuously. And when that page does arrive anyway, it resolves faster, because the runbook already carries a name for the failure, a surface to route it to, and a trace pattern to check it against, from the last time the team ran through exactly this loop.
Sources
- Microsoft Updates Taxonomy of AI System Failure Modes
- Silent Failure in LLM Agent Systems: The Entropy Principle and the Inevitable Disorder of Autonomous Agents
- Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
- The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI
- Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions

