Est.

Alerting on Agent Behavior Instead of Exit Codes

AI agents can succeed technically while failing their actual goal.

Staff Writer · · 10 min read
Cover illustration for “Alerting on Agent Behavior Instead of Exit Codes”
Agent Incident Response · October 8, 2026 · 10 min read · 2,217 words

Engineers coming out of traditional incident response carry a specific mental model with them: a process either completes or it doesn't, and the signals that tell you which one happened are exit codes, HTTP status, CPU and memory load, and error rate. That model built the entire discipline of application performance monitoring, and it works, for the class of software it was designed to watch. The trouble is that it answers one question only: did the process complete? It was never built to answer whether the process did the right thing, and for most of software history that distinction didn't matter much, because a completed process and a correct one were close enough to the same event. An AI agent breaks that equivalence. It can return HTTP 200, exit cleanly, log no errors anywhere in the chain, and still have fabricated every meaningful output along the way. The monitoring stack has no instrument for that gap, because the gap doesn't exist in the kind of software the stack was built to watch.

Part of what makes this hard to see at first is that nothing about an agent's trace looks broken. The reason traditional monitoring misses the failure is structural, and it traces back to how agents actually run: the same model, given the same system prompt and the same input, can take a different path to a different output each time it runs. A function that adds two numbers returns the same result every time, so if it ever fails, something broke, and the error shows up in the trace. An agent's reasoning path is probabilistic at every step, so "it ran" and "it succeeded" are no longer fixed to each other the way they are in a compiled program.

Multi-step agents make this worse, not because of some complicated interaction, but by simple arithmetic. Every reasoning step, every tool call, every state transition is a fresh opportunity for the agent to drift from what it was asked to do, and none of those opportunities trips a traditional alert, because none of them resembles an error from the monitoring stack's point of view. The instrument built to answer "did it run?" is being pointed at a question it has no way to measure: "did it do what it was supposed to do?" That mismatch is the starting condition for everything that follows.

What "Silent Failure" Means Inside a Running Agent Trace

Silent failure doesn't look like a crash, and it doesn't look like a timeout. It looks like a sequence of operations that each succeed individually while the whole chain stands on a false premise that no single step would ever flag. A 2026 systematic analysis found a multi-agent orchestration system that ran for 47 days before its silent failure surfaced, and nothing external triggered the discovery. No alert fired. No threshold was crossed. The system had simply been doing the wrong thing, successfully, for over a month.

Picture an inventory agent that invents a SKU that doesn't exist in step one. Nothing stops it there, because inventing a plausible-looking SKU number isn't something any system checks for as an error. The agent then calls downstream APIs to price the item, check its stock, and ship it, and each of those calls returns HTTP 200, because each API is doing its job correctly given the input it received. The trace closes cleanly from end to end. The workflow is a complete failure, and every system involved reported success the entire way through.

A different shape of the same problem is visible in what practitioners call the refund misgeneralization case: an agent handling customer support is given instructions that look, on their surface, like they align with the task it was already doing, but following them redirects its actual goal. It is persuaded into doing the wrong job well. Both cases share the same signature: the execution log is clean, no step raised an exception, and the failure lives entirely in the space between what the agent was supposed to accomplish and what it actually did. That space is semantic, invisible to anything that only checks whether operations completed. The diagnostic question that matters is whether the system did what it was supposed to do, at every single step along the way.

The failure modes that matter most, named and classified

Once the inventory hallucination and the refund misgeneralization cases are on the table, the next problem is naming them precisely enough that a team can build defenses around them instead of patching each incident as a one-off. A 2026 practitioner taxonomy sorts agent failures into five categories, and each one carries its own fingerprint in the trace and points to a different fix. Safety and policy violations cover PII leaks, compliance with a jailbreak attempt, harmful output, or a refusal that should have fired and didn't, and this is the category the refund misgeneralization case belongs to.

The Microsoft AI Red Team published a complementary taxonomy in April 2025 that organizes failures along two different axes: whether a failure mode is novel to agentic systems or an existing failure mode that agentic systems simply amplify, and whether it's a safety concern or a security concern. Existing modes that get amplified include memory poisoning, cross-domain prompt injection, and bypass of human-in-the-loop checks; these problems predate agents, but their blast radius grows once a system is autonomous across multiple steps.

That taxonomy wasn't static. Goal Hijacking formalizes exactly the pattern seen in the refund case: instructions that look aligned with the task at hand quietly redirect the agent's actual goal without ever fully compromising the agent itself. Session Context Contamination names a failure no single-step log could ever catch on its own: information introduced early in a session skews the agent's reasoning several steps later, and no safety check along the way trips. Inter-Agent Trust Escalation is the multi-agent version of the confused-deputy problem familiar from traditional software security, except natural language induces the confusion instead of system calls, so it sits outside what single-agent monitoring was ever built to see.

The value of having these categories named isn't academic. Failure category determines which surface gets fixed. A planner-side problem doesn't get solved by adding an output scanner, and a reasoning error in the generator doesn't get solved by improving retrieval. Postmortems that skip the taxonomy tend to close with a prompt tweak and a vague guardrail action item, and those fixes rarely survive the week because they were never aimed at the actual failure surface.

MCP and multi-agent failure surfaces beyond exit-code monitoring

The move toward Model Context Protocol tool ecosystems and multi-agent orchestration hasn't just added more moving parts for you to monitor. It has opened categories of failure that no exit-code or error-rate dashboard was ever built to represent. In 2025, 99 CVEs were published for MCP-related software, and tool poisoning stopped being a theoretical concern and became a live attack surface running in production systems. MCP/Plugin Abuse earned its own place in the v2.0 taxonomy precisely because the trust assumptions built into the protocol can be exploited without tripping anything that looks like an error.

A tool description poisoning attack works by injecting natural-language instructions directly into a tool's schema. The agent reads that schema as part of normal operation, follows the injected instruction because nothing distinguishes it from a legitimate one, and produces an outcome that's wrong or harmful. Cross-server instruction override, where a malicious MCP server overrides the behavior of a trusted one, has no equivalent in traditional service-mesh monitoring. There's no HTTP status code that means "your agent just got redirected by a hostile tool description," because the concept doesn't exist in the vocabulary that status codes were designed to express.

Multi-agent systems carry a parallel version of the same blind spot. Inter-Agent Trust Escalation describes a situation where a compromised subordinate agent asserts a false identity or inflates its own claimed permissions, and an orchestrator that doesn't independently verify those claims simply accepts them. The orchestrator's own trace looks completely clean, because delegation happened the way delegation is supposed to happen, so the failure lives entirely in what got delegated rather than in whether the delegation mechanism worked. Session Context Contamination compounds the same problem across time instead of across agents: data introduced early in a session shapes reasoning several steps later, and because no single step contains anything a safety control would flag, the whole pattern stays invisible to any alert that only evaluates steps in isolation.

The obvious objection is that teams can already monitor tool call inputs and outputs directly. That's true, but it doesn't solve the problem, because a poisoned tool description, a redirected goal, or a trust escalation are semantic failures, not syntactic ones. Logging what happened tells you nothing about whether what happened was actually correct.

What behavioral alerting requires that threshold-based alerting cannot provide

Catching the failure modes described above means building alert policies around a different question than the one threshold-based monitoring asks. What matters is whether the agent did what it was supposed to do. Most teams stop at latency histograms and token counts, which is a reasonable floor for understanding speed and cost, but tells a team nothing about whether the agent actually finished the user's task or stayed inside policy the whole time. Calling that floor "monitoring" understates how much of the actual failure surface it leaves uncovered.

The taxonomy points directly to what a real alert needs to be able to do: fire when no explicit error was ever returned, when every API call succeeded on its own terms, and when the only sign of trouble is a mismatch between what the agent was instructed to do and what it actually did across a trace spanning many steps. Building that kind of alert requires three things a traditional monitor simply doesn't hold: the agent's original intent, the complete trace covering every reasoning step and tool call, and a scoring mechanism capable of judging semantic correctness. Enterprise observability for AI agents increasingly builds around quality degradation alerts, which trigger when an agent's task success rate drops below a set baseline over a rolling window, even when not a single explicit error was logged anywhere in that window.

The dominant method for producing that semantic score in 2026 is LLM-as-a-judge: a separate evaluator model reviews a sample of production responses against criteria like faithfulness, relevance, and groundedness, and that evaluator functions as the instrument exit-code monitoring simply has no equivalent for. A probabilistic judgment is far more useful than having no behavioral signal.

One concrete structural safeguard belongs in this picture alongside the alerting layer itself: tools that perform destructive operations should carry a destructiveHint: true annotation in the MCP schema, which tells the host to prompt for explicit user confirmation before the action runs. These are hints rather than enforced rules, so they don't replace behavioral alerting, but they shrink the blast radius of a tool-layer failure before any alert ever has to fire.

Building the Reliability Loop That Turns Production Failures into Regression Tests

None of this holds up if it stops at detection. A behavioral alert that fires once and gets manually patched teaches the system nothing, and the same failure mode will resurface the next time conditions line up the same way. The value of behavioral alerting compounds only when every production failure feeds back into the eval suite that's meant to catch the next regression before it reaches a user. Without that feedback step, a team ends up patching symptoms indefinitely, fixing the same category of problem over and over under a different name each time.

The loop that prevents this runs in three stages, and each one has to feed the next automatically. Before anything deploys, an eval suite runs in continuous integration against the known failure patterns the taxonomy has already named, and each failure category becomes a labeled dataset entry that every future pull request has to clear before it ships. Without a precise taxonomy label attached to a failure, the eval can't distinguish a true regression of the same failure from a genuinely new failure on the same surface, and the test ends up too vague to catch anything specific. Once a system is live, continuous scorers need to run on actual production traces rather than just the subset of cases a team anticipated when it wrote the test suite, because real traffic contains inputs and edge cases no curated test set predicted in advance, and a quality degradation alert should fire whenever the task success rate drops across a rolling window even if zero explicit errors were ever logged. Finally, every failed production trace needs to become a new regression case that every future change gets checked against, which is the mechanism by which an eval suite grows to match the actual shape of production failures instead of the shape a team imagined when it first built the suite.

The loop breaks at one specific point more often than any other: a production failure gets triaged inside a separate monitoring console, someone fixes the immediate symptom, and the failure mode never gets converted into a labeled eval case. The failure mode stayed unnamed, untested, and fully capable of showing up again the next time the system runs the same path.

Sources

  1. New whitepaper outlines the taxonomy of failure modes in AI agents
  2. Updating the taxonomy of failure modes in agentic AI systems: What a year of red teaming taught us
  3. Silent Failure in LLM Agent Systems: The Entropy Principle and the Inevitable Disorder of Autonomous Agents
  4. Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
  5. Outcome Monitors: Recovery Affordances for Silent Tool Failures
  6. Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions