Cloudflare (Oct 7) describes the multi-agent system behind its Managed Defense alert triage (an early beta). Its first single-agent prototype failed in three ways:
- "Context became authority." A detection is a guess, not proof, and the agent blurred the two.
- "Scope drifted." The agent could query the wrong account, time range or source. "You can't rely on a language model prompt to be a boundary."
- "Failure disappeared." A lookup that timed out looked the same as "checked and not found".
The new design:
- Data collection has no AI in it. Fixed code makes versioned API calls and stores every item with its source, version and timestamp. The same snapshot can be replayed, so when two runs disagree, the cause is interpretation, not different inputs.
- Narrow agents. A coordinator runs 4 specialist agents in parallel (traffic, customer history, global telemetry, threat intel). The agent that combines their findings "can't fetch new evidence or choose a classification outside the approved vocabulary".
- Citations are checked in code. Every finding must cite an item in a versioned evidence package. Code checks that each citation exists, belongs to this investigation and supports the claim.
- Three states for missing data: "not checked", "checked, no match" and "checked, evidence supports absence". When the evidence isn't enough, it makes no classification.
- Tenant scope is fixed in code before any model sees data. The global-telemetry agent sees only aggregates.
- Stages are checkpointed (Cloudflare Workflows), so a failed stage reuses evidence that already passed validation instead of starting over.