Security alerts rarely arrive one at a time. A single alert can cause a spike across the environment, requiring a human analyst to decide which alerts are related and what they mean. When multiple arrive at the same time, it can quickly overwhelm even a seasoned security analyst.
Enter the alert paradox. Now, our built-in, multi-AI-agent security operations harness can handle more of this work at Cloudflare scale. Our Cloudflare Managed Defense AI agent harness speeds up the process of gathering data, connecting and aggregating detections, and accounting for missing sources while new alerts continue to arrive.
To further analyze context, we make use of the OpenAI Daybreak Defense Network and our partnership with Anthropic. Cloudflare uses approved OpenAI Daybreak and Anthropic models, including GPT-5. 6 Cyber and Mythos, for deeper model-backed analysis.
Initial analysis and scoring is done with Clef , Cloudflare’s open-source decision model. Collecting evidence to understand what each alert means requires a lot of time. Even in a highly sophisticated Security Information and Event Management (SIEM), too much is still left for a human to review.
Human analysts must gather data, connect and aggregate detections, and account for missing sources while new alerts continue to arrive. Think about every time a human analyst reviews an alert: "Which should we silence? Which action should we take?
Which alert should we ignore? Which should we resolve as false positives? Which should we resolve as true positives?
Which should trigger our incident team?" We address this predicament with our AI agent strategy. Our approach reduces all of those questions and gives Managed Defense Analysts a quick and consolidated view, directly providing insight into related alerts, admitted evidence, visible gaps, and recommended next steps.
The result: cutting back on the time needed to analyze, and creating a hyper focus on actually getting security alerts resolved and mitigations deployed. Why a single agent fails Our first prototype showed the limits of one general-purpose agent. We provided the AI agent the whole investigation.
It produced useful analysis, but it also hallucinated claims the evidence did not support. Telemetry, detector descriptions, policies, and threat intelligence were flattened into one prompt, which caused their distinct roles to merge together. We saw three recurring problems with our first single-shot AI agent harness: Context became authority .
A detection is a hypothesis, not proof that an exploit succeeded or an attack occurred. A broad AI agent can blur that distinction. Scope drifted.
An AI agent can query the wrong account, time range, or source. You can’t rely on a language model prompt to be a boundary. Failure disappeared.
If a lookup times out, the result may not distinguish "not checked" from "checked and not found." To address these challenges we moved evidence collection and scope enforcement into application code, before model analysis begins. Recon first, inference second It's tempting to put an AI agent at every step.
The front half of our harness has none. Before we even call inference, deterministic code runs a fixed set of reconnaissance workflows with versioned API calls. It collects the customer's identity, detection history, traffic baseline, enforcement outcome, and network observations.
Each piece of data is stored with its source, version, and timestamp. Cloudflare sees both the request and the action applied to it. That lets the investigation connect the behavior that triggered an alert with both the control that fired and its outcome.
The fixed recon snapshot also makes evaluation reproducible. If AI agents fetch their own data, two runs may disagree because their inputs changed. Here, the same snapshot can be replayed, so differences between specialist AI agents’ findings come from interpretation rather than retrieval.
Filter noise early Most alerts are not incidents. The same rule often fires repeatedly on a known traffic pattern, and paging Managed Defense Analysts every time makes it easier to miss a real security incident. We needed a lightweight triage model to compare each alert with its reconnaissance data: Has this event been detected for the customer before?
What did Managed Defense Analysts decide previously? Does the traffic look consistent with normal human behavior? Alerts scored with a high likelihood to be false positives skip analysis by the specialist AI agents.
Clef , running on Workers AI, was the perfect fit for this type of fast agentic reasoning. Known high-volume noise is deterministically classified as passive when it arrives. It remains available as context but does not enter the active queue.
Specialist AI agents handle the investigation For alerts that need deeper review, a coordinator AI agent runs four specialist AI agents in parallel: Traffic analysis reviews request behavior, historical changes, and enforcement. Customer context reviews earlier alerts, dispositions, and Managed Defense Analysts’ decisions. Global telemetry compares the activity with privacy-preserving Internet-wide signals.
Threat intelligence checks indicators already admitted to the alert or case. A synthesis AI agent combines their typed findings into one advisory; it can’t fetch new evidence or choose a classification outside the approved vocabulary. Keeping each task narrow makes unsupported claims easier to catch, and recommendations easier to audit.
Global context without customer data A security tool knows what happened inside the environment it is deployed in, but little about the world beyond it. Cloudflare compares an alert with patterns seen across its global network. For example, an IP may be targeting one site, scanning thousands of sites, or appearing for the first time.
Those patterns carry different weights.
Originally published at blog.cloudflare.com


