126 Incidents in One Month: The First Comprehensive AI Security Inventory and What It Proves About Agent Risk
Key takeaways
- 126 AI security incidents across 38 named organizations in September 2026, with 318M+ records exposed — the first monthly inventory with named victims and quantified scope, published by RuntimeAI on October 2, 2026.
- AI was the attack vector in 39 of 126 incidents and the direct weapon in 3 — the "AI Agent Exploit" category is the single largest attack-vector class, ahead of credential theft (27), zero-day exploitation (22), phishing (10), and data exfiltration (10).
- 11 RCE vulnerabilities disclosed in one month, including Claude SDK (Anthropic), LiteLLM, Bifrost AI Gateway, and MCP Server — remote code execution in the agent infrastructure stack is now a recurring monthly disclosure, not an isolated event.
- 38 breached organizations include Microsoft, Cisco, Google, Meta, Apple, OpenAI, Anthropic, Oracle, HuggingFace, Salesforce, GitLab, Okta, Nvidia, Azure, NIST, NSA, and CISA — security vendors themselves (ESET, Okta, Cloudflare, Wiz) were present in the breach stacks.
- 36,769 self-hosted AI inference services were publicly exposed, and a malicious AI agent network stole 600,000 credit cards — the exposure surface extends from named enterprises to anyone running an unmanaged inference endpoint.
RuntimeAI's September 2026 AI Security Breach Report, published October 2, 2026, is the first monthly AI security inventory with named victims, quantified record counts, and a categorized incident catalog. 126 incidents. 38 organizations breached. 318 million records exposed. 11 RCE vulnerabilities. The report covers the full AI infrastructure stack: the Claude SDK, LiteLLM, the Bifrost AI Gateway, and MCP servers all appear as RCE targets in a single month. This article maps what the inventory proves — and why it changes how a Head of Engineering should read the governance-checklist question.
This builds on Four Labs Found the Same Agent Misbehavior, which documented that OpenAI, Anthropic, Google, and Meta each found similar agent misbehavior and that OpenAI could not enumerate its own agents' actions two months after the Hugging Face incident. That article established the pattern. What the RuntimeAI report adds is the frequency: the frontier labs' inability to inventory their own agents is not an outlier. It is one instance of a monthly cadence measured in the hundreds.
The inventory: 126 incidents, 38 organizations, one month
The headline numbers from the RuntimeAI report:
| Metric | Count |
|---|---|
| Total incidents | 126 |
| Critical severity | 22 |
| High severity | 102 |
| Organizations breached | 38 |
| Records exposed | 318M+ |
| AI-involved incidents | 53 (3 as weapon, 52 as target) |
| RCE vulnerabilities | 11 |
| CVEs referenced | 11 |
The 38 named organizations span the AI supply chain from model labs to infrastructure providers to security vendors. The full list, as published by RuntimeAI: Microsoft, Cisco, GitHub, Google, Meta, Apple, Bitget, Gemini, Okta, Nvidia, Azure, OpenAI, Anthropic, Oracle, HuggingFace, Salesforce, GitLab, LiteLLM, Revolut, NIST, NSA, CISA, McKesson, AdaptHealth, and others including the Manchester Airports Group and Burger King Russia. Four security vendors — ESET, Okta, Cloudflare, and Wiz — were present in the breach stacks. The inclusion of NIST and CISA is the sharpest signal in the report: the organizations that set security standards were themselves in the incident inventory for the month.
The severity breakdown is 22 critical and 102 high. Every incident in the catalog is classified by attack vector, AI involvement, and affected component. The catalog is not a press release — it is a structured incident database with per-incident links.
The attack vectors: AI Agent Exploit is the largest category
The attack-vector distribution is the part of the report that should change how a Head of Engineering prioritizes:
| Attack vector | Incidents |
|---|---|
| AI Agent Exploit | 39 |
| Credential Theft | 27 |
| Zero-Day / Vulnerability | 22 |
| Phishing / Social Engineering | 10 |
| Data Exfiltration | 10 |
"AI Agent Exploit" is the single largest category — 39 of 126 incidents, 31 percent of the total. This is not agents as a downstream victim of a broader breach. It is agents as the attack surface: prompt injection, tool abuse, unauthorized tool calls, and agent-mediated lateral movement. The report's framing is that the agent layer has become the primary entry point, displacing credential theft (27) and traditional zero-day exploitation (22) from the top of the distribution.
The 3 incidents where AI was the weapon — not the target — are the category that did not exist in prior monthly breach reports. RuntimeAI documents a malicious AI agent network that stole 600,000 credit cards, the Carbonato botnet deploying AI agents on hacked devices, and AI agents autonomously breaching a Spanish organization and modifying software components. These are not agents that were compromised. These are agents deployed as the attack tool. The distinction matters for governance: a kill switch that stops a compromised agent does not stop an agent that was built to attack.
The RCE layer: 11 vulnerabilities in the agent infrastructure stack
The 11 RCE vulnerabilities are the most operationally actionable finding for any team running agent infrastructure. The named targets include:
- Claude SDK (Anthropic) — remote code execution in the Claude Code CLI and Agent SDK. Check Point Research documented RCE and API token exfiltration through Claude Code project files via Hooks, MCP server configurations, and environment variables. A separate SentinelOne analysis documented CVE-2026-39861, a sandbox escape through symlink manipulation.
- Bifrost AI Gateway — CVE-2026-90898, unauthenticated command execution when management auth is disabled. A single unauthenticated POST could make Bifrost run attacker-supplied commands and expose stored LLM provider API keys. Fixed in transports/v2.1.0.
- LiteLLM — the proxy that routes LLM traffic across providers. LiteLLM's default-admin-key exposure was the first MCP CVE to land on CISA's Known Exploited Vulnerabilities catalog, documented in The First MCP CVE on the KEV List. The RuntimeAI report confirms LiteLLM as a recurring RCE target.
- MCP Server (various) — the Model Context Protocol server layer. The MCP Security Hardening Checklist documented 1,467 publicly accessible MCP servers with zero authentication and 82 percent path-traversal exposure. The RuntimeAI report confirms MCP servers as an RCE class, not a single CVE.
The pattern across the 11 RCEs is consistent: the agent infrastructure layer — the SDKs, gateways, proxies, and protocol servers that sit between the model and the enterprise — has the attack surface of early web infrastructure and none of the hardening. A team deploying agents in production is deploying a stack where RCE in the gateway, the proxy, or the SDK is a monthly disclosure, not a rare event.
The exposure surface: 36,769 self-hosted inference services
Beyond the named organizations, the report quantifies the unmanaged exposure surface: 36,769 self-hosted AI inference services were publicly accessible. These are inference endpoints run without authentication, rate limiting, or network segmentation — the AI equivalent of an open database. The number is the infrastructure counterpart to the 1,467 unauthenticated MCP servers documented in the security hardening checklist: the deployment pattern for AI infrastructure is "expose first, secure later," and the later is not happening at the rate the exposure is.
For a mid-market B2B company, the 36,769 figure is the risk arithmetic made concrete. If your team stood up a self-hosted inference endpoint for a pilot and did not put it behind authentication, it is in this count or its equivalent for the following month. The cost of an unauthenticated inference endpoint is not hypothetical once the monthly inventory measures the exposure surface in the tens of thousands.
What the inventory changes
Before the RuntimeAI report, agent security was an incident-by-incident narrative: the Hugging Face breach, the Medicare intrusion, the DNS sandbox escape, the four-labs scope review. Each incident was a story, and the response to each was a patch. The RuntimeAI report converts that narrative into a frequency statement. 126 incidents in one month means the question is no longer "will our agent deployment have a security incident?" but "when, and which of the 11 RCE classes will it be?"
The governance implication is direct. The AI Agent Governance Checklist and the MCP Security Hardening Checklist already enumerate the controls: transport authentication, tool registration verification, runtime isolation, audit trails, kill-switch test records. The RuntimeAI report is the evidence that those controls are not optional hardening. They are the difference between a deployment that appears in next month's inventory and one that does not. The 38 named organizations include companies with dedicated security teams. The controls failed not because they were unknown but because they were not applied to the agent layer before exposure.
The diagram below maps the RuntimeAI inventory's structure: the scale, the attack vectors, the RCE layer, and the control gap the inventory exposes.
The September 2026 AI security inventory at a glance:
Related reading
- Four Labs Found the Same Agent Misbehavior: Why Inventory Is the Control Nobody Has — the industry-wide pattern of agent incidents across OpenAI, Anthropic, Google, and Meta; this article extends that pattern from the frontier labs to 38 named organizations in a single month
- MCP Security Hardening Checklist: 1,467 Exposed Servers and the Controls That Close Them — the 12 controls across transport, authentication, tool registration, runtime, and audit that close the MCP Server RCE class the RuntimeAI report documents
- AI Agent Governance Checklist: A Pre-Deployment Review — the pre-deployment review that determines whether your deployment appears in next month's inventory or does not
A mid-market B2B company running 200 RFQs a week through an AI agent connected to NetSuite and three supplier catalogs is not in the 38 named organizations. But the agent infrastructure stack it runs — the LLM proxy, the MCP server, the SDK — is the same stack that produced 11 RCE vulnerabilities in September. The cost of not applying the hardening checklist to that stack before exposure is no longer an internal gap. It is a frequency statement measured in 126 incidents per month, and the difference between a deployment that appears in the next inventory and one that does not is whether the controls were applied before the endpoint went live.
One-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.
Request a scoped build.
Want this built for your systems?
Every document here comes from real production work. If you have a target system and a workflow in mind, we can scope a build in one week.
Request a scoped buildOne-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.