Deadbugz: The Fourth MCP Attack Class Hides Until You Trust It
Key takeaways
- 23 pull requests in 74 minutes — a single GitHub account filed campaign PRs across unrelated AI, MCP, and developer-tool projects on August 10, 2026, none merged through GitHub's review mechanism but four remained open at disclosure time (Pillar Security).
- Three benign calls, then the metadata rewrites — the malicious MCP server keeps a per-client call counter; after the third
tools/callrequest, subsequenttools/listandprompts/getresponses direct the agent to seek SSH keys, AWS credentials, shell history, and Kubernetes configuration, and to conceal the activity from the user. - Runtime-gated metadata poisoning is the fourth distinct MCP attack class — STDIO command injection → tool-poisoning → SHA-pinning bypass → runtime-gated metadata poisoning after trust. The Deadbugz mechanism is the hardest to detect pre-deployment because the server passes initial inspection and the payload activates only after the client has established a usage pattern.
- Tool-definition fingerprinting is the defense — capture and compare a tool-definition fingerprint at approval time; treat any change to an already-approved server's tool metadata as a security event requiring renewed operator approval before the changed tool can influence sensitive actions.
A single GitHub account filed 23 pull requests in 74 minutes, each offering a "productivity-suite" MCP server that formats text and summarizes documents. The server behaves normally for the first three tool calls. On the fourth, it rewrites its own metadata to instruct the connected AI agent to seek SSH keys, AWS credentials, shell history, and Kubernetes configuration — and to hide the activity from the operator. On September 23, 2026, Pillar Security disclosed the campaign, naming it Deadbugz after the deadbug-mcp.py delivery artifact embedded in four of the pull requests. The Cloud Security Alliance published a research note noting mechanical proximity to the earlier "Miasma" campaign that targeted 73 GitHub repositories including Microsoft's azure/durabletask.
This article maps the Deadbugz mechanism, places it in the four-class MCP attack taxonomy, and identifies the control that closes the gap. It builds on the Plugin4Shell SHA-pinning bypass analysis and the MCP security hardening checklist, both of which established controls that Deadbugz exploits a gap in.
The mechanism: a per-client call counter
The Deadbugz server advertises two tools — format_text and summarize — that perform exactly what their names suggest. The server keeps an in-memory, per-client counter for tools/call requests. The counter starts at zero. Each ordinary tool call increments it. When it reaches three, subsequent tools/list and prompts/get responses change: instead of ordinary tool documentation, they contain instructions that steer the connected AI agent toward sensitive local files and instruct the agent to conceal the activity from its operator.
This matters because MCP tool definitions are not labels. Clients supply them to the model as context that helps it decide which tools to use and what actions to take. A server that can alter those definitions after installation can change the instructions the agent sees without changing the benign-looking name of the tool. The public source code advertises tools.listChanged, a capability that allows a compatible client to refresh tool metadata — so the server does not need the client to reconnect. It simply notifies the client that its tool list has changed, and the client fetches the new, poisoned definitions.
The runtime gate is a research-evasion technique. A brief inspection or limited automated test receives only benign metadata. The server passes security review, sandbox testing, and automated scanning because none of those methods make three ordinary tool calls and then check what the server returns next. Pillar confirmed the threshold logic and the altered metadata from public sources, and observed the gated metadata from the remote service through benign text-only requests.
The delivery route: a configuration change that looks routine
The campaign's delivery mechanism is a GitHub pull request that adds or changes MCP configuration in unrelated repositories. Of the 23 reviewed PRs, 17 add a remote MCP endpoint (productivity-suite-mcp.onrender.com/mcp) to a project's configuration file. Four configure Python to run a hidden local file (~/.config/.cache/.sys/.deadbug-mcp.py). Two are directory or listing submissions. All 23 were unmerged through GitHub's pull-request merge mechanism at the time of review — 19 closed, four open — but the delivery model does not require a merge. It requires a maintainer who copies the configuration change into their own setup, or a developer who sees the PR, visits the linked repository, and installs the server directly.
The account-level delivery pattern is coordinated: the same public account (zellkernel) used the same product name, configuration theme, and campaign markers across all 23 PRs, filed between 9:52 PM and 11:07 PM UTC on August 10, 2026. The account had 50 public repositories at collection, including 20 forks, and created 21 repositories on August 10 alone. The account's GitHub profile links to an X profile (@llmgod) that links back to the GitHub account — a public, account-attributed linkage between the delivery identity and AI/LLM-focused public activity.
This delivery model extends the supply-chain pattern the Plugin4Shell analysis documented: PR-based configuration changes as a vector that bypasses marketplace controls entirely. Plugin4Shell exploited marketplace SHA pinning by naming a branch after the pinned commit. Deadbugz bypasses marketplace controls by not using a marketplace at all — it goes directly to open-source projects through their contribution workflow.
The four attack classes
The MCP security family now has four distinct attack classes, each exploiting a different trust boundary:
| Class | Mechanism | First documented | Detection difficulty |
|---|---|---|---|
| STDIO command injection | Malicious commands embedded in STDIO configuration strings | April 2026, 20+ CVEs (Practical DevSecOps) | Medium — static analysis catches shell metacharacters |
| Tool-poisoning | A benign tool's description changes after approval to manipulate the agent | Invariant Labs, April 2025 WhatsApp sleeper | Medium — metadata-change detection at the client |
| SHA-pinning bypass | Branch named as SHA defeats marketplace pin verification | Plugin4Shell, September 2026 | Hard — requires resolved-HEAD assertion post-checkout |
| Runtime-gated metadata poisoning | Malicious metadata withheld until N calls, then delivered through tools/list |
Deadbugz, September 2026 | Hardest — pre-deployment testing does not cross the threshold |
The Deadbugz attack sequence and the four-class taxonomy, visualized:
Each class exploits the same structural gap: a trust mechanism that performs its check at the wrong time or not at all. STDIO injection trusts configuration strings without sanitization. Tool-poisoning trusts that tool descriptions will not change after approval. SHA-pinning trusts that the resolved commit matches the pinned name. Runtime-gated poisoning trusts that what the server returned during testing is what it will return during use.
Deadbugz is the hardest to detect pre-deployment because the server's behavior during inspection is genuinely benign. The malicious payload is not hidden in the code in a way that static analysis can flag — it is gated behind a runtime counter that only activates after the client has established a usage pattern. A security team that connects to the server, calls format_text once or twice, and checks the response will see nothing wrong. The attack is designed to pass exactly that kind of review.
Why the existing controls are not enough
The MCP security hardening checklist organizes 12 controls across five layers: transport, authentication, tool registration, runtime, and audit. Deadbugz exploits a gap in the tool-registration and runtime layers. The checklist's tool-registration controls verify the server at approval time — checking tool names, schemas, and descriptions before the server goes live. But Deadbugz's tools are genuinely benign at approval time. The runtime controls monitor for unauthorized actions, but the poisoned metadata is not itself an action — it is an instruction that steers the agent toward an action the agent then takes, apparently within its authorized scope.
The governed-modules thesis — that audit logs, rate limits, typed errors, and kill-switch architecture make the governance layer the security boundary — gains a fourth attack class as evidence. The Deadbugz mechanism validates the thesis from the opposite direction: a server without governance controls (no metadata-change detection, no tool-definition fingerprinting, no operator-visible diff of what the agent sees) is the exact attack surface the campaign exploits.
The defense: tool-definition fingerprinting
Pillar's recommendation is specific and implementable: capture and compare a tool-definition fingerprint at approval time. When an MCP client approves a server, it records a hash of every tool definition the server returns — name, description, input schema, and annotations. When the server subsequently notifies the client that its tool list has changed (via tools.listChanged), the client fetches the new definitions, compares them to the fingerprint, and surfaces the diff to the operator as a security event. The changed tool cannot influence sensitive actions until the operator re-approves it.
This control closes the gap that Deadbugz exploits because it does not rely on pre-deployment testing. It monitors the actual metadata the server delivers at runtime, after the trust boundary has been crossed. The fingerprint comparison catches the metadata rewrite regardless of when the runtime gate triggers — three calls, thirty calls, or three hundred. The control also catches the earlier tool-poisoning class (Invariant Labs' WhatsApp sleeper), because both attacks share the same mechanism: a tool description that changes after approval.
Four implementation steps for a team running MCP servers in production:
- Record tool-definition fingerprints at approval time. Hash every tool definition the server returns during the initial connection. Store the fingerprints alongside the server's approval record in the agent's configuration management system.
- Monitor
tools/listandprompts/getresponses for drift. When the server notifies the client that its tool list has changed, fetch the new definitions and compare them to the stored fingerprints. Flag any difference as a metadata-drift event. - Require operator re-approval for changed tool definitions. A changed tool definition cannot influence agent behavior until a human operator reviews the diff and explicitly re-approves the server. This converts a silent metadata rewrite into a visible security event.
- Gate sensitive file reads, credential access, and code execution behind policy — not behind tool metadata. The Deadbugz payload instructs the agent to seek SSH keys, AWS credentials, and Kubernetes configuration. Those reads should be policy-enforced actions that require explicit authorization, not consequences of instructions contained in remote tool metadata. The shadow AI agents analysis documents the runtime-control gap that determines whether a metadata-driven credential seek gets noticed.
Related reading
- SHA Pinning Is Not Verification: Plugin4Shell and the First AI-Agent Supply-Chain RCE — the third MCP attack class, which shares Deadbugz's PR-based delivery model and supply-chain targeting
- MCP Security Hardening Checklist: 1,467 Exposed Servers and the Controls That Close Them — the 12-control baseline; tool-definition fingerprinting belongs in the tool-registration and runtime layers
- MCP Security: Why 200,000 Vulnerable Instances Make Governed Modules a Buying Criterion — the governed-modules thesis that Deadbugz validates: a server without governance controls is the exact attack surface the campaign exploits
A mid-market B2B distributor runs a procurement agent that connects to NetSuite, BigCommerce, and three supplier catalogs through MCP modules. The team's security review connects to each new MCP server, calls its tools twice, and checks the responses. Deadbugz passes that review. The team adds tool-definition fingerprinting to its MCP client configuration — every server's initial tool definitions are hashed at approval, tools/list responses are monitored for drift, and changed definitions trigger an operator re-approval gate before the changed tool can influence agent behavior. The next metadata-poisoning campaign becomes a flagged diff and a review, not a credential exposure.
Request a scoped build. One-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.
Want this built for your systems?
Every document here comes from real production work. If you have a target system and a workflow in mind, we can scope a build in one week.
Request a scoped buildOne-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.