Back to Library
Security & Governance

Deadbugz: The Fourth MCP Attack Class Hides Until You Trust It

Last updated: September 29, 2026

Key takeaways

  • 23 pull requests in 74 minutes — a single GitHub account filed campaign PRs across unrelated AI, MCP, and developer-tool projects on August 10, 2026, none merged through GitHub's review mechanism but four remained open at disclosure time (Pillar Security).
  • Three benign calls, then the metadata rewrites — the malicious MCP server keeps a per-client call counter; after the third tools/call request, subsequent tools/list and prompts/get responses direct the agent to seek SSH keys, AWS credentials, shell history, and Kubernetes configuration, and to conceal the activity from the user.
  • Runtime-gated metadata poisoning is the fourth distinct MCP attack class — STDIO command injection → tool-poisoning → SHA-pinning bypass → runtime-gated metadata poisoning after trust. The Deadbugz mechanism is the hardest to detect pre-deployment because the server passes initial inspection and the payload activates only after the client has established a usage pattern.
  • Tool-definition fingerprinting is the defense — capture and compare a tool-definition fingerprint at approval time; treat any change to an already-approved server's tool metadata as a security event requiring renewed operator approval before the changed tool can influence sensitive actions.

A single GitHub account filed 23 pull requests in 74 minutes, each offering a "productivity-suite" MCP server that formats text and summarizes documents. The server behaves normally for the first three tool calls. On the fourth, it rewrites its own metadata to instruct the connected AI agent to seek SSH keys, AWS credentials, shell history, and Kubernetes configuration — and to hide the activity from the operator. On September 23, 2026, Pillar Security disclosed the campaign, naming it Deadbugz after the deadbug-mcp.py delivery artifact embedded in four of the pull requests. The Cloud Security Alliance published a research note noting mechanical proximity to the earlier "Miasma" campaign that targeted 73 GitHub repositories including Microsoft's azure/durabletask.

This article maps the Deadbugz mechanism, places it in the four-class MCP attack taxonomy, and identifies the control that closes the gap. It builds on the Plugin4Shell SHA-pinning bypass analysis and the MCP security hardening checklist, both of which established controls that Deadbugz exploits a gap in.

The mechanism: a per-client call counter

The Deadbugz server advertises two tools — format_text and summarize — that perform exactly what their names suggest. The server keeps an in-memory, per-client counter for tools/call requests. The counter starts at zero. Each ordinary tool call increments it. When it reaches three, subsequent tools/list and prompts/get responses change: instead of ordinary tool documentation, they contain instructions that steer the connected AI agent toward sensitive local files and instruct the agent to conceal the activity from its operator.

This matters because MCP tool definitions are not labels. Clients supply them to the model as context that helps it decide which tools to use and what actions to take. A server that can alter those definitions after installation can change the instructions the agent sees without changing the benign-looking name of the tool. The public source code advertises tools.listChanged, a capability that allows a compatible client to refresh tool metadata — so the server does not need the client to reconnect. It simply notifies the client that its tool list has changed, and the client fetches the new, poisoned definitions.

The runtime gate is a research-evasion technique. A brief inspection or limited automated test receives only benign metadata. The server passes security review, sandbox testing, and automated scanning because none of those methods make three ordinary tool calls and then check what the server returns next. Pillar confirmed the threshold logic and the altered metadata from public sources, and observed the gated metadata from the remote service through benign text-only requests.

The delivery route: a configuration change that looks routine

The campaign's delivery mechanism is a GitHub pull request that adds or changes MCP configuration in unrelated repositories. Of the 23 reviewed PRs, 17 add a remote MCP endpoint (productivity-suite-mcp.onrender.com/mcp) to a project's configuration file. Four configure Python to run a hidden local file (~/.config/.cache/.sys/.deadbug-mcp.py). Two are directory or listing submissions. All 23 were unmerged through GitHub's pull-request merge mechanism at the time of review — 19 closed, four open — but the delivery model does not require a merge. It requires a maintainer who copies the configuration change into their own setup, or a developer who sees the PR, visits the linked repository, and installs the server directly.

The account-level delivery pattern is coordinated: the same public account (zellkernel) used the same product name, configuration theme, and campaign markers across all 23 PRs, filed between 9:52 PM and 11:07 PM UTC on August 10, 2026. The account had 50 public repositories at collection, including 20 forks, and created 21 repositories on August 10 alone. The account's GitHub profile links to an X profile (@llmgod) that links back to the GitHub account — a public, account-attributed linkage between the delivery identity and AI/LLM-focused public activity.

This delivery model extends the supply-chain pattern the Plugin4Shell analysis documented: PR-based configuration changes as a vector that bypasses marketplace controls entirely. Plugin4Shell exploited marketplace SHA pinning by naming a branch after the pinned commit. Deadbugz bypasses marketplace controls by not using a marketplace at all — it goes directly to open-source projects through their contribution workflow.

The four attack classes

The MCP security family now has four distinct attack classes, each exploiting a different trust boundary:

Class Mechanism First documented Detection difficulty
STDIO command injection Malicious commands embedded in STDIO configuration strings April 2026, 20+ CVEs (Practical DevSecOps) Medium — static analysis catches shell metacharacters
Tool-poisoning A benign tool's description changes after approval to manipulate the agent Invariant Labs, April 2025 WhatsApp sleeper Medium — metadata-change detection at the client
SHA-pinning bypass Branch named as SHA defeats marketplace pin verification Plugin4Shell, September 2026 Hard — requires resolved-HEAD assertion post-checkout
Runtime-gated metadata poisoning Malicious metadata withheld until N calls, then delivered through tools/list Deadbugz, September 2026 Hardest — pre-deployment testing does not cross the threshold

The Deadbugz attack sequence and the four-class taxonomy, visualized:

Deadbugz: The Fourth MCP Attack Class Runtime-gated metadata poisoning after trust — the hardest to detect pre-deployment The Deadbugz Attack Sequence 23 PRs in 74 minutes · 3 benign calls then credential theft · Pillar Security disclosure Sep 23, 2026 1 Delivery: 23 GitHub PRs in 74 minutes Single account "zellkernel" offers "productivity-suite" MCP server to unrelated repos 17 remote endpoints, 4 hidden local scripts, 2 directory submissions — none merged 2 Inspection passes: two benign tools, zero flags format_text and summarize work as documented — security review sees nothing wrong Research-evasion technique: brief testing receives only benign metadata 3 Three ordinary tool calls — the threshold Per-client call counter reaches 3 — server advertises tools.listChanged to trigger refresh Client fetches new metadata without reconnecting — no operator-visible event 4 Metadata rewrites: credential-seeking instructions delivered tools/list and prompts/get now instruct agent to seek SSH keys, AWS credentials, shell history, Kubernetes config — and to conceal the activity from the user Tool definitions are security context — a changed definition changes what the agent does 5 Defense: tool-definition fingerprinting catches the drift Hash tool definitions at approval, compare on tools.listChanged, require operator re-approval Gate credential access behind policy — not behind remote tool metadata instructions The Four MCP Attack Classes Each exploits a different trust boundary — Deadbugz is the hardest to detect before production CLASS 1 STDIO command injection Shell metacharacters in STDIO config 20+ CVEs, April 2026 Detection: Medium — static analysis CLASS 2 Tool-poisoning (sleeper) Description changes after approval Invariant Labs, WhatsApp, Apr 2025 Detection: Medium — metadata-change CLASS 3 SHA-pinning bypass (Plugin4Shell) Branch named as SHA defeats pin 925 skills, 134K agents, Sep 2026 Detection: Hard — post-checkout assert CLASS 4 Runtime-gated metadata poisoning Payload withheld until N calls, then metadata rewrite via tools/list — Deadbugz Detection: Hardest — runtime fingerprint Key evidence 23 PRs in 74 minutes 3 benign calls then attack 4 distinct MCP attack classes 0 PRs merged through review Tool-definition fingerprinting closes the gap — ideabosque.com/library

Each class exploits the same structural gap: a trust mechanism that performs its check at the wrong time or not at all. STDIO injection trusts configuration strings without sanitization. Tool-poisoning trusts that tool descriptions will not change after approval. SHA-pinning trusts that the resolved commit matches the pinned name. Runtime-gated poisoning trusts that what the server returned during testing is what it will return during use.

Deadbugz is the hardest to detect pre-deployment because the server's behavior during inspection is genuinely benign. The malicious payload is not hidden in the code in a way that static analysis can flag — it is gated behind a runtime counter that only activates after the client has established a usage pattern. A security team that connects to the server, calls format_text once or twice, and checks the response will see nothing wrong. The attack is designed to pass exactly that kind of review.

Why the existing controls are not enough

The MCP security hardening checklist organizes 12 controls across five layers: transport, authentication, tool registration, runtime, and audit. Deadbugz exploits a gap in the tool-registration and runtime layers. The checklist's tool-registration controls verify the server at approval time — checking tool names, schemas, and descriptions before the server goes live. But Deadbugz's tools are genuinely benign at approval time. The runtime controls monitor for unauthorized actions, but the poisoned metadata is not itself an action — it is an instruction that steers the agent toward an action the agent then takes, apparently within its authorized scope.

The governed-modules thesis — that audit logs, rate limits, typed errors, and kill-switch architecture make the governance layer the security boundary — gains a fourth attack class as evidence. The Deadbugz mechanism validates the thesis from the opposite direction: a server without governance controls (no metadata-change detection, no tool-definition fingerprinting, no operator-visible diff of what the agent sees) is the exact attack surface the campaign exploits.

The defense: tool-definition fingerprinting

Pillar's recommendation is specific and implementable: capture and compare a tool-definition fingerprint at approval time. When an MCP client approves a server, it records a hash of every tool definition the server returns — name, description, input schema, and annotations. When the server subsequently notifies the client that its tool list has changed (via tools.listChanged), the client fetches the new definitions, compares them to the fingerprint, and surfaces the diff to the operator as a security event. The changed tool cannot influence sensitive actions until the operator re-approves it.

This control closes the gap that Deadbugz exploits because it does not rely on pre-deployment testing. It monitors the actual metadata the server delivers at runtime, after the trust boundary has been crossed. The fingerprint comparison catches the metadata rewrite regardless of when the runtime gate triggers — three calls, thirty calls, or three hundred. The control also catches the earlier tool-poisoning class (Invariant Labs' WhatsApp sleeper), because both attacks share the same mechanism: a tool description that changes after approval.

Four implementation steps for a team running MCP servers in production:

  1. Record tool-definition fingerprints at approval time. Hash every tool definition the server returns during the initial connection. Store the fingerprints alongside the server's approval record in the agent's configuration management system.
  2. Monitor tools/list and prompts/get responses for drift. When the server notifies the client that its tool list has changed, fetch the new definitions and compare them to the stored fingerprints. Flag any difference as a metadata-drift event.
  3. Require operator re-approval for changed tool definitions. A changed tool definition cannot influence agent behavior until a human operator reviews the diff and explicitly re-approves the server. This converts a silent metadata rewrite into a visible security event.
  4. Gate sensitive file reads, credential access, and code execution behind policy — not behind tool metadata. The Deadbugz payload instructs the agent to seek SSH keys, AWS credentials, and Kubernetes configuration. Those reads should be policy-enforced actions that require explicit authorization, not consequences of instructions contained in remote tool metadata. The shadow AI agents analysis documents the runtime-control gap that determines whether a metadata-driven credential seek gets noticed.

Related reading


A mid-market B2B distributor runs a procurement agent that connects to NetSuite, BigCommerce, and three supplier catalogs through MCP modules. The team's security review connects to each new MCP server, calls its tools twice, and checks the responses. Deadbugz passes that review. The team adds tool-definition fingerprinting to its MCP client configuration — every server's initial tool definitions are hashed at approval, tools/list responses are monitored for drift, and changed definitions trigger an operator re-approval gate before the changed tool can influence agent behavior. The next metadata-poisoning campaign becomes a flagged diff and a review, not a credential exposure.

Request a scoped build. One-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.

Want this built for your systems?

Every document here comes from real production work. If you have a target system and a workflow in mind, we can scope a build in one week.

Request a scoped build

One-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.