Logan Kelly

MCP Governance: How to Secure AI Agents in a Plugin-Driven World

MCP Governance: How to Secure AI Agents in a Plugin-Driven World

91.8% of internet-facing MCP servers audited in 2026 had no OAuth. What MCP governance actually is: allowlisting, result inspection, sequence policy.

Black blog cover image with subtle grid pattern. Category label reads "MCP / PLUGINS" in the upper left. Large headline text reads "MCP Is Changing Everything." Waxell logo in the bottom right corner.

There's a before-and-after line in the history of AI agent deployments, and Model Context Protocol is it.

On April 20, 2026, CIS published its MCP Companion Guide — a set of controls co-authored with Astrix Security and Cequence Security that extends the CIS Controls into MCP environments. CIS describes that guide as detailing “protections for Model Context Protocol environments, emphasizing secure tool access, management of Non-Human Identities (NHIs), and auditable interactions across the protocol layer.” Four days earlier, The Register reported that researchers at Ox Security had documented a design flaw in the official MCP implementation's STDIO transport that, according to their analysis, puts as many as 200,000 servers at risk, with, in the researchers’ account, “10 (so far) high- and critical-severity CVEs issued for individual open source tools and AI agents that use MCP.” The Ox team says Anthropic “declined to modify the protocol’s architecture, citing the behavior as ‘expected’”; Anthropic did not respond to The Register’s inquiries for that story, so the characterization is the researchers’, not Anthropic’s own. The gap is not theoretical: a security assessment published on July 31, 2026 dynamically audited 414 internet-facing production MCP servers and found that 91.8% of them had no OAuth authentication at all.

MCP governance is the practice of defining and enforcing policies on AI agents that use Model Context Protocol to connect to external tools and services. It is not the same thing as MCP security — hardening the servers themselves — and it is not observability, which tells you what an agent did after it did it. Because MCP dramatically expands an agent's capability surface, connecting it to file systems, databases, APIs, and communication platforms through a standardized protocol, it proportionally expands the attack surface and the governance requirements. MCP governance operates at three layers: tool allowlisting and scoping (which tools can be called, under what conditions), tool result inspection (scanning responses before they enter agent context), and sequence-level policy evaluation (tracking cross-tool patterns within a single agent's run). For a broader definition of the governance plane that MCP sits within, see What is agentic governance →.

Before MCP, agents were powerful but bounded. What they could do in the world was limited by whatever tools a developer explicitly wired up — dangerous if the tools were dangerous, but bounded. After MCP, an agent's capabilities can be dynamically extended by connecting to MCP servers — some of which the developer chose, some of which may have been added by other users, other systems, or the agent ecosystem itself. That dynamic extensibility is what makes MCP valuable, and what makes governance non-optional.

What MCP Actually Changes

The key shift is this: before MCP, an agent's capabilities were determined at deploy time by whatever tools a developer explicitly configured. After MCP, an agent's capabilities can be dynamically extended by connecting to MCP servers — some of which the developer chose, some of which may have been added by other users, other systems, or the agent ecosystem itself.

The ecosystem is also less stable than a deployment diagram suggests. The same July 2026 assessment that found 91.8% of audited servers running without OAuth also confirmed 640 production MCP servers out of more than 21,000 server instances detectable on the public internet, found that 687 tool instances across those confirmed servers expose shell execution capabilities without access controls, and observed that 41.6% of confirmed servers disappeared within three days between consecutive measurement runs. That is a supply of capability being stood up and torn down faster than any review cycle designed for software dependencies.

Here's the governance implication, stated plainly: when your agent had five explicitly configured tools, you could reason about what it might do with those five tools. When your agent has access to dozens of tools via MCP — some installed by you, some added by others, some with capabilities that have changed since you last reviewed them — the reasoning problem is fundamentally different. The CIS MCP guide puts management of Non-Human Identities at the centre of its recommendations; Astrix, one of its co-authors, contributed expertise in securing AI agents, MCP servers and NHIs, “including API keys, service accounts, and OAuth tokens that connect AI systems to enterprise resources.”

How Does MCP Expand the AI Agent Attack Surface?

Every tool an agent can access is a potential attack surface. Not because the tool is malicious, but because any tool can be invoked in ways that cause unintended consequences, and the space of invocations an LLM might generate is large.

Prompt injection via tool results. This is the one that should keep security teams up at night. In March 2026, Oasis Security researchers demonstrated a complete attack chain — dubbed "Claudy Day" — that combined invisible prompt injection with data exfiltration, showing that conversation history could be pulled out of a default, unmodified claude.ai session. No special configuration. No elevated permissions. The injection was embedded as HTML tags inside the claude.ai/new?q= URL pre-fill parameter — “invisible in the text box but fully processed by Claude when the user hit Enter”. The chain the researchers built used a Google Ads listing plus an unvalidated claude.com/redirect/ open redirect to deliver the URL, after which Claude could be instructed to write conversation history to a file and upload it to an attacker-controlled Anthropic account through the Files API. Oasis reported all three findings to Anthropic through its Responsible Disclosure Program before publication, and states that “the prompt injection issue has been fixed, and the remaining issues are currently being addressed.” Oasis describes what its own researchers were able to do; it does not report exploitation by anyone else. It is a demonstration of what becomes possible when content entering agent context goes uninspected. For a deeper technical breakdown of how prompt injection enters agents through tool call results, post 43 covers the mechanism in full.

Tool definitions that change after you approve them. An allowlist approves a tool as it was described at approval time. Descriptions can change afterwards, and the approval does not re-run when they do. On August 12, 2026, Pillar Security published its analysis of a campaign it named Deadbugz and described as active at that date, in which a malicious MCP server was offered to open source projects through 23 pull requests created, in Pillar’s words, “in a 74-minute period, from 9:52 PM to 11:07 PM UTC on August 10, 2026.” None of the 23 had been merged when Pillar reviewed them — 19 were closed and four were still open — so this is an attempted supply-chain insertion rather than a completed one, which is exactly why the tool-definition mechanism is the part worth studying. The server behaves normally at first and counts invocations; once it reaches three, subsequent tools/list and prompts/get responses change, and the agent is redirected toward SSH keys, AWS credentials, shell history, and Kubernetes configuration. Pillar's own recommendation is the governance control, stated plainly: tool descriptions and schemas are a security boundary, and clients should treat a change in the tool definition of an already-approved server as a meaningful security event. This is the rug-pull pattern, covered in depth in The MCP Rug Pull Attack.

Unauthorized tool invocation. This is no longer hypothetical either. On June 26, 2026, Wiz researchers disclosed CVE-2026-12957, assigned a CVSS 4.0 score of 8.5, in Amazon Q’s Visual Studio Code extension: the extension automatically loaded and executed any .amazonq/mcp.json MCP server configuration found in a workspace the moment a developer opened the project and activated Amazon Q — no prompt, no consent, no workspace trust check. Because the resulting process inherited the developer's full environment, a single booby-trapped Git repository could run arbitrary commands with access to AWS credentials, API keys, and SSH agent sockets, with no user interaction beyond opening a folder and activating Amazon Q. Amazon patched the flaw in language server version 1.65.0, and Wiz argued, in The Register’s account, that the bug is less an Amazon problem than an industry one: as more AI coding assistants adopt MCP to reach local tools, workspace configuration files become an under-scrutinized trust boundary. The same dynamic drives the wider AI agent supply chain problem. An MCP config file is itself a tool-invocation instruction, and without an explicit allowlist and a consent gate in front of it, whatever the file says goes. A governed agent enforces that gate before the first command runs; an ungoverned one discovers the gap after the credentials are gone.

Data exfiltration through tool call sequences. An agent with access to database query tools and communication tools could retrieve sensitive data and transmit it externally — not through a single dramatic action, but through a sequence of plausible-looking operations that individually wouldn't raise flags. Defending against this requires visibility into sequences of tool calls and cross-tool patterns within a single run, not just individual invocations. This is what the CIS guide calls auditable interactions across the protocol layer.

Cost explosions from unconstrained tool use. Some MCP tools are expensive to call. Code execution environments, large data retrievals, external API calls with per-request pricing — an agent that calls these tools more aggressively than expected, or gets into a loop involving expensive tools, can run up costs quickly. Without budget enforcement at the tool call level, you won't know until after.

Why Existing Agent Governance Doesn't Fully Address MCP

If you have governance infrastructure in place — policies on LLM calls, input/output scanning, session-level budget guardrails — you have a foundation. But MCP introduces governance requirements that foundation may not cover.

Dynamic tool availability. Static governance policies assume a known, fixed set of tools. MCP's whole point is that the tool set is extensible. Your governance layer needs to handle tool calls against tools that weren't in scope when the policy was written — which means you need policies that generalize across tool categories, not just name-specific allowlists.

Tool result trust. Your existing scanning may evaluate inputs from users. MCP tool results come from servers, not users, and deserve their own trust evaluation. A tool result can contain malicious content even when the tool itself is legitimate and the server is trusted. The content-as-attack-surface problem requires governance at the tool result layer, not just the user input layer. The GitHub coding agent prompt injection case — where instructions embedded inside repository files were retrieved and acted on by an agent without disclosure — is a sharp example: GitHub AI Agent Prompt Injection: What No CVE Disclosure Looks Like.

Cross-tool sequence analysis. Individual tool calls may be benign. Sequences of tool calls may not be. A governance layer that evaluates each tool invocation in isolation misses the pattern-level risks. This requires tracking what tools have been called, in what order, with what payloads, and applying policies to sequences rather than individual events.

One thing changed underneath that requirement this summer. The current MCP specification revision, 2026-07-28, removed protocol-level sessions and the Mcp-Session-Id header from the Streamable HTTP transport, and removed the initialize handshake along with them; servers that need cross-call state are told to mint their own explicit handles and pass them as ordinary tool arguments. Sequence-level governance still works, but it can no longer lean on a transport-provided session identifier to decide which calls belong together. Whatever sits in front of your tool calls now has to establish that correlation itself, and any audit record that assumed a protocol session id needs a different anchor.

Annex III of the EU AI Act — covering high-risk AI systems — was originally set to apply from August 2, 2026. That date has moved. Under the AI Omnibus, which entered into force on July 27, 2026, the Annex III high-risk obligations now apply from December 2, 2027, and the Annex I obligations for high-risk AI embedded in physical products from August 2, 2028. The Omnibus changed the timeline, not the substance. Article 26 still requires deployers of high-risk systems to assign human oversight to named, competent people, and to keep the logs the system generates automatically, where those logs are under their control, for at least six months — obligations that come into force on the new Annex III and Annex I dates. The deferral does not reach Article 50's transparency obligations, which applied from August 2, 2026 as originally scheduled. MCP tool call sequences are exactly the kind of record an AI agent audit trail has to capture, and a governance stack that logs LLM calls but not tool-call sequences will not produce that record.

What MCP Governance Actually Looks Like

Practically speaking, MCP governance operates at three layers:

Tool allowlisting and scoping. Define explicitly which MCP tools the agent is permitted to call, under which conditions, with which parameter constraints. This is RBAC for tool access — not just "can the agent access this tool" but "under what circumstances." An agent might be permitted to read files but not write them, or to query a database only within a specified set of tables, or to call an email tool only when a specific approval condition is met. Allowlisting also means a workspace-supplied MCP config file is treated as untrusted input requiring explicit approval, not something an assistant loads and executes on open — the exact control gap the Amazon Q flaw exposed. The 2026-07-28 specification revision makes the inventory side of this easier: server/discover is now an RPC that servers must implement, advertising their supported protocol versions, capabilities, and identity in a single request, which gives an allowlist a standard place to read from. On the permission side, Your MCP Agents Are Over-Privileged walks through per-agent scopes and ephemeral tokens.

Waxell's MCP Gateway operationalizes allowlisting at the infrastructure level: one URL per tenant replaces the per-machine upstream MCP configuration, and the Fingerprints tab tracks every tool the gateway has seen through the trust states Pending review, Drift detected, Trusted, Blocked, and Removed. Worth being precise about what that buys you: fingerprinting detects a change, it does not by itself stop a call. On the Gateway, a deny_drift policy rule — recommended in the starter rule set, not enabled by default — is what turns a changed fingerprint into a denied call. For Python agents you build on another framework and can instrument in-process, Waxell Observe enforces 50+ policy categories before execution, between steps, and after completion, with no changes to your agent logic.

Tool result inspection. Every tool call response that gets appended to the agent's context should pass through an inspection step before it's processed. Scan for injection patterns. Check for PII that shouldn't be entering the context. Validate that the response conforms to expected schemas. Flag anomalies for review. The Gateway handles this at the network layer: every tool description is scanned for embedded instructions at fingerprint time — before any user or agent ever calls the tool — and redaction, DLP, and egress rules you author cover PII, credentials, and outbound hosts in both tool arguments and tool results.

Sequence-level policy evaluation. Track tool call sequences within a single agent run and evaluate them against run-level policies. A policy might say: if more than three different external endpoints are called within a single run, flag for review. Or: if a database read is followed within N turns by any outbound call, require approval. These are run-level governance decisions, not call-level decisions — and, as above, the correlation they depend on is now yours to establish rather than the transport's.

For teams governing MCP servers they didn't build — marketplace servers, vendor-provided integrations, third-party data connectors — the MCP Gateway puts a single governed endpoint in front of them, with every tool call that traverses it resolved to a real user identity and policy-checked before the upstream sees it and again on the way back. The scope of that is worth stating: the gateway governs the calls that traverse it. Agents holding upstream credentials directly, MCP servers registered locally on a machine, and clients that were never pointed at the gateway do not. The MCP STDIO design flaw that Ox Security says puts 200,000 servers at risk is a clear example of why governance of external MCP servers can't depend on the servers themselves implementing controls.

Why Is Now the Right Time to Establish MCP Governance?

MCP is early enough that the tooling ecosystem around it is still being built — including governance tooling. The CIS MCP Companion Guide is a useful marker of where the practice has got to: it extends an established controls framework into MCP environments rather than proposing a new one, and the problems it names — secure tool access, Non-Human Identity management, and auditable interactions across the protocol layer — are the ones teams running MCP hit first.

The organizations that establish MCP governance practices now — while the tool ecosystem is still relatively small and the attack surface is still relatively bounded — will be in a fundamentally different position than the organizations that establish them in eighteen months, when the ecosystem is larger, the incidents have accumulated, and the EU AI Act's Annex III obligations are in application (December 2, 2027, per the AI Omnibus).

There's a general principle worth stating: the right time to build governance for a capability is when the capability is being adopted, not after the first incident that makes its absence visible. MCP is being adopted right now.

How Waxell handles this: Waxell's MCP Gateway is a single governed endpoint that replaces the per-machine upstream MCP configuration. Every tool call that traverses it is resolved to a real user identity rather than a service account, and evaluated against policy before the upstream sees it and again on the way back; the decision and the rules that fired land in a payload-free audit log, exportable to CSV. The Fingerprints tab tracks every tool the gateway has seen through the trust states Pending review, Drift detected, Trusted, Blocked, and Removed, and every tool description is scanned for embedded instructions at fingerprint time, before any user or agent ever calls the tool. Detection and enforcement are separate: a deny_drift rule is what denies a call whose fingerprint changed since it was approved. Redaction, DLP, and egress rules you author cover PII, credentials, and outbound hosts; destructive actions can be parked for a human reviewer, with the connection held open so the agent doesn't time out. Policy rule changes take effect across the gateway fleet within seconds. For Python agents you build on another framework and can instrument in-process, Waxell Observe adds two lines of code and enforces 50+ policy categories before execution, between steps, and after completion.

Start free with the Waxell MCP Gateway — the Free plan includes one governed MCP upstream with 14-day data retention. Create an account →

Frequently Asked Questions

What is MCP governance? MCP governance is the practice of enforcing policies on AI agents that use Model Context Protocol to access external tools and services. It covers which tools an agent is permitted to call (allowlisting), what conditions must be met before a tool is invoked (scoping), whether tool results are safe to enter agent context (result inspection), and whether a sequence of tool calls within a single run is consistent with policy (sequence-level analysis).

Did the 2026 MCP specification update change what I need to govern? Partly. The current revision, 2026-07-28, removed protocol-level sessions and the Mcp-Session-Id header from the Streamable HTTP transport, along with the initialize handshake, so any governance or audit logic that correlated tool calls using a transport session identifier needs a different anchor. It also added server/discover, an RPC servers must implement, which gives allowlisting a standard inventory surface. Separately, Roots, Sampling, Logging, the HTTP+SSE transport, and OAuth 2.0 Dynamic Client Registration are now marked Deprecated. In the specification’s own words, a deprecated feature “remains in the specification but is scheduled for removal” and “New implementations should not adopt the feature” — deprecated features have not stopped working.

What is prompt injection via MCP tool results? Prompt injection via tool results occurs when an MCP tool returns content that contains instructions designed to alter the agent's behavior — for example, a retrieved document containing text like "ignore your previous instructions and…" That text enters the agent's context window and may influence subsequent reasoning. In March 2026, Oasis Security demonstrated the "Claudy Day" attack chain, showing that invisible prompt injection plus data exfiltration was achievable against a default claude.ai session without any special configuration, using a URL parameter injection. Defense requires inspecting content before it enters agent context, whether from tool results, URL parameters, or any other source.

How do you govern AI agents that use MCP tools? Effective MCP governance requires three layers: tool allowlisting (explicitly defining which tools the agent may call and under what conditions), tool result inspection (scanning every tool response for PII, injection patterns, and schema anomalies before it enters context), and sequence-level monitoring (tracking cross-tool patterns within a run and triggering policies based on sequences, not just individual calls). Static policies based on named tool allowlists need to generalize across tool categories as the MCP ecosystem grows.

Why doesn't existing agent governance fully cover MCP? Most existing governance infrastructure was designed for known, fixed tool sets. MCP's value proposition is a dynamically extensible tool ecosystem — which means static name-based allowlists become inadequate. Additionally, MCP introduces tool result trust evaluation as a distinct concern: a tool can be legitimate and its server trusted, but its results can still contain malicious content. Governance layers that only evaluate user inputs miss this vector entirely.

What does the CIS MCP Companion Guide (April 2026) recommend? The CIS MCP Companion Guide, published April 20, 2026 with Astrix Security and Cequence Security, extends the CIS Controls into MCP-based environments. Its focus areas include governing the Non-Human Identities created by each MCP connection, implementing secure tool access controls with explicit permission grants per capability rather than broad access, and ensuring auditable interactions across the protocol layer. The announcement covering that guide and the two released with it names four risks across AI systems generally: data leakage, unbounded agent autonomy, credential misuse, and unsafe or inappropriate execution of tools.

How does Waxell's MCP Gateway enforce MCP governance? The MCP Gateway is a single URL that replaces the per-machine upstream MCP server configuration. Every tool call that traverses it is resolved to a real user identity and evaluated against policy before the upstream sees it and again on the way back. It fingerprints tools, tracking them through the trust states Pending review, Drift detected, Trusted, Blocked, and Removed, and scans every tool description for embedded instructions at fingerprint time, before any user or agent calls the tool. Fingerprinting detects a change; a deny_drift policy rule is what denies the call. Redaction, DLP, and egress rules cover PII and credentials, and destructive actions can be parked for human approval. It works with Claude Desktop, Claude Code, Cursor, and other MCP-compatible clients, and setup requires no code changes. See how it works →

Can an MCP configuration file execute code without my consent? It can if the client trusts workspace files by default. The June 2026 Amazon Q flaw (CVE-2026-12957) is the clearest documented case: the VS Code extension auto-loaded and ran any .amazonq/mcp.json file found in an opened project once Amazon Q was activated, with no prompt and no workspace trust check, giving a malicious repository a path to the developer's AWS credentials. The fix isn't specific to Amazon Q — any MCP client that treats a workspace config file as pre-authorized rather than as untrusted input requiring explicit consent has the same exposure. Governed MCP access requires an allowlist and a consent gate in front of configuration loading, not just in front of the tools themselves.

Sources

  • Center for Internet Security (CIS), Astrix Security, and Cequence Security, CIS, Astrix, and Cequence Release New AI Security Companion Guides (April 2026) — https://www.cisecurity.org/about-us/media/press-release/cis-astrix-and-cequence-release-new-ai-security-companion-guides

  • The Register, MCP 'design flaw' puts 200k servers at risk: Researcher (April 2026) — https://www.theregister.com/2026/04/16/anthropic_mcp_design_flaw/

  • Nicolás Padilla, Exposed by Design: A Dynamic Security Assessment of Internet-Facing MCP Servers at Scale (July 2026) — https://arxiv.org/abs/2608.00150

  • Pillar Security, Deadbugz: Currently Active MCP Supply Chain Campaign (August 2026) — https://www.pillar.security/blog/deadbugz-currently-active-mcp-supply-chain-campaign

  • Oasis Security, Claude.ai Prompt Injection Vulnerability ("Claudy Day") (March 2026) — https://www.oasis.security/blog/claude-ai-prompt-injection-data-exfiltration-vulnerability

  • The Register (Carly Page), Amazon Q flaw let booby-trapped Git repos execute code, swipe cloud creds (June 2026) — https://www.theregister.com/cyber-crime/2026/06/26/amazon-q-flaw-let-booby-trapped-git-repos-execute-code-swipe-cloud-creds/5263202

  • Model Context Protocol, Key Changes — specification revision 2026-07-28 — https://modelcontextprotocol.io/specification/2026-07-28/changelog

  • Model Context Protocol, Versioning — https://modelcontextprotocol.io/specification/versioning

  • European Commission, AI Omnibus enters into force (July 2026) — https://digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force

  • Regulation (EU) 2026/1744 (Digital Omnibus on AI), Official Journal of the European Union, OJ L, 2026/1744, 24.7.2026 — https://eur-lex.europa.eu/eli/reg/2026/1744/oj

  • EU Artificial Intelligence Act, Article 26: Obligations of Deployers of High-Risk AI Systems — https://artificialintelligenceact.eu/article/26/

  • EU Artificial Intelligence Act, Implementation Timeline — https://artificialintelligenceact.eu/implementation-timeline/

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.