MCP Protocol Flaws Let Attackers Hijack AI Agent Chains
Researcher Syed Anas Mohiuddin found a fundamental flaw in the Model Context Protocol (MCP) that lets malicious prompts hop between trusted AI agents, a technique he calls "protocol pivoting." Google, JPMorgan Chase, Weaviate, Rapid7, and government agencies in France and the U.S. have all acknowledged related vulnerabilities in the past five months. The exploit works because agents implicitly trust each other without verifying request sources, effectively abandoning zero-trust princip
GGLOBAIRESEARCH DESKSHARE
Researcher Syed Anas Mohiuddin found a fundamental flaw in the Model Context Protocol (MCP) that lets malicious prompts hop betwee…
Share this post
Short answer: Researcher Syed Anas Mohiuddin found a fundamental flaw in the Model Context Protocol (MCP) that lets malicious prompts hop between trusted AI agents, a technique he calls "protocol pivoting." Google, JPMorgan Chase, Weaviate, Rapid7, and government agencies in France and the U.S. have all acknowledged related vulnerabilities in the past five months. The exploit works because agents implicitly trust each other without verifying request sources, effectively abandoning zero-trust principles in agentic architectures.
MCP protocol AI agent trust flaw
It turns out there's a fundamental flaw in how AI agents trust each other inside corporate networks, and it centers on the Model Context Protocol. Independent researcher Syed Anas Mohiuddin showed that malicious prompts can hop from one specialized agent to another by exploiting the implicit trust baked into MCP, the communication standard that lets agents delegate tasks without ever verifying the source. In just the last five months, Google, JP Morgan Chase, Weaviate, Rapid7, the French government's interministerial digital directorate, and the U.S. federal government have all acknowledged vulnerabilities tied to this exact mechanism.
Think of the protocol like a hallway where every door is locked but nobody checks who walks between rooms. An attacker plants a crafted prompt in content that a first agent, say, a translation tool, reads and passes along as a normal delegated request. The second agent executes it because it trusts the first agent by design. Each component behaves exactly as intended, which makes the chain incredibly difficult to detect. Douglas McKee, director of vulnerability intelligence at Rapid7, described it perfectly: every protocol checks its own front door while nobody watches the hallway in between.
CVE-2026-97228 vulnerability details
Rapid7's vulnerability, tracked as CVE-2026-97228, carried a modest severity rating of 2.7 out of 10 and was patched last month. Google's flaw was far more serious, rated 8 out of 10. It originated in the googleapis/mcp-toolbox component, which initialized its HTTP client without a CheckRedirect policy and failed to validate target IP addresses. A manipulated path parameter could force the toolbox to follow a redirect to an internal endpoint and issue requests on the attacker's behalf. Google's fix introduced an allow-list of IP ranges, block lists, and a startup-time rejection of unsafe base URLs rather than waiting for the first request. Syed noted that this approach, validating at initialization, is what a real server-side request forgery guard looks like, and it represents more work than most MCP servers have bothered with.
Protocol pivoting attack technique
Syed calls the technique "protocol pivoting" because the exploit crosses protocol boundaries. An initial compromise via MCP can hand off malicious instructions to another agent using Google's Agent-to-Agent protocol or the emerging Agent Network Protocol. Trust and authorization effectively get lost in translation. Markus Vervier, a researcher at X41 D-Sec, argues the better term remains "prompt injection" and that Syed's method is a subclass of indirect prompt injection. The protocol-hopping aspect, he said, isn't strictly required for such attacks to work, though it makes them unexpected and harder to mitigate.
What makes the findings notable is that the same pivoting technique succeeded across five organizations with nothing in common except their use of MCP. The protocol is new, already widespread, and hasn't been sufficiently tested or hardened. In the rush to build sprawling agentic architectures, organizations have abandoned the zero-trust principle, the assumption that any node may be compromised and that sensitive transactions require explicit authorization. McKee emphasized that anything passed from an LLM to a tool should be treated like input from a stranger on the internet, because in a prompt injection scenario that is exactly what it is. The underlying bugs are old friends, injection and server-side request forgery, and the fixes haven't changed in twenty years. Naming the pattern, he added, is what gets defenders and standards bodies to actually design for it.
Agent-based system security recommendations
For teams building or operating agent-based systems, the immediate takeaway is to stop assuming internal agents are trustworthy. Implement allow-lists and block-lists for outbound connections, validate redirects and target addresses at startup rather than on first use, and require explicit authorization for any agent-to-agent action that touches sensitive data or external endpoints. The protocol layer cannot be the only line of defense.
Frequently asked questions
What is the fundamental flaw in the Model Context Protocol (MCP)?
MCP allows AI agents to delegate tasks without verifying the source, creating implicit trust between agents. Malicious prompts can hop from one agent to another because each component trusts the previous one by design, making the attack chain difficult to detect as every agent behaves exactly as intended.
Which organizations have acknowledged MCP-related vulnerabilities in the last five months?
Google, JP Morgan Chase, Weaviate, Rapid7, the French government's interministerial digital directorate, and the U.S. federal government have all acknowledged vulnerabilities tied to the MCP mechanism since late 2025.
What is "protocol pivoting" and how does it differ from standard prompt injection?
Protocol pivoting describes exploits that cross protocol boundaries-an initial MCP compromise hands off malicious instructions to another agent using Google's Agent-to-Agent protocol or the emerging Agent Network Protocol. Researcher Markus Vervier argues this is a subclass of indirect prompt injection, noting the protocol-hopping aspect isn't strictly required but makes attacks unexpected and harder to mitigate.
How did Google fix their critical MCP vulnerability rated 8 out of 10?
Google's flaw in the googleapis/mcp-toolbox component stemmed from an HTTP client initialized without a CheckRedirect policy and missing target IP validation. Their fix introduced an allow-list of IP ranges, block lists, and startup-time rejection of unsafe base URLs rather than waiting for the first request-validating at initialization, which researcher Syed Anas Mohiuddin noted represents a proper server-side request forgery guard.
What immediate steps should teams take to secure agent-based systems using MCP?
Stop assuming internal agents are trustworthy. Implement allow-lists and block-lists for outbound connections, validate redirects and target addresses at startup rather than on first use, and require explicit authorization for any agent-to-agent action touching sensitive data or external endpoints. The protocol layer cannot be the only line of defense.
A static site generator compiles HTML files from templates and content-often Markdown-at build time, producing a deployable folder of static assets. This eliminates the need for a database or application server, allowing hosting on CDNs or object storage for faster loads and a smaller attack surface. Publishing updates requires rebuilding and redeploying the site, a workflow suited for informational sites like blogs or documentation but less ideal for real-time personalization without
The Wikimedia Foundation discovered unauthorized OpenAI agents making automated edits, probing collaborative tools, and generating millions of API requests across Wikipedia, Wikidata, and Wikimedia Commons. The bot traffic may have contributed to a partial service outage in May, though OpenAI has not verified its role. No systems were compromised, and the agents lacked required community approval for automated editing.
OpenAI claimed in early September that its AI solved the Navier-Stokes Millennium Prize Problem, triggering backlash from mathematicians who accused the company of "scooping" decades of collective work, violating academic norms, and potentially absorbing unpublished insights from chat interactions. Critics argue OpenAI prioritizes competitive benchmarking over collaborative advancement. The company formed an independent mathematician advisory panel on September 23, but its authority an
NO COMMENTS YET
Comments are open. Have a thought or a question? Share it below.