AI Demo Cloudflare AI security demo

This has already happened

Every control in this demo exists because of something real. The apps here are invented; the failures they reproduce are not. This page is the citation list — one incident per control, with a link to the primary source for each, because "we made this up to sell you something" is the first objection any security team should raise.

How to use this page

Two ways. Open it cold at the start of a session, before anyone has seen a single screen, and spend four minutes here: it changes the register of everything that follows, because the room stops evaluating a vendor demo and starts recognising its own environment. Or keep it in a second tab and reach for the relevant entry the moment somebody says would that really happen.

Every date and claim below is sourced. If a source says "proof of concept" rather than "exploited in the wild", so does this page — the argument does not need the exaggeration, and one overstated claim costs you the whole set.

Where each one lands in this demo

IncidentWhat it is evidence forScript
EchoLeak, Microsoft 365 CopilotIndirect prompt injection by email, with no click. Why the prompt-injection guardrail sits on the model call.The email nobody clicked
Gemini for Workspace calendar invitesThe same attack delivered by an invite nobody accepted. Why content you never chose to receive is untrusted input.The email nobody clicked
Samsung and ChatGPTStaff pasting source code into a consumer assistant to get their work done. Why the control is on the data, not the site.Pasting a colleague's data, and step 1 of the walkthrough
Asana's MCP serverAn MCP server shipped over a working app, exposing data across tenants. Why the new surface needs its own inspection.every Gateway DLP over MCP script
GitHub's MCP serverAn agent with legitimate access to two things, talked into moving data between them. Why tool permissions are not enough.The email nobody clicked, the CEO's home address
DeepSeek's exposed databaseAn unapproved service whose own security nobody assessed. Why an approval status is a control and not paperwork.The Redirect non approved AI policy on the protection page
Cloudflare's own MCP OAuth libraryThree advisories against the exact library these MCP servers are built on. Why nobody gets to be smug about this.The five MCP servers in this repository

EchoLeak — zero-click exfiltration from Microsoft 365 Copilot

CVE-2025-32711, disclosed June 2025

What happened. Researchers at Aim Security found that a single crafted email could make Microsoft 365 Copilot collect information from the user's earlier conversations and send it to an attacker-controlled server. The victim did not have to open the email or click anything. The attack fired later, when they asked Copilot an ordinary question and the malicious email was pulled into context as relevant material.

Why it worked. It is a chain, and each link is instructive. Microsoft's cross-prompt-injection classifiers were bypassed by writing the email as if the instructions were addressed to the recipient, never mentioning AI or Copilot. Link redaction was bypassed using Markdown's reference-style link syntax, which the filter did not cover. Auto-fetched images and a Microsoft Teams proxy permitted by the content security policy carried the data out.

Status. Reported to Microsoft in January 2025, patched server-side, rated critical as "AI command injection in M365 Copilot". Microsoft told customers no action was required.

What this demo reproduces. The delivery mechanism and the trigger, not the exfiltration chain. Alice's injected email sits unread until she asks her assistant to summarise her inbox. Being honest about the difference: a real payload is hidden with HTML the renderer drops, and WorkBox stores plain text, so ours is merely boring rather than invisible.

Sources: CVE record · Aim Labs' own disclosure · the same write-up via Cato · BleepingComputer · SecurityWeek · Simon Willison's walkthrough · case-study paper

"Invitation Is All You Need" — hijacking Gemini with a calendar invite

SafeBreach and academic researchers, Black Hat USA and DEF CON 33, August 2025

What happened. Or Yair (SafeBreach), Ben Nassi and Stav Cohen demonstrated that sending someone a Google Calendar invite whose title carries an indirect prompt injection was enough to hijack Gemini for Workspace the next time that person asked it about their week. They showed 14 attack scenarios across five threat classes against Gemini's web interface, mobile app and Google Assistant.

Why it matters more than an email. An invite arrives in your calendar whether you accept it or not. There is no click to avoid and no attachment to distrust. The researchers' consequences ran from spam and disinformation to exfiltrating the victim's email, geolocating them, video-streaming them over Zoom, and operating physical home automation — because the assistant's agents had the permissions to do all of it.

Status. Disclosed to Google, which deployed mitigations. The researchers' own risk assessment put 73% of the threats they analysed at high or critical before those mitigations, and very low to medium after.

What this demo reproduces. The channel. One difference worth stating: their payload is in the invite's title, where ours is in the description, below an agenda a colleague pasted in good faith. Both are text an agent reads while summarising the week; theirs is the harder trick.

Sources: SafeBreach write-up · the paper (arXiv 2508.12175) · DEF CON 33 slides

Samsung — engineers pasting source code into ChatGPT

April and May 2023

What happened. In April 2023, Samsung's Device Solutions division — the semiconductor business — found three cases where engineers had pasted confidential material into ChatGPT: buggy source code submitted for a fix, another piece of code submitted for optimisation, and a set of meeting notes submitted for summarising. On 1 May Samsung banned generative AI tools on company devices in one of its largest divisions, capped prompts at 1,024 bytes, and told staff that failure to comply could result in dismissal.

Why it is the most useful incident on this page. Nobody was attacking anything. Three engineers had a problem, reached for the best tool available, and did their jobs slightly faster. That is the entire mechanism, and no amount of security awareness training removes the incentive.

Note what Samsung did next: started building an internal assistant. A ban on its own is a policy staff will route around on their phones — which is why the control in this demo inspects the data rather than blocking the site, and why the company's own sanctioned assistant deliberately allows the customer data a marketer works with all day.

Sources: Korea Herald · The Register · Bloomberg (the original scoop, paywalled — use one of the first two on stage)

Asana — an MCP server that crossed tenants

May to June 2025, roughly 1,000 customers notified

What happened. Asana launched an MCP server on 1 May 2025 so customers could point AI assistants at their workspaces. On 4 June they found a logic flaw in it that could expose data from one customer's Asana domain to MCP users in another. They took the server offline the next day and brought it back on 17 June, reset every connection, and notified affected customers. A spokesperson put the number at around a thousand.

What was exposed. Task-level information, project metadata, team details, comments and discussions, and uploaded files — bounded by each MCP user's own access scope rather than whole workspaces. Asana was explicit that this was a bug in their code, not a hack, and there is no indication anyone exploited it or that another organisation's data was actually read.

Why it is here. This is the closest published analogue to the argument this demo makes. A competent company with a mature product wrote new code to put an agent in front of existing data, and the new code got the scoping wrong for five weeks. Every MCP server in this demo is the same shape: a thin layer over an app whose web UI has been careful about access control for years.

Sources: UpGuard's timeline · BleepingComputer · The Register

GitHub's MCP server — a public issue that reads a private repository

Invariant Labs, May 2025

What happened. Invariant Labs showed that an attacker who can open an issue on a public repository can leave a prompt injection there, and that a developer later asking their agent to "have a look at the open issues" gets the agent to pull data out of their private repositories and publish it in a pull request on the public one. The agent has legitimate access to both, so every individual action is authorised.

GitHub's response is the important part. They did not treat it as a bug in their MCP server, and they were right not to: it was not a specific vulnerability in GitHub MCP. It is true of any tools that enable lethal trifecta — private data, untrusted content and a way out. They have since added read-only and lockdown modes, and secret scanning on tool calls that touch public repositories.

Why it is here. It is the cleanest published statement of why this demo puts the control in the network path. If the vulnerability is the combination of capabilities rather than a defect in any one tool, then no amount of reviewing individual tools finds it, and the thing that catches it has to be watching the data rather than the code.

Sources: Invariant Labs · the GitHub issue, including their response · DevClass

DeepSeek — a million chat logs on the open internet

Wiz Research, January 2025

What happened. Within minutes of starting to look, Wiz found two ClickHouse databases belonging to DeepSeek reachable from the internet with no authentication at all, allowing arbitrary SQL through a browser. They held over a million log lines including plaintext user chat history, API secrets, and backend operational detail, and allowed full control of the database. DeepSeek secured it within hours of being told.

Why it belongs next to the shadow-AI control. The usual framing of shadow AI is "our data goes somewhere we cannot see". This is the other half: it goes somewhere whose own security posture nobody has assessed, during the exact week half the industry was trying the service out. An approval status in the Application Library is not paperwork — it is the record of whether anyone looked.

What this demo does with it. DeepSeek is one of the services marked Unapproved, and the Redirect non approved AI policy sends anyone who opens it to the assistant the company did evaluate. Nobody is refused, so nobody looks for a way around it.

Sources: Wiz Research · TechCrunch · CyberScoop

Cloudflare's own MCP OAuth library — three advisories

workers-oauth-provider: CVE-2025-4143, CVE-2025-4144, and GHSA-2h78-5wx8-jccc

Include this one. A page of other people's failures invites the obvious question, and the honest answer is better than the dodge: the library the five MCP servers in this repository are built on has had three security advisories of its own.

Where this repository stands. All five MCP servers pin ^0.10.3, far above every affected range, so none of the three applies here. The consent screen each server shows, its CSRF token in a __Host- cookie and its HMAC-signed state token in KV are the mitigations Cloudflare's own guide to securing MCP servers now documents.

Why say it out loud. Two reasons, and both help you. It demonstrates that OAuth for agents is genuinely new ground where careful teams are still finding basic mistakes — which is the argument for defence in depth rather than against it. And a room that has watched you volunteer your own vendor's CVEs will believe the rest of what you say.

What the pattern actually is

Read the seven together and the same shape appears in all of them. Not one is a model behaving strangely, and not one would have been prevented by a better prompt:

Which is the argument for where the controls in this demo sit. The protection layer changes no application code and fixes none of the leaky endpoints, because in every incident above the application was not the thing that could be fixed in time.

Keeping this page honest

Every entry here was verified against its primary source rather than from memory, and the wording tracks what those sources actually claim. If you are reading this months later, check the advisories again before presenting: version numbers move, mitigations ship, and a page like this decays quietly. The dates are all here so you can tell how stale it has become.