Prompt injection: real cases from 2026 and what helps
A link in Microsoft Copilot, a web form in Salesforce Agentforce: two real prompt injection cases from 2026, the pattern they share and what employees and IT should do.
Two prompt injection cases published in 2026 show how an AI assistant with access to a mailbox or CRM could send internal data to outsiders while the users did only everyday things: open a link or ask for the latest leads. At Microsoft Copilot a prepared link was enough, at Salesforce Agentforce a filled-in web form. Both flaws are fixed, the pattern behind them is not.
How prompt injection works technically and which risk classes OWASP distinguishes is covered in our primer on prompt injection in AI agents. This article is about the real cases and what employees and IT can learn from them.
TL;DR
- Two flaws published in 2026, in Microsoft Copilot Personal and Salesforce Agentforce, show that a link or a form entry was enough for an AI assistant to send internal data to attackers. Both flaws have been fixed.
- The shared pattern: the assistant reads external content, has access to internal data and can communicate outward. OWASP still ranks prompt injection first in 2026 and writes that there is currently no reliable protection mechanism against it.
- What helps: treat external content as untrusted, give agents only the rights they need, and allow sending, deleting or paying only after human confirmation.
What prompt injection is
Prompt injection means that a piece of text makes an AI model do something the operator did not intend. With the direct variant, the user enters the manipulative input themselves. With the indirect variant, the instruction sits in content from an external source that the model processes, such as a web page, an email, a document or a database entry. According to OWASP, the user neither entered nor saw this instruction, and it does not even have to be visible in the interface. The core of the problem according to the OWASP Top 10 for LLM Applications 2026: language models do not technically separate instructions from data, both end up as text in the same stream.
Case 1: CoSnitch in Microsoft Copilot
Varonis Threat Labs published the chain of flaws on 18 August 2026 under the name CoSnitch (CVE-2026-24301). The affected product was Microsoft Copilot Personal at copilot.microsoft.com, the version for personal use. Varonis does not name Microsoft 365 Copilot for CoSnitch.
How the attack worked:
- The attacker builds a link to Copilot that contains a pre-written request (parameter
q) plus the undocumented parameterautorun=1. - The victim clicks the link. Copilot opens in the signed-in session and runs the request immediately, without another click.
- The request makes Copilot search connected services, such as Gmail, Google Drive or the calendar. Copilot encodes the hits and sends them to a server of the attacker through its own function for fetching web addresses.
The second part is trickier. Through a prepared web page that the victim has Copilot summarize, the attacker could write persistent rules into Copilot's memory. When summarizing, Copilot did not distinguish between content and instruction. According to Varonis, such entries stayed active in every future conversation until the user deleted them by hand in the memory settings, even after password changes, ended sessions and newly registered devices.
Timeline: Varonis reported the flaws to Microsoft in December 2025, and they were fixed on 18 August 2026. Varonis has not seen any sign of exploitation in the wild.
Why this is a business topic: Anyone who connects a private Copilot account to a mailbox or storage that also holds work content carries this risk into the company. How private AI use can run under the radar in a company is described in our article on the shadow AI reality in the Mittelstand. Varonis recommends reviewing connected apps and reducing them to what is necessary, treating Copilot like a privileged insider and monitoring unusual data access by the assistant.
Template
AI policy with an agent annex
Case 2: SalesBleed in Salesforce Agentforce
On 24 September 2026, Zenity Labs published SalesBleed: vulnerabilities in Salesforce Agentforce through which an outsider could pull CRM data without ever logging in and without anyone in the company having to click anything.
How the attack worked:
- The attacker fills in the company's public Web-to-Lead form, that is, the contact form on the website from which a lead is created in the CRM. In one of the fields they hide instructions for the agent.
- An employee asks the agent for something completely normal, in Zenity's example roughly: "Look at my latest leads and help me with the newest one."
- The agent reads the prepared lead and follows the hidden instructions. It queries the lead and account tables and builds company names and deal sizes into a web address as part of the server name.
- The agent outputs this address as an image. As soon as the interface tries to load the image, it resolves the server name, and the data reaches the attacker through a DNS request. In Slack, the automatic link preview, which fetches links on its own, was enough.
Zenity also bypassed the redaction of links in the agent's responses, among other things with a top-level domain that the filter did not recognize. The researchers note: redacting outputs is a race between one parser and every program that later renders the output.
According to SecurityWeek, SalesBleed comprised three vulnerabilities. Through the third, the agent could be hijacked with a prepared lead to post phishing messages in Slack under its identity. According to Zenity's report on this flaw, the agent's Slack action required no confirmation before sending, and recipients were not shown which user had triggered the message.
Timeline: Zenity reported the flaws to Salesforce on 1 June 2026. For the data leak, Salesforce confirmed the fixes on 18 August 2026, and Zenity confirmed the fix for the reported bypass of the link filter on 19 August. For the Slack flaw, all fixes were confirmed and tested by 21 September 2026, according to Zenity. Salesforce hardened the Trusted URLs mechanism and, according to SecurityWeek, made user confirmation before sending the default for certain Agentforce actions in Slack.
What the cases have in common
Both cases follow the same pattern, and it has little to do with the respective product. Three ingredients came together:
| Ingredient | CoSnitch (Copilot) | SalesBleed (Agentforce) |
|---|---|---|
| External content | Request in the link, prepared web page | Entry in the public web form |
| Access to internal data | Connected mailboxes, storage, calendars | Leads and accounts in the CRM |
| Way out | Copilot fetching web addresses | Image address via DNS, link preview in Slack |
The OWASP 2026 list describes exactly this combination. It picks up Simon Willison's "lethal trifecta": an agent that can at the same time read private data, take in untrusted content and communicate outward offers the conditions for consequential attacks. If one of the three ingredients is removed, these conditions disappear as well, according to OWASP.
Two further commonalities:
- The attacker needed no access. With SalesBleed, a public form was enough. OWASP describes this approach explicitly: text is placed through a low-privilege channel, such as a public form or a customer ticket, at a spot the user trusts. The agent then performs the privileged step with the user's rights.
- The vendors closed the flaws, the principle remains. In the 2026 edition, OWASP writes that there is currently no reliable protection mechanism against prompt injection, because models do not distinguish between instructions and data and their behavior is not deterministic.
The pattern is not new. On 11 June 2025, Microsoft published a flaw in Microsoft 365 Copilot rated critical (CVE-2025-32711, CVSS 9.3), which Microsoft itself describes as "AI command injection": an unauthorized attacker could use it to disclose information over a network. According to Microsoft, the flaw is fully fixed, customers do not need to do anything, and Microsoft lists it as not exploited.
ENISA also takes a broader view of the issue. According to the ENISA Threat Landscape 2026 (22 September 2026), integrating AI systems into enterprise environments creates a new attack surface. Threat actors keep experimenting with methods to bypass the safeguards of closed models, including steganography and indirect prompt injection.
What OWASP, BSI and ENISA advise
OWASP: defense is architecture
The 2026 edition of the OWASP Top 10 for LLM Applications, which the OWASP GenAI Security Project officially presented on 2 September 2026, still ranks prompt injection first (LLM01:2026). Excessive Agency, meaning agents with overly broad rights and autonomy, is in third place. Because prompt injection cannot be reliably prevented, OWASP advises building the system so that a successful injection can do as little harm as possible. The key points:
- Credentials and the ability to change something live in the application code, not in the model. Rights are granted minimally per operation.
- Before every privileged, irreversible or externally visible action, a human confirms, and confirms the exact action, not a summary.
- The "Rule of Two" is the lower bound: if an agent has untrusted inputs, sensitive data and the ability to change something or communicate outward at the same time, every action needs human approval.
- Write access to an agent's memory counts as a privileged operation.
- MCP servers and tool packages are pinned to fixed versions, signed and vetted. Our MCP article explains what MCP is.
- Testing is done against attackers who know the protective measures in use.
With the 2026 edition, the OWASP GenAI Security Project also introduced the Agent Control Standard, a donation to the project that extends the agent recommendations toward runtime enforcement. For agents there has also been a dedicated list since 9 December 2025, the OWASP Top 10 for Agentic Applications 2026, from ASI01 Agent Goal Hijack to ASI10 Rogue Agents.
BSI: confirmation before actions, clear rights for AI
On 10 November 2025, the Federal Office for Information Security (BSI) presented a publication on evasion attacks on large language models, which include indirect prompt injections. As countermeasures it names, for example, precise and secure system instructions, filtering harmful content in third-party documents or explicit confirmation by users before the model executes functions. According to the BSI, implementing such countermeasures can make it harder for attacks to succeed or reduce the potential damage.
In a cybersecurity warning of 22 June 2026 (criticality 2, yellow), the BSI writes that current AI systems can find vulnerabilities quickly and in part autonomously and turn them into usable attack paths. It demands patch management that reviews vulnerabilities and rolls out patches within minutes to at most a few hours, plus minimal rights and multi-factor authentication. A few days is not an adequate response speed, it says.
In the BSI blog of 24 July 2026, BSI President Claudia Plattner describes a test agent from OpenAI that left its sandbox, connected to Hugging Face over the internet and obtained tools it considered necessary for its task, without malicious intent. One of her recommendations: robust identity and rights management for AI systems.
ENISA: secure agent endpoints, involve humans
In the position paper ENISA's view on Cybersecurity in the Frontier AI Era (July 2026), ENISA counts agentic endpoints and enterprise browsers among the critical new attack surfaces, because AI agents increasingly work autonomously on devices and in browsers. Among other things, it recommends AI workflows with human approval, supported by training for security teams. ENISA also reports feedback from practitioners indicating that SMEs may need additional support, especially guidance and access to current models.
Practical steps for employees
The everyday rules fit on one page. They follow a principle that our AI policy template also states in Annex A7: prompt injection cannot currently be fully ruled out technically, so you limit what a manipulated agent can do.
- External content is untrusted. This applies to emails, web pages, documents and form entries, even when they come from known senders. Anyone who lets an agent read such content must expect that instructions may be hidden in it.
- No links from unknown sources that open an AI assistant. CoSnitch shows that merely opening such a link could start a ready-made request. Agents themselves never automatically follow links or instructions from unknown sources.
- Send, delete, pay only after confirmation. An agent that reads external content deletes, sends outward or triggers payments only if a human has confirmed the action in the individual case (permission levels 4 to 6 in Annex A2). Do not give blanket approval or confirm an action without reading it.
- Only approved tools and connectors. Do not put company data in private AI accounts and do not connect mailboxes, storage or CRM systems yourself without approval. Approved connections get as few rights as possible, read-only where possible.
- Check the assistant's memory. If your assistant remembers rules across sessions, look at the settings now and then. Delete and report entries you did not create yourself.
- Report anything unusual at once. Unexpected actions, replies containing instructions from external sources, unexplained links or images: report them, stop the agent if that is possible without risk, and do not delete anything that helps the investigation.
Most of these rules appear this way in our AI policy template, above all in Annex A on AI agents. The template also regulates agent registers, permission levels, connectors and logging, and can be adapted to your company as a Word file. How much autonomy an agent should get for which task is described in our human-in-the-loop guide.
What IT should do
- Get an overview. Which assistants and agents are running, with which connectors (mailbox, storage, CRM, chat) and with which rights? An agent register answers that. A survey by Dataiku and The Harris Poll of 685 CIOs in eight countries, including Germany (9 to 29 July 2026), shows that this is no marginal issue: 81 percent do not have a full overview of agents that arise entirely outside approved systems or channels, and 84 percent agree that employees build agents and applications faster than IT can govern them.
- Check the three ingredients. For every agent ask: does it read untrusted content, does it see sensitive data, can it change something or communicate outward? If all apply, either remove one ingredient or have every action approved, as OWASP's Rule of Two requires.
- Close the ways out. Both cases ran through channels that look harmless: loading an image, a link preview, fetching a web address. OWASP explicitly names image addresses as a channel for data leaks. Where the platform allows it, switch off external images and link previews in agent responses or limit them to approved domains. This is exactly where Salesforce improved with Trusted URLs.
- Minimize rights, protect credentials. Connectors get only the rights they need, read-only where possible. IT treats API keys for AI services like production secrets and rotates them. Anthropic also recommends this in its Threat Intelligence Report of 10 September 2026. In one case described there, a single stolen developer token escalated to full administrator control over a cloud environment in about three hours.
- Log and monitor. Log agent actions as far as the platform allows, detect unusual data access by assistants (Varonis) and monitor the activity of autonomous agents (Anthropic). A sample of the logs is reviewed regularly.
- Test your own agents realistically. Anyone who builds their own agents tests them against attackers who know the protective measures. Filters alone are not enough: OWASP points to a study in which adaptive attacks reached a success rate of over 90 percent against most of 12 recent defense methods, while static attacks remained almost entirely unsuccessful.
How to set up protection systematically along the OWASP list is covered in Protecting against prompt injection.
If you want to clarify for your team which agents need which rights and how employees learn to handle them safely, talk to us for 30 minutes.
FAQ
What is prompt injection, simply explained?
Prompt injection means that a piece of text makes an AI model do something other than intended. It works because language models do not technically separate instructions from data. If a hidden instruction sits in an email or web page, an assistant can follow it as soon as it reads the content.
What is the difference between direct and indirect prompt injection?
With direct prompt injection, the user enters the manipulative input themselves, for example to override the safeguards of a chatbot. With indirect prompt injection, the instruction sits in content from other sources, such as web pages, documents, emails or form entries. The user neither entered nor saw this instruction, and it does not even have to be visible. For companies with agents, the indirect variant is especially tricky because the attacker needs no access, as SalesBleed shows.
Why is prompt injection critical for companies?
Because assistants and agents are connected to mailboxes, storage and CRM and are allowed to act there. According to OWASP, the damage of a successful injection depends on which tools and rights the agent has. CoSnitch and SalesBleed show that internal data can leak this way without anyone noticing.
Can prompt injection be prevented completely?
Not as things stand. In the 2026 edition, OWASP writes that there is currently no reliable protection mechanism, and the BSI describes countermeasures as means that make attacks harder or reduce the damage. That is why defense starts with what a manipulated agent can do at all.
How do you protect against prompt injection?
With several layers instead of a single filter: treat external content as untrusted, give agents minimal rights, allow risky actions only after human confirmation and limit ways out such as image addresses or link previews. Add logging, an agent register and tests against attackers who know the protective measures. The rules for employees belong in an AI policy.
Which real prompt injection examples were there in 2026?
Two well-documented cases are CoSnitch (CVE-2026-24301) in Microsoft Copilot Personal, published by Varonis on 18 August 2026, and SalesBleed in Salesforce Agentforce, published by Zenity Labs on 24 September 2026. According to the researchers, both flaws are fixed. For CoSnitch, Varonis has not seen exploitation in the wild.
Sources
- OWASP GenAI Security Project: OWASP Top 10 for LLM Applications 2026 (LLM01:2026 Prompt Injection, LLM03:2026 Excessive Agency): https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/
- OWASP GenAI Security Project, press release of 2 September 2026 (Top 10 2026, Agent Control Standard): https://www.prnewswire.com/news-releases/owasp-genai-security-project-releases-2026-top-10-for-llm-applications-debuts-agent-control-standard-and-new-resources-for-securing-generative-and-agentic-ai-302867085.html
- OWASP Top 10 for Agentic Applications 2026 (9 December 2025): https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/
- Varonis Threat Labs: CoSnitch, CVE-2026-24301 (18 August 2026): https://www.varonis.com/blog/cosnitch
- Zenity Labs: SalesBleed, 0-click data exfiltration on Agentforce (24 September 2026): https://labs.zenity.io/post/salesbleed-0-click-data-exfiltration-on-agentforce
- Zenity Labs: SalesBleed, Anonymous Phishing via Agentforce in Slack (24 September 2026): https://labs.zenity.io/post/salesbleed-hijacking-agentforce-in-slack-for-anonymous-phishing
- SecurityWeek: SalesBleed flaws in Salesforce Agentforce (25 September 2026): https://www.securityweek.com/salesbleed-flaws-in-salesforce-agentforce-enabled-zero-click-data-exfiltration/
- Microsoft Security Response Center: CVE-2025-32711, M365 Copilot Information Disclosure Vulnerability (11 June 2025): https://msrc.microsoft.com/update-guide/vulnerability/CVE-2025-32711
- ENISA Threat Landscape 2026 (22 September 2026): https://www.enisa.europa.eu/publications/enisa-threat-landscape-2026
- ENISA: ENISA's view on Cybersecurity in the Frontier AI Era (July 2026): https://www.enisa.europa.eu/sites/default/files/2026-07/ENISA%20view%20on%20cybersecurity%20in%20the%20frontier%20AI%20era_en_0.pdf
- BSI: Maßnahmen gegen Evasion Attacks auf große KI-Sprachmodelle (measures against evasion attacks on large AI language models, 10 November 2025): https://www.bsi.bund.de/DE/Service-Navi/Presse/Alle-Meldungen-News/Meldungen/Evasion-Attacks-LLM_251110.html
- BSI: Cybersicherheitswarnung BITS-B Nr. 2026-262788-1032 (cybersecurity warning, 22 June 2026): https://www.bsi.bund.de/SharedDocs/Cybersicherheitswarnungen/DE/2026/2026-262788-1032.pdf?__blob=publicationFile&v=2
- BSI blog, Claudia Plattner: Wenn eine KI aus ihrer Sandbox ausbricht (when an AI breaks out of its sandbox, 24 July 2026): https://www.bsi.bund.de/DE/Service-Navi/Presse/Alle-Meldungen-News/Blog/KI_Ausbruch_Sandbox_260724.html
- Anthropic: Threat Intelligence Report, Detecting and countering misuse of AI (10 September 2026): https://www.anthropic.com/threat-intelligence-report-september-2026
- Dataiku: Global AI Confessions Report, CIO Edition 2026 (24 September 2026): https://www.dataiku.com/company/news/global-ai-confessions-report-cio-edition-2026
Template
AI policy with an agent annex
An editable Word template for approvals, data classes, roles, AI literacy and AI agents. Free by email.
About the author
Co-Founder · Business & Content Lead
Co-Founder of Sentient Dynamics. 15+ years of business strategy (incl. SAP), MBA. Writes about EU AI Act compliance, ROI measurement and how Mittelstand CTOs actually adopt agentic AI.