TL;DR: AI email summarisation can turn attacker-supplied text into trusted-looking “security alert” content inside Copilot workflows, with behaviour varying across Outlook and Teams surfaces, according to Permiso Security. The risk is trust transfer, because users often treat assistant output as system-generated even when it is attacker-shaped, and that breaks existing email security assumptions.
At a glance
What this is: Permiso Security shows that AI email summaries can be manipulated by attacker-supplied text to produce trusted-looking phishing content inside Copilot interfaces.
Why it matters: IAM and security teams need to treat AI summarisation as a new trust boundary because users may act on assistant output with less scrutiny than on the original message.
Context
AI email summarisation is now part of the phishing surface, not just a productivity feature. When a user asks Copilot to summarise untrusted email, the model can interpret attacker-shaped text as instructions and return content in a trusted-looking interface.
That changes the governance problem for human identity and access programmes. The question is no longer only whether an email was blocked or flagged, but whether an AI-generated summary can manufacture trust and steer user action before the user reaches the original message.
Permiso Security's testing shows that different Copilot surfaces do not behave identically, so organisations are dealing with more than one trust boundary inside the same workflow.
Key questions
Q: What breaks when AI email summaries can be shaped by attacker-controlled text?
A: The control that breaks is the assumption that users will inspect the original email before acting. AI summaries can reframe untrusted content as a polished system-like alert, so the trust decision moves from the inbox to the summary surface. That is why summarisation needs its own abuse testing and policy boundary.
Q: Why do AI-generated email summaries create a phishing risk?
A: Because they can transfer credibility from the assistant to attacker-supplied content. Users tend to trust polished, system-looking output more than raw email, so a malicious instruction can become a believable action prompt even when the original message looked suspicious.
Q: How should security teams evaluate Copilot summary interfaces?
A: Treat each summary interface as a separate control point and test whether it flags injected instructions, ignores them, or converts them into trusted-looking action prompts. The goal is to understand which surfaces can be abused to present phishing content with system authority.
Q: What should organisations do when AI summaries can draw from multiple Microsoft 365 sources?
A: Limit retrieval scope to the minimum data needed for the task and review whether internal context can be folded into externally triggered summaries. If the assistant can combine email with collaboration data, the attack surface expands from phishing into context-assisted deception.
Technical breakdown
Cross prompt injection turns email text into model instructions
Cross prompt injection happens when attacker-controlled content embedded in an email influences an LLM prompt at summary time. The model is not executing code, but it is still pattern-matching and instruction-following inside a mixed-trust input stream. That means the boundary between message content and system instruction becomes ambiguous unless the application isolates or sanitises untrusted text before summarisation. In practice, the attack works because the assistant is asked to transform the email, and the malicious instruction is already inside that email. The output can therefore reflect attacker intent while still looking like ordinary productivity assistance.
Practical implication: Treat summarisation prompts as untrusted input handling, not as a harmless UI convenience.
Why Copilot summary surfaces behave like separate control planes
The article shows that Outlook Summarize, the Outlook Copilot pane, and Teams Copilot do not share the same failure mode. That matters because each surface has its own rendering path, context access, and guardrail behaviour, so one interface may reject injected instructions while another echoes them. Security teams should think of these as distinct policy enforcement points rather than one feature. When controls differ by surface, the user experience still looks unified, but the risk posture is not unified. That is exactly where attackers look for the weakest trust boundary.
Practical implication: Inventory each AI summary surface separately and test them as separate control points.
Trust transfer is the real phishing primitive
The phishing problem is not just that Copilot can be influenced. It is that AI output is perceived as system-authored, polished, and therefore more credible than raw email text. That creates a model-mediated phishing pattern in which the attacker uses the assistant's voice, layout, and authority to compress user scepticism and increase action rates. In that sense, the summary panel becomes the social-engineering asset, not just the delivery channel. The summary can present a button, banner, or alert that looks operationally legitimate even though it is attacker-shaped.
Practical implication: Assume users will trust assistant output more than email body text and design controls around that trust asymmetry.
Threat narrative
Attacker objective: The attacker wants the victim to trust a Copilot-generated action prompt enough to click through and disclose information or follow a malicious link.
- Entry occurs when an attacker places hidden or low-visibility instruction text inside an otherwise routine email destined for Copilot summarisation.
- Credential or context abuse follows when the assistant processes that text and incorporates it into a summary that can also draw from Microsoft 365 sources such as Teams, OneDrive, or SharePoint.
- Impact arrives when the user trusts the Copilot-generated alert or action prompt and clicks a link or follows a call to action that supports phishing or data exfiltration.
Breaches seen in the wild
- CoPhish OAuth phishing via Copilot Studio: Datadog showed Copilot Studio agents on a Microsoft domain can front OAuth consent phishing and forward stolen tokens; no victims reported.
- Gemini AI Breach — Google Calendar Prompt Injection: Gemini AI assistant prompt injection attack leaks sensitive Google Calendar data.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Trust transfer is now a first-class identity risk: when users act on AI-generated summaries, they are responding to a trust signal that did not exist in the original email. That shifts the security problem from message filtering to interface authority. The practitioner implication is that identity-aware controls must account for where trust is formed, not just where content originates.
Summary surfaces are not functionally equivalent: Outlook Summarize, Outlook Copilot pane, and Teams Copilot should be governed as distinct policy environments because their guardrails and failure modes diverge. A control that works in one summarisation path may not hold in another. The implication is that governance must be surface-specific rather than brand-level.
Model-mediated phishing is the right named concept for this pattern: the attacker does not need to break authentication or deliver malware, only to shape an assistant into presenting malicious intent with system-like credibility. That is a governance shift because the abuse target is the credibility layer itself. Practitioners should treat credibility as an attack surface in human IAM and security awareness programmes.
Email security assumptions break when the assistant becomes the messenger: traditional controls assume the user inspects the message before taking action, but Copilot can become the first thing the user sees. That collapses the distinction between content review and trust establishment. The implication is that email governance now has to include AI-assisted rendering paths, not just inbound message controls.
Cross-app retrieval raises the stakes beyond classic phishing: once the assistant can blend email with context from Teams, OneDrive, or SharePoint, the summary can inherit internal data and amplify the attack's credibility. That creates a broader identity exposure problem because the output can present internal context as if it were an ordinary alert. The practitioner conclusion is that retrieval scope and summarisation scope must be governed together.
From our research library:
- The IBM/Ponemon 2025 Cost of a Data Breach Report found that phishing-initiated breaches cost an average of $4.8M each.
What this signals
AI summary tools should be governed as trust-bearing interfaces, not just productivity features. Once the assistant can present attacker-shaped text in an authoritative tone, the security team has to manage credibility leakage as part of the control model.
Model-mediated phishing: this is the name worth using for attacks that exploit assistant output rather than raw message content. The practical implication is that awareness, filtering, and retrieval controls all need to extend into the AI summary layer.
For practitioners
- Test each Copilot summary surface separately Validate Outlook Summarize, the Outlook Copilot pane, and Teams Copilot as independent trust boundaries. Record whether each surface flags, ignores, or echoes appended instruction text and whether any surface blends internal context into the summary.
- Red-team summary-output credibility Run phishing simulations that place low-visibility instruction text inside benign-looking mail and measure whether users trust AI-generated alerts more than the original message. Focus on whether the summary changes user action rates.
- Constrain cross-workspace retrieval Review which Microsoft 365 sources Copilot can surface during email summarisation and reduce retrieval scope where the summary does not need Teams, OneDrive, or SharePoint context.
- Tune email and DLP controls for AI rendering paths Update Safe Links, DLP, sensitivity labels, and user-warning patterns so they still trigger when malicious intent is delivered through a summary panel rather than the raw email body.
Key takeaways
- AI summaries can become the phishing payload when attacker-supplied text is rendered with assistant authority.
- The control gap is not only in inbox filtering but in the trust that users place in AI-generated output.
- Practitioners need separate governance for each summary surface, plus tighter limits on what the assistant can retrieve while summarising mail.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI09 — Human-Agent Trust Exploitation | The attack exploits user trust in assistant output rather than code execution. |
| Recommendation — Test whether assistant output can be trusted by users more than the source content and tighten that trust boundary. | ||
| OWASP API Security Top 10 | API6 — Unrestricted Access to Sensitive Business Flows | The summary flow can expose business actions and context through assistant-driven output. |
| Recommendation — Constrain summary-driven flows that can expose sensitive actions or context to the user. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | Copilot retrieval scope and summary surfaces depend on permissions and authorisations. |
| Recommendation — Review entitlement scope for AI summary workflows and reduce access to unnecessary Microsoft 365 sources. | ||
| NIST SP 800-53 Rev 5 | IA-5 — Authenticator Management | The article concerns trust and control over prompts and action paths rather than credentials, but authorization flow management still matters. |
| Recommendation — Govern AI-assisted workflows with explicit control over issuance and use of action-bearing links or prompts. | ||
Key terms
- Cross Prompt Injection: Cross prompt injection is an attack where hostile instructions are hidden inside content an AI system is asked to process, such as an email, document, or chat message. The model treats attacker text as input and may follow it during summarisation, retrieval, or response generation, even though the user never intended that content to become an instruction.
- Model-Mediated Phishing: Model-mediated phishing is a social engineering pattern where an attacker uses an AI assistant to deliver the lure instead of sending the lure directly. The assistant becomes the voice, formatting layer, or authority cue, which can make malicious instructions seem more trustworthy than the original email or message.
- Trust Transfer: Trust transfer is the security failure that occurs when users give AI-generated output more credibility than the raw content it was based on. In practice, the assistant’s polished tone, layout, or system-like framing can make attacker-shaped content feel legitimate, which increases the chance of unsafe user action.
- Summary Surface: A summary surface is any interface where an AI system turns raw content into condensed output for a user, such as email summaries or chat-based recap panels. Different surfaces may enforce different guardrails, so they should be governed and tested independently.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 24, 2026.
Updated on October 11, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org