Join our Newsletter — 33% off our NHI Course

What breaks when AI email summaries can be shaped by attacker-controlled text?

The control that breaks is the assumption that users will inspect the original email before acting. AI summaries can reframe untrusted content as a polished system-like alert, so the trust decision moves from the inbox to the summary surface. That is why summarisation needs its own abuse testing and policy boundary.

How attacker-shaped summaries break the email trust boundary

The core failure is that the summary becomes a new decision surface. If the model can be nudged into omitting qualifiers, amplifying urgency, or recasting a request as routine, the user no longer evaluates the sender, the raw message, and the surrounding context together. That changes the security property from message inspection to summary trust, which is much easier to exploit.

It also means the control is not just “did the summary look accurate,” but “did the summary preserve the parts of the email that matter for a safe action.” The risk is highest when summaries are treated as a surrogate for the inbox rather than a convenience layer over it. In practice, that is where benign-looking phrasing can hide payment requests, approval nudges, credential prompts, or policy exceptions.

Why this is a prompt-injection and social-engineering problem

Attacker-controlled text in an email can act like prompt injection when the summariser ingests it without a hard boundary between untrusted content and system instructions. The model may not execute code, but it can still be steered into changing emphasis, suppressing warnings, or adopting the attacker’s framing. That makes the summary a social-engineering amplifier even when the original message was obviously suspicious.

When this happens, the attacker is not trying to beat technical email filtering alone. They are trying to influence interpretation at the moment of action. A summary that sounds authoritative, neutral, or operational can lower scepticism more effectively than a noisy phishing message, especially for busy users who rely on the assistant to triage at speed.

What controls need to change when summaries become an input to action

Summarisation should be treated as a security-sensitive transformation, not a convenience feature. The model must preserve provenance cues, message uncertainty, and visible links back to the original content, because the summary should never be the only artefact users see before acting. If the summary is allowed to decide what is “important,” then the assistant is effectively performing risk triage and needs abuse testing accordingly.

This is where policy boundaries matter. A safe design assumes the summary can be influenced, so it limits the actions a summary can authorise on its own, keeps raw-message access one click away, and subjects the summariser to the same kind of adversarial review you would apply to any trust-facing control. For an adjacent governance perspective on how attacker-driven context poisoning and identity abuse can alter AI behaviour, see OWASP Agentic AI Top 10 and MITRE ATLAS adversarial AI threat matrix.

Risk and Threat Considerations

The risk is not limited to incorrect summarisation, it is misdirection at the point where users decide whether to click, reply, approve, or escalate. If attacker text can shape the summary, then the attacker can reduce friction, hide caveats, and manufacture false confidence, which turns a productivity feature into an attack surface.

Failure mechanism: The summariser ingests untrusted email text and elevates attacker-authored phrasing over sender provenance or original context, so the assistant output becomes more persuasive than the source message.

Impact: Users may approve fraudulent requests, miss warning signs, or trust a synthetic summary more than the email itself, creating a faster path from phishing to action.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
OWASP Agentic AI Top 10 ASI06 — Memory & Context Poisoning Attacker-shaped text can poison the model's local context and distort summary output.
ASI09 — Human-Agent Trust Exploitation The attack relies on users trusting a polished assistant summary over the source email.
ASI03 — Identity & Privilege Abuse A shaped summary can steer privileged user actions and approvals through trust abuse.
Recommendation — Test summary prompts against injected instructions and preserve hard trust boundaries for untrusted text. Design review flows so users confirm critical actions from the original message, not the summary alone. Constrain assistant-driven approvals and require explicit human confirmation for high-impact actions.
MITRE ATLAS Adversarial Machine Learning The scenario is an adversarial manipulation of an AI system's interpretation of input text.
Recommendation — Harden the summarisation pipeline against input manipulation and adversarial prompt influence.
NIST AI RMF Govern AI governance is needed to manage trust boundaries, testing and oversight for summarisation features.
Recommendation — Define approval, testing and monitoring controls for summaries that influence user decisions.

Practitioner Guidance

What to prioritise: Preserve source visibility and uncertainty markers before you optimise summary quality. If the summary can trigger action, require a clear path back to the original email and make the raw message the authoritative record for approval decisions.

What to verify: Test whether the summary still exposes sender identity, request type, urgency claims, and any conflicting details from the original message. If those elements disappear under paraphrase, the control is too brittle to trust for action-taking workflows.

Practitioner takeaway: The right design assumption is not that summaries will be perfectly faithful, but that they will be influenceable, so the safe workflow is one where a summary can assist triage without becoming the sole basis for trust.