The control assumption that only text can drive high-risk agent behaviour breaks down. A single manipulated image or audio clip can become a trigger for approvals, transfers or record changes, so security teams have to govern input-to-action paths rather than just prompt quality.
When text is no longer the only trigger path
Once an agent can interpret screenshots, PDFs or voice, the security problem shifts from prompt hygiene to input-to-action security. The control assumption that the dangerous instruction must appear as plain text no longer holds. Teams need to treat any user-visible or machine-readable input as potentially executable intent if the agent can turn it into an approval, transfer, deletion or disclosure.
That matters because multimodal input expands the attack surface without necessarily changing the user experience. A benign-looking invoice image, a screenshot of an approval request or a voice note can carry the effective command, while the agent’s downstream action still lands in the same business system. The right question is not “was the prompt well written?” but “what classes of input can cause an agent to take action?”
In practice, that means the agent’s decision boundary has to sit around the whole ingestion pipeline: capture, transcription, OCR, summarisation, classification, policy check and execution. If any one of those stages can be influenced, then the agent may be steered without ever receiving a traditional text prompt.
What breaks across screenshots, PDFs and voice
The biggest break is between content understanding and authority. A screenshot can smuggle an approval context, a PDF can package a fraudulent instruction with convincing formatting, and a voice clip can impersonate a familiar approver or colleague. If the agent treats these as trusted business artifacts, it may bypass the scrutiny that would normally apply to a typed request.
Another break is provenance. Multimodal systems often optimise for comprehension, not source validation, so they may fail to distinguish a real document from a manipulated one, or a live speaker from a replayed clip. That is why per-action authorisation for AI agents matters: the agent should not be allowed to convert an understood instruction into an irreversible action unless the request is both authorised and contextually appropriate.
The third break is operational, not just technical. Evidence, audit, and user intent all become harder to reconstruct when the triggering input is an image, file or audio stream. If your logging only records the final text prompt, you lose the material context needed to explain why the agent acted, whether the action was justified, and whether the same pattern could recur.
How to govern input-to-action paths safely
The practical fix is to govern the action path, not just the model prompt. That means classifying which multimodal inputs are allowed to influence which actions, requiring step-up approval for high-impact requests, and separating interpretation from execution so the model cannot directly trigger the most sensitive operations.
It also means narrowing trust at the content boundary. Voice commands for routine, reversible tasks may be acceptable, but the same channel should not be able to authorise payments, change records or release data without a stronger verification step. Likewise, a screenshot can be used as evidence for a workflow, but it should not by itself become the basis for a privileged decision.
For teams building controls, agent observability and incident response should include the original input modality, the transformed intermediate artefacts and the final action taken. That gives responders a way to tell whether the agent was legitimately instructed, socially engineered or manipulated through an untrusted file or recording.
Risk and Threat Considerations
Multimodal agents create a higher-impact social engineering path because the attacker no longer has to win a text-only prompt battle. If the model trusts screenshots, PDFs or audio too readily, the input channel itself becomes a phishing surface, with the agent acting as the compromised user would have acted.
Failure mechanism: A manipulated image, document or voice clip is accepted as authoritative intent, passes transcription or OCR checks, and is converted into a privileged action before a human can verify the source.
Impact: Attackers can trigger approvals, transfers, record changes, access grants or data disclosure while leaving only a weak forensic trail, especially if the organisation does not log the original modality and provenance.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Multimodal inputs can drive agent actions through abused authority. |
| ASI09 — Human-Agent Trust Exploitation | Screenshots, PDFs and voice prompts can socially engineer agent trust. | |
| ASI02 — Tool Misuse | Manipulated inputs can steer tools toward unsafe approvals or changes. | |
| Recommendation — Require per-action approval and constrain agent privileges before execution. Add source verification for non-text inputs before acting on them. Limit which tools each modality can reach and enforce action gating. | ||
| NIST AI RMF | GOVERN — Govern | The subject is about governing how multimodal inputs influence agent actions. |
| MAP — Map | Teams need to map modalities, workflows and downstream business impacts. | |
| MANAGE — Manage | Controls must manage multimodal agent risk over time and across workflows. | |
| Recommendation — Define policies for allowed modalities, approvals and escalation paths. Map input channels to the actions they can influence and rank by impact. Monitor multimodal trigger paths and update controls when new abuse patterns appear. | ||
Practitioner Guidance
What to prioritise: Inventory every action an agent can take from non-text input and rank those actions by blast radius. The highest-risk cases are not “interesting prompts”, they are irreversible business operations with weak human verification.
What to verify: Confirm that the system can prove which modality produced the trigger, which transformation layer interpreted it, and which policy allowed the final action. If you cannot trace those three points, you do not have enough assurance for high-impact automation.
Decision rule: If the input can originate outside your trusted workflow, treat it as untrusted intent until a separate control validates the request. Use human approval, step-up authentication or an out-of-band confirmation for anything that changes money, records or access.
Practitioner takeaway: The key design choice is not whether multimodal ai is useful, it is whether any single image, PDF or voice clip can still cause a material action without a second trust check.
Related resources from NHI Mgmt Group
- What breaks when AI agents are allowed to act on untrusted prompts without runtime guardrails?
- What breaks when AI agents are allowed to act inside privileged CI/CD workflows?
- What breaks when AI agents can act without a verified human behind them?
- What breaks when AI coding agents can act before a trust prompt appears?