Pause the transaction, verify through a separate channel, and require a control that does not depend on the same media path. If the request concerns credentials, payments, or access, escalate it through a predefined approval route instead of relying on the apparent authenticity of the message.
Why synthetic voice and video requests need a different control path
synthetic voice and video raise the assurance problem in a practical way: the message may look and sound authentic while bypassing the normal cues people rely on. The right response is to treat the request as unverified until it is confirmed through a separate channel and a different control, especially when the request can move money, reveal credentials, or change access.
The key issue is not whether the media is convincing, but whether the organisation can confirm the requester without depending on the same voice, video, chat, or inbox that may already be compromised. That means the receiving team should pause execution, preserve the request for review, and avoid making the first-line operator the only verifier.
What separate verification should actually look like
Separate verification should use a channel and decision path that are independent of the original request path. A callback to a known number, a pre-registered approval workflow, or an out-of-band confirmation from a trusted account can work if it is established in advance and not selected ad hoc during the incident.
The stronger the consequence, the less acceptable it is to rely on informal confirmation. For credential resets, payment instructions, or access changes, the control should require a second-person approval or a pre-defined authority chain. When the request cannot be validated quickly, the safer choice is delay, not improvisation.
Organisations also need to distinguish between verifying the content of the request and verifying the authority to act on it. A message may contain accurate project details and still be fraudulent. Good process design checks both: does the person or system have the right to ask, and is the request coming through a trusted route?
Where organisations usually fail under pressure
The most common failure is channel collision, where the same medium is used both to deliver the request and to confirm it. That creates a single point of failure for impersonation, deepfake abuse, and social engineering. Once staff are trained to “confirm by replying in the same thread,” the attacker only needs to own that thread.
Another weak point is exception handling. Teams often know the policy in normal cases, but high urgency, executive pressure, or operational disruption pushes them toward shortcuts. The fix is not just awareness, but a process that is still workable when the requester sounds urgent, familiar, or authoritative.
For organisations handling payments or privileged access, hardening the decision path matters as much as media detection. Controls such as separation of duties, two-person approval, and step-up review reduce the chance that one convincing synthetic clip can complete the full transaction chain.
Risk and Threat Considerations
Synthetic voice and video create a direct impersonation risk because they can exploit trust in human recognition, urgency, and authority. The consequence is highest when the request can trigger credential exposure, financial transfer, or access provisioning without an independent approval step.
Failure mechanism: The attacker or impersonator uses a convincing voice or video to steer the target into acting inside the same communication channel, where the apparent authenticity blocks normal challenge and the request is executed before verification occurs.
Impact: The result can be unauthorised payments, account takeover, privileged access changes, or the release of sensitive information, especially where staff are conditioned to treat familiar voice and face as proof.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5, NIST SP 800-63 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI09 — Human-Agent Trust Exploitation | Synthetic voice/video exploits trust in a requester's apparent identity. |
| Recommendation — Require out-of-band verification before acting on high-impact requests. | ||
| NIST SP 800-53 Rev 5 | IA-5 — Authenticator Management | High-risk requests can expose credentials and need controlled verification flows. |
| AC-6 — Least Privilege | Limits the damage if a synthetic request succeeds in changing access or privilege. | |
| Recommendation — Protect credential reset and release processes with stronger approval checks. Restrict request handlers to the minimum permissions needed to process approvals. | ||
| NIST SP 800-63 | Digital Identity Guidelines | Separate verification should rely on stronger authenticators and assurance than a voice or video call. |
| Recommendation — Use phishing-resistant authenticators for step-up verification of sensitive actions. | ||
| NIST CSF 2.0 | PR.AA-05 — Identity and Access Management | The scenario requires independent verification before granting access or approving actions. |
| Recommendation — Enforce separate approval paths for access, payments, and credential changes. | ||
Practitioner Guidance
What to prioritise: Build one response rule for high-risk requests, then make it easy to execute under stress. The rule should be simple enough that frontline staff can apply it without debating whether the media is real.
What to verify: Confirm that the verification channel is truly independent of the original path and that approvers are pre-enrolled. If the control depends on a number, inbox, or account copied from the request itself, it is not independent.
Decision rule: If the request changes money, credentials, or access, treat the message as untrusted until a separate approval path is completed. If that path is unavailable, hold the action rather than substituting a weaker check.
Practitioner takeaway: The control objective is not to detect every synthetic clip perfectly, but to ensure that one believable request cannot complete a high-impact transaction without an independent human or workflow check.
Related resources from NHI Mgmt Group
- How should organisations verify payment instructions when a video meeting or voice call could be synthetic?
- What breaks when organisations rely on voice or video to verify executives?
- What breaks when organisations still rely on voice or video verification for password resets?
- When should organisations rely on multiple verification methods instead of a single voice or video check?