They focus on whether the model explicitly produced the banned content or named the obvious target. In practice, a model can still enable harm by returning adjacent terms, search prompts, or destination paths that lead users to the same outcome. That is a referral problem, not just a generation problem.
Why This Matters for Security Teams
Harmful content referrals are easy to miss because they sit between policy failure and user intent. A model may never state the prohibited material directly, yet still guide a user toward it through euphemisms, synonyms, adjacent search terms, oblique source names, or destination paths that remove only one step from misuse. That means the control question is not simply whether a banned answer was generated, but whether the system materially enabled the harmful outcome.
This is why current guidance from the NIST AI Risk Management Framework matters: teams need to assess harmful output as a lifecycle risk across design, deployment, and monitoring, not as a one-off content filter. The same logic appears in the OWASP Agentic AI Top 10, where tool use, delegation, and indirect actions can create exposure even when the text looks harmless on its face.
Security teams often overvalue keyword blocking and underinvest in outcome-based review, which is why harmful referrals slip through policy gates that look strong on paper. In practice, many teams encounter referral abuse only after users have already followed the breadcrumb trail to misuse, rather than through intentional detection of the pathway itself.
How It Works in Practice
Operationally, a harmful referral often appears as a safe-looking response that still contains actionable pointers. Those pointers might include related search terms, names of tools, categories of websites, encoded hints, or stepwise suggestions that help the requester infer the prohibited target. The risk increases when the model is allowed to browse, call tools, or chain prompts, because referral content can be amplified by external retrieval and automation.
Effective controls need to inspect more than final text. Teams should evaluate the full interaction path, including prompt, retrieved context, model output, and downstream tool calls. The NIST AI 600-1 Generative AI Profile and the MITRE ATLAS adversarial AI threat matrix both support this broader view because they push teams to think about adversarial manipulation, abuse paths, and system-level controls.
- Classify outputs by intent and effect, not just by explicit banned terms.
- Detect oblique referrals such as redirects, aliases, code words, and search instructions.
- Log retrieval sources and tool invocations so reviewers can reconstruct the harm path.
- Apply policy checks before and after generation, especially for agentic workflows.
- Escalate uncertain cases to human review when the model provides a plausible route to misuse.
Where the system includes autonomous actions, the problem becomes closer to agentic security than simple content moderation. The CSA MAESTRO agentic AI threat modeling framework and the OWASP Top 10 for Agentic Applications 2026 both reinforce that indirect enablement, not only direct generation, should be treated as a security event. These controls tend to break down when retrieval is unconstrained and the model can surface search-friendly pointers faster than reviewers can judge the downstream harm.
Common Variations and Edge Cases
Tighter referral controls often increase false positives and review overhead, requiring organisations to balance harm reduction against usability and operational cost. That tradeoff is especially visible in customer support, compliance assistance, and research workflows where legitimate redirection is expected and where there is no universal standard yet for what counts as an impermissible referral.
One common edge case is dual-use language. A model may provide terms that are legitimate in one context and harmful in another, so static blocklists rarely suffice. Another is “safe completion drift,” where the model starts with a compliant answer and then adds one extra hint that becomes the harmful bridge. Teams should therefore define referral risk at the interaction level, not sentence by sentence.
Agentic systems raise a further complication: a model may never produce the harmful referral in the visible answer, but may expose it through a tool call, a retrieved document snippet, or a chained agent response. That is why the intersection between content safety and agent governance matters. Best practice is evolving, but current guidance suggests treating hidden pathways, not just surfaced text, as part of the attack surface. For teams building governance around that boundary, the NIST AI Risk Management Framework remains the broad anchor, while the Anthropic report on AI-orchestrated cyber espionage is a useful reminder that seemingly indirect model outputs can still operationalize abuse.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, MITRE ATLAS and CSA MAESTRO address the attack and risk surface, while NIST AI RMF and NIST AI 600-1 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | Risk management must cover harmful referrals across the AI lifecycle. | |
| NIST AI 600-1 | GenAI profile supports controls for unsafe outputs and misuse pathways. | |
| OWASP Agentic AI Top 10 | A1 | Agentic systems can turn harmless-looking text into harmful action paths. |
| MITRE ATLAS | Adversarial AI tactics include prompt manipulation and harmful guidance paths. | |
| CSA MAESTRO | Threat modeling for agentic AI should include indirect harmful referrals. |
Assess referral harm as a lifecycle risk and monitor outputs, tools, and user impact continuously.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on August 20, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org