Platform teams should classify prompts locally before any cloud call, then route sensitive requests to a private model and generic requests to a frontier model. The key control is keeping embeddings and routing logic on infrastructure you control so secrets, internal plans, and regulated data are not exposed during the decision process. That preserves usability while reducing context leakage risk.
Why Sensitive Prompt Routing Needs a Local Classification Layer
The routing decision is part of the security boundary, not just an efficiency choice. If the platform sends prompts straight to a public model provider, the provider can see the full context before any sensitivity decision is made. A local classification layer lets the platform decide what can leave the environment, what must stay private, and what should be redacted or transformed before external processing.
This matters because prompts often contain more than a user request. They can include internal project names, operational plans, customer data, incident details, or embedded secrets from prior system context. Routing logic that runs locally can inspect those signals without exposing them to a third party first, which preserves the usefulness of shared model access while keeping the trust boundary where the team controls it.
Good routing also separates classification from generation. The classifier should make a narrow decision about sensitivity, destination, and handling, while the model call should only receive the minimum context required for that task. That keeps the architecture understandable and reduces the chance that a convenience feature quietly becomes a data-exposure path.
How Private Models, Public Models, and Redaction Fit Together
Most platform teams need more than a binary private versus public choice. A practical design uses a private model for sensitive prompts, a public frontier model for low-risk prompts, and a redaction or summarisation step when the request is useful but not fully safe to send as-is. The goal is not to block AI use, but to match the model to the data class and the purpose of the request.
The control only works if the routing inputs are trustworthy. That means embeddings, prompt classifiers, policy rules, and feature extraction should remain on infrastructure you control, and the route decision should be made before any cloud call that could echo the full prompt. Once context has been shipped outward, the platform has already lost the most important opportunity to prevent leakage.
For sensitive workloads, the safer pattern is to minimise what crosses the boundary rather than rely on the provider to ignore it. This is especially important when prompts are assembled from multiple systems, because the combined request may be far more sensitive than any single field looked at in isolation.
What Actually Leaks, and Why It Is Hard to Notice
Leakage is not limited to obvious secrets. Corporate context can slip out through instructions, examples, attached documents, retrieved snippets, metadata, and even the structure of the question itself. In practice, the risk is often “prompt overexposure,” where the platform sends more background than the model needs because the routing layer is too coarse or too late in the flow.
That is why routing policies need explicit handling for regulated data, confidential business context, and operational secrets. When the decision is based on visible attributes only after the prompt has been normalised locally, the platform can keep the model call narrow and avoid accidental disclosure through logs, tracing, or downstream provider processing.
For teams building or reviewing this pattern, the AI Infrastructure Workload Identity Guide is useful because it frames the control boundary around the systems that move AI workloads and their associated secrets. It also helps clarify why model routing, storage, and execution identity should be treated as part of the same operational trust chain.
Risk and Threat Considerations
The main risk is that the routing layer becomes a data-exposure chokepoint. If classification happens after the prompt has already left the environment, or if the classifier itself depends on external services, sensitive context can leak even when the final model selection looks correct. The same problem appears when logging, tracing, or observability tools record full prompts before the policy decision is applied.
Failure mechanism: The platform forwards full prompt context to a public provider, shared logging pipeline, or external enrichment service before local sensitivity classification and routing complete.
Impact: Secrets, internal plans, customer data, or regulated information can be exposed outside the intended trust boundary, creating confidentiality, compliance, and incident-response risk.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Limits what prompt routing and model access paths can see or send. |
| AU-2 — Event Logging | Prompt routing needs logging that records decisions without exposing raw sensitive content. | |
| Recommendation — Apply least privilege to the routing layer and restrict prompt access to the minimum needed. Log routing decisions and redact sensitive prompt content from audit records. | ||
| ISO/IEC 27001:2022 | A.5.12 — Classification of information | Prompt sensitivity routing depends on classifying information before external sharing. |
| A.8.11 — Data masking | Redaction or masking is a key control when only partial prompt content should leave the boundary. | |
| Recommendation — Classify prompts and attached context before any external model call. Mask or redact sensitive context before sending prompts to external providers. | ||
| CIS Controls v8 | CIS-3 — Data Protection | Sensitive prompts and embedded context need controlled handling and exposure reduction. |
| Recommendation — Protect prompt data with local classification, masking, and controlled egress. | ||
Practitioner Guidance
What to verify: Confirm that classification, redaction, and routing all run locally and that no pre-routing telemetry captures raw prompt content. If any part of the decision path depends on an external service, treat that path as a potential disclosure point.
Decision rule: If a prompt can influence business decisions, contain internal context, or be reconstructed into sensitive information, route it as sensitive by default unless the policy can prove it is safe to externalise. Use the public model only when the minimum necessary context is clear.
What good looks like: The platform can explain, for any request, why it was sent to a private model, a public model, or a redaction path, and the explanation is supported by local policy evidence rather than provider-side inspection.
Practitioner takeaway: The safest routing designs treat prompt classification as a local control over data exposure, not a convenience feature, because once corporate context leaves your boundary, the loss is often irreversible.
Related resources from NHI Mgmt Group
- How should teams preserve AI context across devices and model providers?
- How should security teams prevent sensitive data from leaking through AI prompts and copilots?
- How should security teams implement model capability checks in AI applications that route across multiple providers?
- How should public sector security teams use AI without increasing exposure to sensitive data?