Organisations should move beyond standard vendor questionnaires and assess how the vendor’s AI actually behaves with their data. That means identifying whether data is used for training or fine-tuning, where exposure could occur, how outputs are generated, and whether the model’s use is explainable and auditable. The goal is to measure real data risk, not just contractual assurance.
What Vendor AI Risk Evaluation Has to Prove
When a third-party product uses generative models on customer data, the key question is not whether the vendor has an AI policy. It is whether the vendor can show how customer data is handled at each stage of the model flow, from prompt ingestion to storage, retrieval, logging, and output generation. That is why contractual assurances alone are weak evidence unless they are backed by operational controls and traceable system behaviour.
Strong evaluations focus on data boundaries and model behaviour together. If customer content can be retained, reused for training, surfaced in logs, or exposed through downstream integrations, the risk is materially different from a vendor that processes data transiently and can prove segregation, retention limits, and auditability.
- Identify whether customer data is used for training, fine-tuning, retrieval augmentation, or human review.
- Confirm where prompts, outputs, embeddings, and logs are stored and how long they persist.
- Check whether the vendor can explain how outputs are generated and whether those steps are auditable.
- Assess whether the product enforces data segregation across tenants, environments, and support workflows.
A useful test is simple: could the vendor reconstruct who accessed what data, when it was processed, and whether it influenced a model response? If the answer is no, the organisation is evaluating a promise rather than a control.
Why Contract Terms Are Not Enough on Their Own
Vendor questionnaires often stop at policy language such as “we do not train on customer data unless permitted” or “we apply industry-standard safeguards.” Those statements matter, but they do not resolve exposure if the product architecture still allows broad retention, operator access, or indirect reuse through logging and telemetry. The practitioner concern is the mismatch between legal intent and technical reality.
This is especially important where generative AI is embedded inside a broader SaaS workflow. A vendor may have secure core infrastructure while still exposing customer data through support tooling, analytics pipelines, partner integrations, or model debugging paths. The real evaluation therefore has to include the full data path, not just the model endpoint.
- Require the vendor to separate policy claims from implementation evidence.
- Review whether support staff, subprocessors, or model providers can access customer prompts or outputs.
- Ask for retention, deletion, and tenant-isolation details for all AI-related data stores.
- Validate whether export, backup, and disaster-recovery copies follow the same protections.
In practice, the most common failure is assuming that “no training” means “no risk.” Even without training, customer data can still be exposed through logs, prompt history, retrieval indexes, or cross-environment access paths.
How to Turn Vendor Claims into a Defensible Assessment
Organisations should evaluate vendor AI risk as a control test, not a procurement formality. That means asking for evidence that can be reviewed, compared, and revisited if the product changes. For higher-risk use cases, the vendor should be able to demonstrate explainability, audit logging, access restrictions, incident handling, and clear boundaries on data use.
Where the vendor cannot provide that evidence, the decision is not necessarily to reject the product, but to constrain usage. That may mean limiting the data class allowed in prompts, blocking sensitive fields, disabling history features, or requiring contractual and technical safeguards before production rollout.
- Start with data classification and decide which customer data is prohibited, restricted, or allowed.
- Require proof for retention, deletion, logging, and support access controls.
- Test the product with realistic prompts to see whether outputs leak sensitive or cross-tenant information.
- Reassess whenever the vendor changes model provider, data handling, or feature set.
For organisations that want a broader identity and secret-risk lens on third-party exposure, the patterns in NHI Mgmt Group’s Ultimate Guide to Non-Human Identities help frame how unmanaged access paths and overexposed data can expand risk across vendor ecosystems. Breach case studies such as Vercel Context.ai OAuth Supply Chain Breach and Salesloft OAuth token breach show how third-party AI or integration paths can expose customer data when access boundaries are weak.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
CIS Controls v8, NIST AI RMF, NIST AI 600-1 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| CIS Controls v8 | CIS 3 — Data Protection | Customer data handling and retention in vendor AI workflows directly concerns data protection. |
| CIS 15 — Service Provider Management | Evaluating third-party AI products is a service-provider risk and assurance problem. | |
| Recommendation — Classify customer data used by vendor AI and enforce handling rules for storage, retention, and exposure. Require evidence of AI data handling, logging, and support-access controls from the vendor. | ||
| NIST AI RMF | GV.1 — Govern AI Risk Management Strategy | Vendor AI risk needs governance over acceptable use, accountability, and review criteria. |
| MAP.2 — Map Context and Stakeholders | The answer depends on mapping where customer data flows and who can access it. | |
| MEASURE.2 — Analyze and Assess AI Risks | The core task is measuring real data risk from model behaviour, not just policy claims. | |
| Recommendation — Define AI vendor risk thresholds, approval criteria, and review ownership before deployment. Map customer-data flows, data classes, and stakeholders across the vendor AI lifecycle. Assess retention, reuse, logging, and output leakage with testable evidence. | ||
| NIST AI 600-1 | GV.1 — Governance of GenAI Systems | Generative AI on customer data requires governance over use, oversight, and accountability. |
| MAP.1 — Use Context and Intended Function | Assessing how the vendor AI actually behaves requires understanding the intended function and deployment context. | |
| MEASURE.2 — Testing, Evaluation, and Monitoring | The page emphasizes testing actual behaviour, explainability, and auditability of model outputs. | |
| Recommendation — Establish governance for vendor GenAI features that process customer data. Document the intended use, data classes, and operating context for each GenAI feature. Test outputs, logging, and leakage behavior before approving customer-data use. | ||
| NIST CSF 2.0 | GV.RM-03 — Risk Management Strategy Established and Maintained | Vendor AI risk is a third-party risk decision that should be governed within the organisation's strategy. |
| ID.SC-01 — Supply Chain Risk Management Process | Third-party generative AI products are a supply-chain dependency requiring structured risk review. | |
| Recommendation — Set organisational thresholds for acceptable vendor AI data risk and review them regularly. Apply supply-chain review to vendor AI data flows, subprocessors, and AI providers. | ||
Practitioner Guidance
What to verify: Ask the vendor to show, not just state, whether customer data is retained, reused, logged, or made available to operators. If they cannot evidence the path from prompt to output, treat the product as higher risk until controls are clearer.
Decision rule: If the AI feature can process sensitive customer data, require a documented data-flow map and a testable retention and access model before allowing broad use. If the vendor cannot bound that use, restrict the feature to low-sensitivity data or block it entirely.
What practitioners underestimate: The biggest exposure often sits outside the model itself, in support access, telemetry, retrieval layers, and copied data in adjacent systems. That is why the assessment has to cover the full product ecosystem, not only the model statement in the contract.
Practitioner takeaway: Treat vendor AI risk as an evidence problem, the organisation should only trust customer-data use when the vendor can prove how the data is contained, audited, and prevented from taking on a life of its own.
Related resources from NHI Mgmt Group
- How should organisations assess third-party AI risk in vendor contracts?
- How should security teams implement ISO 42001 certification for AI systems that use customer data and third-party tools?
- How should security teams build an AI-BOM for cloud AI systems that use managed models, retrieval data, and third-party services?
- Why does model genealogy matter when organisations assess the risk of third-party AI models?