Join our Newsletter — 33% off our NHI Course
Home FAQ AI Security How should security teams evaluate AI chat tools…
AI Security

How should security teams evaluate AI chat tools that claim encryption?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated August 21, 2026 Domain: AI Security

Start by separating transport security from runtime privacy. Ask where plaintext exists, who can access it, whether saved history is zero-access or merely stored, and whether the vendor trains on prompts. If those controls are not explicit, the product is not suitable for sensitive data.

Why This Matters for Security Teams

Encryption claims often sound decisive, but they can hide very different protection models. A chat tool may encrypt data in transit while still keeping plaintext in memory, logging prompts for support, or retaining conversation history in ways that create exposure. Security teams need to evaluate the full data path, not just the marketing label, because the real question is whether sensitive prompts remain confidential during processing, storage, retention, and operator access.

That distinction matters because AI chat tools are frequently used for code review, incident response drafting, policy analysis, and summarising internal material. If a vendor cannot clearly explain how plaintext is handled, whether administrators can inspect content, and whether prompts are used for training, the control design is incomplete. That is a governance problem as much as a technical one, and it maps naturally to the NIST Cybersecurity Framework 2.0 categories for data protection, access control, and risk management.

In practice, many security teams encounter weak AI data handling only after employees have already uploaded sensitive material into a tool that was never reviewed through procurement or security architecture.

How It Works in Practice

Start by testing the vendor’s encryption claim against the life cycle of a prompt. Transport encryption protects data in transit between the user and the service. It does not answer what happens after the request reaches the model host, how long data persists, or which roles can access logs and telemetry. For security evaluation, the important questions are: is data encrypted at rest, is the key management model customer-controlled or vendor-controlled, is the service zero-retention by default, and is history searchable by vendor staff?

Security teams should ask for answers in writing and map them to operational controls. A useful review usually covers:

  • Where plaintext exists during inference, logging, monitoring, and support troubleshooting.
  • Whether prompts, outputs, and attachments are retained separately or as a single record.
  • Whether training, fine-tuning, or product improvement can use customer inputs by default.
  • Whether administrators, support engineers, or subcontractors can access content.
  • Whether the tool supports tenant isolation, export controls, deletion, and audit logging.

For procurement and risk teams, the right evidence includes a data flow diagram, retention schedule, encryption architecture, access model, and DPA or terms of service that state whether prompts are used for model training. If the tool supports enterprise controls, compare those controls to the organisation’s AI acceptable use policy and data classification scheme. Current guidance suggests treating prompts as sensitive data whenever users might paste secrets, regulated records, or internal strategy into the chat interface.

That review should also distinguish product claims from implementation reality. A vendor may advertise encryption, but if the service uses content for abuse detection, model quality assurance, or human review, the confidentiality posture is still weakened. OWASP guidance for LLM applications is useful here because it forces teams to look beyond transport security and assess prompt injection, data leakage, and unsafe output handling alongside privacy controls. These controls tend to break down when a tool is integrated quickly into browser extensions, SaaS copilots, or shadow IT workflows because the organisation loses visibility into where prompts are copied, cached, or re-shared.

Common Variations and Edge Cases

Tighter privacy controls often increase cost, reduce convenience, or limit model features, requiring organisations to balance usability against confidentiality. Some tools offer strong enterprise claims but only for paid tiers, while consumer versions retain prompts by default. Others provide customer-managed keys but still keep operational metadata that can expose usage patterns or business context.

There is no universal standard for this yet, so security teams should avoid assuming that “encrypted” means “private enough for sensitive data.” The strongest posture usually comes from a combination of transport encryption, at-rest encryption, zero-retention defaults, clear non-training commitments, and restricted administrative access. In regulated environments, the assessment should also check whether the service creates records that fall under sectoral retention or eDiscovery obligations, because deleting chat history may not remove all copies from backups, logs, or legal archives.

For high-risk use cases such as legal, finance, HR, or incident response, current guidance suggests preferring tools with explicit enterprise isolation and auditability rather than consumer chat products with broad content rights. The NIST Cybersecurity Framework 2.0 is still a good baseline for aligning these decisions to governance and monitoring, but the specific privacy test should be whether the service can prove that sensitive prompts are not reused, overshared, or retrievable by unauthorised parties.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST CSF 2.0, NIST AI RMF and NIST AI 600-1 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0PR.DS-1The question centers on data confidentiality claims and how sensitive prompts are protected.
NIST AI RMFAI risk governance is needed to assess whether the tool handles prompts safely and transparently.
OWASP Agentic AI Top 10Prompt leakage and unsafe tool behavior are common AI application risks in chat interfaces.
NIST AI 600-1GenAI profile guidance helps teams evaluate prompt privacy, logging, and misuse risks.
MITRE ATLASAML.TA0001Adversarial AI threat modeling helps assess abuse paths and data exposure in chat systems.

Apply AI risk governance to document retention, access, training use, and confidentiality assumptions.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on August 21, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org