Join our Newsletter — 33% off our NHI Course
Home› Glossary› AI Security› OWASP Top 10 for LLM Applications
AI Security

OWASP Top 10 for LLM Applications

← Back to Glossary
By NHI Mgmt Group Updated September 7, 2026 Domain: AI Security

OWASP Top 10 for LLM Applications is a community framework that ranks major security risks affecting systems built with large language models. It helps teams prioritise threats such as prompt injection, data leakage, and excessive agency. Security leaders use it as a practical reference for design reviews and testing.

Expanded Definition

OWASP Top 10 for LLM Applications is a community-driven risk catalogue for systems that use large language models in production. It is not a product checklist and it is not a formal standard; it is a practical way to describe the most important failure patterns that arise when an application delegates interpretation, generation, or tool use to an LLM.

The term covers prompt handling, retrieval-augmented workflows, model outputs, tool invocation, and surrounding controls such as filtering, access boundaries, and logging. It excludes generic AI concepts that do not materially change the security profile of the application. A common boundary mistake is to treat “LLM application security” as only a model problem, when many of the highest-impact failures emerge in orchestration, data flow, and trust decisions around the model.

For readers comparing authority sources, OWASP’s page is best used as a risk vocabulary for builders, while NIST AI Risk Management Framework provides a broader governance lens for AI risk management.

Examples and Use Cases

Teams usually encounter this framework when they are reviewing how an LLM is embedded into a real workflow rather than when they are testing a model in isolation. The risks become concrete once the application can read data, answer users, call tools, or influence downstream decisions.

  • A customer-support chatbot is given access to knowledge base content and must resist prompt injection embedded in retrieved documents.
  • An internal assistant summarizes tickets or documents and must avoid leaking sensitive data from prompts, context windows, or outputs.
  • An agentic workflow uses LLM output to trigger actions, which creates a need to test tool abuse, unsafe delegation, and excessive autonomy.
  • A RAG system is tuned to retrieve the “right” source, but the real issue is whether source selection, filtering, and citation handling are trustworthy.
  • A development team uses the framework during design reviews to decide where human approval, rate limiting, or output validation is required.

Implementation trade-offs often appear between usability and restraint: tighter controls reduce harmful behavior, but they can also make the assistant less useful if they block legitimate context or action.

Security Implications

Misunderstanding the OWASP Top 10 for LLM Applications can lead teams to secure the model while leaving the application exposed. That creates false confidence, because attackers and abuse paths often exploit the gap between what the model “knows” and what the application allows it to do.

Common consequences include prompt injection that changes system behavior, sensitive-data exposure through overbroad context, untrusted retrieval that contaminates answers, and tool misuse that turns text generation into real-world action. In practice, the failure condition is often weak separation between user input, system instructions, retrieval content, and execution authority.

A practitioner should watch for symptoms such as surprising tool calls, inconsistent refusal behavior, output that mirrors hidden instructions, or repeated leakage of internal prompts and stored content. These are usually signs that the control boundary is in the application layer, not the model layer.

The main blast radius is operational trust: once the assistant can be steered into wrong answers or unauthorized actions, the impact can spread into customer support, engineering, finance, or incident response workflows.

Domain and Governance Relevance

This term matters most in AI application governance, where teams need a shared language for recurring failure classes. It helps security, engineering, and product owners discuss whether a given LLM feature is informational only, advisory, or capable of taking actions that require stricter control.

For NHI-adjacent systems, the relevance becomes sharper when an LLM or agent can use tokens, service accounts, APIs, or delegated tools. At that point, the question is no longer only whether the model responds safely, but whether the surrounding identity and authorization model prevents excessive agency and unintended access.

That is why the framework is useful in design review, threat modeling, and control validation. It gives organisations a way to align risk discussions around real application behavior instead of treating all AI usage as a single category.

Used well, it supports governance decisions about ownership, testing, and escalation thresholds without pretending that every LLM feature belongs to the same risk tier.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI 600-1 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-03 — Secrets and Credential ExposureLLM apps often expose or misuse machine credentials through prompts and tool access.
Recommendation — Reduce exposed credentials in prompts, retrieval, and tool outputs.
OWASP Agentic AI Top 10A2 — Tool and Action AbuseAgentic LLM features can be steered into unsafe tool calls or delegated actions.
Recommendation — Constrain tool execution paths and verify high-impact actions before they run.
MITRE ATLASAML.TA0002 — Prompt InjectionPrompt injection is a core adversarial technique against LLM applications.
Recommendation — Map observed prompt-manipulation patterns to ATLAS and test those attack paths directly.
NIST AI 600-1GenAI — Generative AI ProfileThis term concerns generative AI application risks and governance choices.
Recommendation — Apply the GenAI profile to align testing, oversight, and risk treatment for LLM deployments.
CIS Controls v86 — Access Control ManagementLLM applications need access boundaries around data, tools, and execution authority.
Recommendation — Enforce least-privilege access for every service and integration the LLM can reach.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 7, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org