Skip to the FAQ
AI security · OWASP 2025

The OWASP Top 10 for LLMs, explained.The ten risks in applications built on language models: how each one is exploited, what shuts it down and what changes when the model runs on your own servers.

An assistant that reads email, a search tool over your documentation or an agent with database access all obey a model that cannot tell your instructions from the text that reaches it from outside. Almost everything else follows from that.

10risks
2025current edition
0patches for injection
5risks that shift on-prem

OWASP, the foundation that has spent more than twenty years publishing the web security Top 10 much of the industry relies on, has kept a separate list since 2023 for applications built on large language models (LLMs). The current edition is the 2025 one, and the full text is on the OWASP GenAI project site. Here we go through it risk by risk, with an example of how each one is exploited and what closes it in practice. We also add two things vendor guides tend to skip: what changes when the model runs on your own infrastructure, and what to log so you hear about an attack before someone else tells you.

The root flaw


An LLM application hands the model a single block of text that mixes the system instructions (the rules written by whoever built it), the user's message and, very often, documents, web pages or tool results added along the way. The model reads all of it as text, which is why there is no technical boundary between ‘this is an order from my developer’ and ‘this is the content of an email I have been asked to summarise’.

If the email says ‘forward the latest report to this address’, an agent with mailbox access may well do it.

Keep that in mind and the list makes sense on its own: half the risks describe how malicious text gets in, and the other half how much damage it can do once inside.

The ten risks, one by one


LLM01 · Prompt injection

How it is exploited

Directly, when the user types ‘ignore everything above and…’, or indirectly, which is the dangerous kind: the instruction hides in a PDF, a web page the agent reads or an email the assistant summarises, and the user never sees it.

What shuts it down

Nothing completely. You limit the damage: external content flagged as untrusted, least privilege, authorisation outside the model and output validation. It gets its own section below.

LLM02 · Sensitive information disclosure

How it is exploited

In a RAG assistant (one that searches your documentation before answering), payroll files end up in the same index as the holiday policy and anyone can ask about them.

What shuts it down

Applying the user's permissions at retrieval time, before any text reaches the model, indexing only what is needed and reviewing what data shows up in answers and logs.

LLM03 · Supply chain

How it is exploited

A third-party model, set of weights, Python library or dataset arrives tampered with or under a licence that does not allow your use. Weights are huge, opaque files, so nobody looks inside.

What shuts it down

Verified sources, pinned versions, checked hashes and formats that do not run code when loaded (safetensors or GGUF instead of pickle).

LLM04 · Data and model poisoning

How it is exploited

Someone slips manipulated data into training, fine-tuning or the RAG documents and manages to skew answers or plant a backdoor triggered by a specific word.

What shuts it down

Knowing where every document comes from, controlling who can add content to the index and rerunning a set of known cases after every change.

LLM05 · Improper output handling

How it is exploited

Model output is rendered on a web page, run as an SQL query or used as a command without checks. It is the same old injection with a new middleman: XSS, SQL injection or remote code execution.

What shuts it down

Treating model output like any user input: validate it, escape it and use parameterised queries.

LLM06 · Excessive agency

How it is exploited

An agent whose tools have more permissions than it needs, or that chains actions without anyone confirming them. That turns an injection from an odd answer into an email sent, a file deleted or a payment made.

What shuts it down

Least privilege per tool, authorisation checked in the service that performs the action and human confirmation for anything irreversible.

LLM07 · System prompt leakage

How it is exploited

Getting the model to repeat its instructions is easy. The problem is what some teams put in them: API keys, internal server names or business rules.

What shuts it down

Assuming those instructions will end up public, keeping secrets out of them and moving the rules that matter into controls that do not depend on the model.

LLM08 · Vector and embedding weaknesses

How it is exploited

The RAG vector database does not keep each customer's data apart, or lets the wrong people insert documents, and search returns what it should not, or what an attacker has crafted to show up.

What shuts it down

Separate indexes or permission filters on every query, restricted write access and a log of which fragments were used in each answer.

LLM09 · Misinformation

How it is exploited

The model is confidently wrong, invents references or accepts figures that do not add up, and that ends up in a report or on the website with nobody checking it.

What shuts it down

Answers that cite their source, an assistant that abstains when it finds none and human review for anything that gets published or decided.

LLM10 · Unbounded consumption

How it is exploited

A malicious user or a loop in an agent sends the bill soaring or takes the service down, and with enough well-chosen requests an attacker can copy the model's behaviour.

What shuts it down

Per-user quotas, input and output size limits, a maximum time per request and alerts when usage goes beyond normal.

If you work with agents

In December 2025 OWASP also published the Top 10 for Agentic Applications 2026, focused on systems that plan and act on their own. It does not replace this list: it extends it for the case where LLM01 and LLM06 meet.

0

patches that close prompt injection.

Which is why OWASP keeps it at number one

Prompt injection in detail


Almost every team's first idea is to filter suspicious phrases, and it falls short, because an instruction can be written a thousand ways, in another language, encoded or spread across several documents. Dedicated classifiers such as Prompt Guard help as one more layer and stop the crude attempts, but the approach that holds up is to take for granted that injection will happen and design so that, when it does, the damage is small.

Path of an indirect injection and the points where it is cut off System instructions written by your team User message ‘summarise this email’ Email, PDF or web page ‘forward the report to…’ hidden instruction LLM one block of text 1 2 3 Tool send email 4 · SIEM 1 least privilege 2 authorisation outside the model 3 human confirmation
The malicious instruction comes in mixed with the rest of the text. You cannot stop the model from reading it, but you can stop the tool from carrying it out: three cut-off points before the action and a log that raises the alarm if one of them fails.

In practice that comes down to six design decisions, and none of them relies on the model behaving itself:

  1. External content goes in flagged as such and kept apart from the instructions when the request is built, so the model and the filters know what is data.
  2. Each tool gets the smallest permission that does the job: an agent that summarises emails has no need to send them.
  3. Authorisation is checked outside the model, in the service that performs the action and with the real user's identity, never the agent's.
  4. Output is validated against what is expected before it is used, just like a form.
  5. Anything irreversible (sending, deleting, paying) asks a person to confirm.
  6. Everything is logged somewhere someone actually looks, because an injection that works throws no errors and can go unnoticed for weeks.

And all of it gets tested. A set of known injections, run every time the model, the instructions or the documents change, is what separates a defence you believe works from one you know works.

If the model runs on your servers


More and more companies run the model on their own infrastructure so that data stays in-house, and that is a good reason (we cover it in which server you need for a local LLM). But a local model inherits no protection from sitting inside your data centre: the risks get redistributed, and five of the ten move.

RiskWith an external APIWith the model on-prem
LLM02 Sensitive informationData goes to the provider and you depend on their contract.It stays in-house, but RAG permissions are still your job.
LLM03 Supply chainThe provider chooses and protects the weights.You download them: format, origin and hash become your problem.
LLM04 PoisoningYou do not control base training.If you fine-tune, the dataset and who touches it are your responsibility.
LLM07 System promptIt travels to the provider with every request.It stays inside, though the model can still repeat it.
LLM10 Unbounded consumptionAbuse shows up on the per-token bill.There is no bill, but the GPU saturates and the service goes down for everyone.

The supply chain point deserves a specific warning: weight files saved with pickle, PyTorch's classic format, can run code the moment they are loaded. Downloading a model from any old repository and opening it amounts to running a stranger's program. Safetensors and GGUF store only the numbers, which is why they are worth insisting on. Consumption works in a similar way: inference servers such as vLLM or Ollama do have limits (concurrent requests, maximum context length, queue size), but their defaults are tuned for throughput rather than for withstanding abuse, so they need adjusting.

Prompt injection and excessive agency, on the other hand, do not move an inch. A local agent with write access to the database is exactly as dangerous as one in the cloud.

1

pickle weight file is enough to run code when the model loads.

Safetensors or GGUF, every time

What to log so you find out in time


An injection that works leaves no errors in any log: the model does what it was asked, the tool reports success and the user notices nothing. That is why detection has to be built on purpose, recording for every request who made it, which document fragments were retrieved (with their identifiers), which tools were called and with what parameters, which validations failed and how many tokens it used.

All of that goes to the SIEM, the system where security logs and alerts are centralised (QRadar, Wazuh or whichever you use). From there you can write rules that make sense for an LLM application:

  • An agent calls a sending or writing tool right after reading an external document.
  • An answer includes fragments from documents belonging to a group the user is not part of.
  • A user's or agent's consumption jumps well above its average for the past week.
  • The injection classifier or output validation keeps rejecting requests from the same account.

No SIEM ships with these rules, because they depend on how your application is built, but they are what turns a penetration test finding into an alert the security team sees on the day it happens.

Where to start


If you have an AI application live or about to launch, the order we suggest starts with the map: which data comes in, which tools it can use and where the secrets are, because without it you cannot prioritise anything. Next come tool and document permissions (LLM02, LLM06 and LLM08 tend to cause the most expensive scares) and usage limits, and finally the testing and logging we have just described.

Do not forget the classic side either: many LLM applications are exposed as just another API, and the controls in the OWASP API Security Top 10 still apply there.

We work through all of this with security and development teams in our AI security course: OWASP for LLMs and red teaming, attacking and defending a deliberately vulnerable application. And if what you need is to design secure agents from the start, that is our secure AI agent integration service.

Frequently asked questions


What is the OWASP Top 10 for LLMs?

It is the list of the ten most important security risks in applications that use language models, published by the OWASP foundation. The current edition is the 2025 one and runs from LLM01, prompt injection, to LLM10, unbounded consumption.

How does it differ from the classic OWASP Top 10?

The classic Top 10 covers web applications in general. The LLM list focuses on what changes when a language model is involved, such as prompt injection, agents with too many permissions, RAG or unbounded consumption, and both apply at the same time.

Can prompt injection be prevented completely?

Not today. What you can do is limit the damage with least privilege, authorisation outside the model, output validation, human confirmation for irreversible actions and monitoring.

Is a local model more secure?

It keeps data from going to an external provider, but it hands your team the supply chain of the weights and control over consumption, and it does not protect against prompt injection or agents with too many permissions.

What is the OWASP Top 10 for Agentic Applications?

It is a companion list OWASP published in December 2025 for systems that plan and carry out actions on their own. It extends the Top 10 for LLMs into agents rather than replacing it.

SIXE

Got an assistant or an agent up and running?

Tell us what data it reads, which tools it can use and with what permissions, and we will tell you where we would start locking it down. If what you want is for your team to know how to attack and defend it, that is what the SXIA07 course is about.

AI security training and services · sixe.eu
SIXE