Skip to content
Red Cell

What Is AI Red Teaming? Prompt Injection, Jailbreaks and Agentic Abuse

Kenneth Brown

Definition · Published · 4 min read

The short definition

The National Institute of Standards and Technology (NIST) defines a red team as "a group of people authorized and organized to emulate a potential adversary's attack or exploitation capabilities against an enterprise's security posture."

AI red teaming applies that to systems built on a large language model (LLM): assistants, copilots, retrieval applications and agents. An authorized operator works out what the system can be made to do, what it can reach when it does, and what happens downstream when its output is trusted, then shows the evidence.

Test the system, not just the model

A deployed assistant is rarely just a model. It is an application with a system prompt, an identity, a set of tools it can call, a store it can read from, and a user on the other end whose input becomes part of its instructions.

That distinction decides what an assessment is worth. Model providers already test their models for harmful output. What they cannot test is your deployment: the service account your agent runs as, the documents your retrieval layer can surface for a given user, or the internal application programming interface (API) your assistant can call. Those are where the consequential failures live, and they are specific to how your system is wired.

What an AI red team looks for

The Open Worldwide Application Security Project (OWASP) publishes a Top 10 for LLM Applications, which is useful shared vocabulary. The failures an assessment works through map onto it closely:

FailureWhat is assessedOWASP 2025
Prompt injectionWhether instructions in user input, or in content the system retrieves, are obeyed as instructionsLLM01 Prompt Injection
Safeguard bypassHow reliably the guardrails the application depends on hold when a request is reframed or spread across turnsLLM01 Prompt Injection
Sensitive information exposureWhat the system reveals about its configuration, its data sources and other users' recordsLLM02 Sensitive Information Disclosure, LLM07 System Prompt Leakage
Tool misuse and excessive permissionsWhat actions the agent can take, with whose identity, and what a stray instruction can make them doLLM06 Excessive Agency
Cross-user and cross-tenant isolationWhether one user's data or context can reach another user or tenant through the assistantLLM02 Sensitive Information Disclosure, LLM08 Vector and Embedding Weaknesses
Memory and content poisoningWhether content written today changes the system's behaviour for someone else tomorrowLLM04 Data and Model Poisoning
Unsafe output handlingWhat the surrounding application does with the model's output: rendering, executing, querying or forwarding itLLM05 Improper Output Handling
Approval bypassWhether actions that need human approval can be avoided, or approved on the basis of a misleading summaryLLM06 Excessive Agency
Common failure classes in deployed AI systems, with the closest OWASP Top 10 for LLM Applications (2025) entry.

Prompt injection

A model cannot reliably tell an instruction from data. Anything the application feeds it arrives with the same standing as its system prompt unless the architecture separates them. Direct prompt injection comes from the user typing into the assistant. Indirect prompt injection is more dangerous: the instruction hides in content the system reads on someone's behalf, such as a support ticket, a shared document, an email or a web page, and fires when an unsuspecting user asks the assistant about it.

Jailbreaks and safeguard bypass

A jailbreak is a request crafted to get around the rules a system is meant to follow. For a business application, the question is not whether the model can be coaxed into saying something off-script, but which of the application's safeguards are enforced in code and which exist only as wording in a prompt. Safeguards that live only in the prompt are advisory, and an assessment should tell you which of yours are which.

Agentic abuse

Once an assistant can call tools, a content problem becomes an action problem. An agent that can send email, issue refunds, change tickets or write to a system is an authenticated user of everything it can reach. Agentic abuse covers getting an agent to take those actions for the wrong person: calling a tool with another user's identifier, chaining tools into something no single tool would allow, or satisfying a human approval step with a summary the agent wrote itself.

How an AI red team engagement runs

An AI assessment follows the same authorization discipline as any offensive engagement: written scope, rules of engagement and signed authorization come first. The work itself then runs in six stages:

  1. Inventory. Map what is actually deployed: entry points, system prompts, retrieval sources, tools, the identity each runs as, and where output goes.
  2. Threat model. Decide which failures would matter for this system and this business, and which trust boundaries carry the weight.
  3. Agreed test cases. Turn the threat model into cases you sign off before execution, including anything that must be left alone.
  4. Controlled evaluation. Run the cases against the agreed environment, within the intensity and stop conditions in the rules of engagement.
  5. Findings. Report what reproduced, what it reached, and the control that removes it, with evidence for each case.
  6. Remediation validation. Re-run the affected cases after the fix, and hand over the ones worth keeping as regression checks.

What you should receive

  • A threat model of the deployed architecture that your team keeps using after the engagement.
  • Reproducible evaluation cases your engineers can run themselves, with the expected safe behaviour stated next to the observed one.
  • Evidence for each finding: transcripts, tool invocations and responses.
  • Control recommendations mapped to the failure each one closes, not to a generic checklist.
  • Regression checks worth running continuously, in a form you can add to your own pipeline.

If the findings need to feed a governance process, the NIST AI Risk Management Framework and its Generative Artificial Intelligence Profile are the usual reference points.

Questions to ask an AI red teaming provider

Questions for an AI red teaming provider

  • Do you test the deployed application, its tools, permissions and retrieval, or only the model's responses?
  • Who performs the testing, and how much of it is manual versus automated prompt suites?
  • Which failure classes are in scope, and how do they map to the OWASP Top 10 for LLM Applications?
  • How will you avoid triggering real consequential actions, such as refunds or emails, during testing?
  • Will we receive reproducible test cases and regression checks we can run ourselves?
  • Is remediation validation included, and within what window?

Sources

  1. 01NIST, Glossary: red teamAccessed
  2. 02OWASP, Top 10 for LLM Applications (2025)Accessed
  3. 03MITRE, ATLASAccessed
  4. 04NIST, AI Risk Management FrameworkAccessed
  5. 05NIST, AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileAccessed

Frequently asked questions

Is AI red teaming the same as testing the model for harmful content?
No. Model providers test their models for harmful output. AI red teaming of your application tests what your deployment allows, such as which tools the assistant can call, whose data it can retrieve and what the surrounding code does with its output. The same model can be safe in one deployment and dangerous in another.
Can automated tools or benchmarks replace an AI red team?
They are a useful input, not a replacement. Automated prompt suites find known patterns quickly, but the serious failures usually depend on how your system is wired, such as a tool running with too much permission or a retrieval boundary enforced in the wrong place. Finding those takes a person reading the architecture.
Do you need our system prompt and tool definitions?
An assessment can run entirely from the outside, but it finds more with the system prompt, the tool definitions, the retrieval configuration and the identity each component runs as. Many of the important questions are permission questions, and permissions are faster to read than to infer.
Can an AI assessment run against production?
Yes, under the same conditions as any production testing, with an agreed window, intensity, escalation contacts and stop conditions. Where an agent can take consequential actions, such as refunds or emails, the rules of engagement say in advance which tools may fire and which are stubbed.
How does this relate to a normal penetration test?
It is the same discipline applied to a new kind of entry point. Many AI findings end in conventional vulnerabilities, such as broken authorization or injection into a downstream system, reached through the model. Where an application has both, it often makes sense to test them in the same engagement.

Written by

Kenneth Brown

Published by Red Cell.

Related services: AI red teaming, Web application and API testing, Help me define the scope

Tell us what you need to test.

Send the systems, the timing and the constraints. You get a scoped proposal, not a sales sequence.