What Is AI Red Teaming? Prompt Injection, Jailbreaks and Agentic Abuse
Kenneth Brown
Definition · Published · 4 min read
The short definition
The National Institute of Standards and Technology (NIST) defines a red team as "a group of people authorized and organized to emulate a potential adversary's attack or exploitation capabilities against an enterprise's security posture."
AI red teaming applies that to systems built on a large language model (LLM): assistants, copilots, retrieval applications and agents. An authorized operator works out what the system can be made to do, what it can reach when it does, and what happens downstream when its output is trusted, then shows the evidence.
Test the system, not just the model
A deployed assistant is rarely just a model. It is an application with a system prompt, an identity, a set of tools it can call, a store it can read from, and a user on the other end whose input becomes part of its instructions.
That distinction decides what an assessment is worth. Model providers already test their models for harmful output. What they cannot test is your deployment: the service account your agent runs as, the documents your retrieval layer can surface for a given user, or the internal application programming interface (API) your assistant can call. Those are where the consequential failures live, and they are specific to how your system is wired.
What an AI red team looks for
The Open Worldwide Application Security Project (OWASP) publishes a Top 10 for LLM Applications, which is useful shared vocabulary. The failures an assessment works through map onto it closely:
Prompt injection
A model cannot reliably tell an instruction from data. Anything the application feeds it arrives with the same standing as its system prompt unless the architecture separates them. Direct prompt injection comes from the user typing into the assistant. Indirect prompt injection is more dangerous: the instruction hides in content the system reads on someone's behalf, such as a support ticket, a shared document, an email or a web page, and fires when an unsuspecting user asks the assistant about it.
Jailbreaks and safeguard bypass
A jailbreak is a request crafted to get around the rules a system is meant to follow. For a business application, the question is not whether the model can be coaxed into saying something off-script, but which of the application's safeguards are enforced in code and which exist only as wording in a prompt. Safeguards that live only in the prompt are advisory, and an assessment should tell you which of yours are which.
Agentic abuse
Once an assistant can call tools, a content problem becomes an action problem. An agent that can send email, issue refunds, change tickets or write to a system is an authenticated user of everything it can reach. Agentic abuse covers getting an agent to take those actions for the wrong person: calling a tool with another user's identifier, chaining tools into something no single tool would allow, or satisfying a human approval step with a summary the agent wrote itself.
How an AI red team engagement runs
An AI assessment follows the same authorization discipline as any offensive engagement: written scope, rules of engagement and signed authorization come first. The work itself then runs in six stages:
- Inventory. Map what is actually deployed: entry points, system prompts, retrieval sources, tools, the identity each runs as, and where output goes.
- Threat model. Decide which failures would matter for this system and this business, and which trust boundaries carry the weight.
- Agreed test cases. Turn the threat model into cases you sign off before execution, including anything that must be left alone.
- Controlled evaluation. Run the cases against the agreed environment, within the intensity and stop conditions in the rules of engagement.
- Findings. Report what reproduced, what it reached, and the control that removes it, with evidence for each case.
- Remediation validation. Re-run the affected cases after the fix, and hand over the ones worth keeping as regression checks.
What you should receive
- A threat model of the deployed architecture that your team keeps using after the engagement.
- Reproducible evaluation cases your engineers can run themselves, with the expected safe behaviour stated next to the observed one.
- Evidence for each finding: transcripts, tool invocations and responses.
- Control recommendations mapped to the failure each one closes, not to a generic checklist.
- Regression checks worth running continuously, in a form you can add to your own pipeline.
If the findings need to feed a governance process, the NIST AI Risk Management Framework and its Generative Artificial Intelligence Profile are the usual reference points.
Questions to ask an AI red teaming provider
Sources
Frequently asked questions
Is AI red teaming the same as testing the model for harmful content?
Can automated tools or benchmarks replace an AI red team?
Do you need our system prompt and tool definitions?
Can an AI assessment run against production?
How does this relate to a normal penetration test?
Written by
Kenneth Brown
Published by Red Cell.
Related services: AI red teaming, Web application and API testing, Help me define the scope
Related articles
- DefinitionHow to Read a Penetration Testing ReportWhat each part of a penetration testing report tells you, how to read severity against real reachability, and what to check before you sign off.
- DefinitionHow to Write a Penetration Testing RFP (with Checklist)What to put in a penetration testing request for proposal: scope, constraints, deliverables, retest terms and evaluation criteria, with a checklist.
- DefinitionPenetration Testing Pricing: What Drives Cost in 2026What sets the cost of a penetration test: scope, complexity, depth, environment, compliance evidence, retesting and timing, and how to compare quotes.