AI security testing is sold under several names, including AI red teaming, LLM penetration testing and AI security assessment, and the same name can mean very different scopes. Before comparing firms, decide what you need tested: the model's behavior, the application and agents built around it, the training and deployment pipeline, or an adversary simulation where AI is one path among many. Then decide whether you want a point-in-time engagement or continuing testing as the system changes.
When proposals arrive, compare them on four things: the attack classes in scope, whether tools and agents are exercised as deployed or only described, what share of the testing is automated, and whether fixes are retested. Our explainer on what AI red teaming is sets out the vocabulary, and the broader offensive security firms shortlist covers the same firms across their full service range.
Scoped engagements under a written statement of work, rules of engagement and signed authorization. The assessment is drawn against the architecture you have deployed, with evaluation cases you sign off before execution and critical findings raised through a named escalation contact when found.
Scope
The application, agent, permissions, connected tools and data flows around a deployed model: direct and indirect prompt injection, guardrail reliability, sensitive information exposure, tool misuse, cross-user and cross-tenant isolation, persistent memory, unsafe output handling and approval bypass.
Buyer fit
Teams whose assistant, copilot, retrieval application or agent has been given access to real systems and real data, and who want it tested as a deployed system rather than as a model in isolation.
What to confirm
— Which tools may fire during testing and which are stubbed
— How much configuration you can share, such as the system prompt, tool definitions and retrieval setup
— Which evaluation cases you want handed over as regression checks
Red Cell takeOur own entry: a San Diego firm that tests the system around the model, delivers reproducible evaluation cases, and includes a retest of the agreed findings.
Publicly describes hands-on AI and LLM assessments tailored to the client's environment, delivered as separate methodologies or combined: application penetration testing of LLM endpoints, cloud penetration testing of the AI stack, and AI-focused red team, purple team and tabletop exercises.
Scope
Publicly lists prompt injection, model extraction, data poisoning, resource exhaustion, supply chain compromise, trust boundaries, isolation patterns and secrets management, plus jailbreaks, context leak chains and denial-of-wallet risks. Its page references the OWASP lists for LLM, machine learning and agentic applications, and MITRE ATLAS.
Buyer fit
Organizations that want AI testing spanning the application, the cloud infrastructure beneath it and a full adversary operation against the model lifecycle.
What to confirm
— Which of the three methodologies a proposal includes
— Whether red team scenarios touching training data or pipelines are run against production or a copy
— Whether retesting of fixed findings is included
Red Cell takeA fit for teams that want application, cloud and red team coverage of an AI system from one firm.
Publicly describes an AI Red Teaming service built on realistic attack simulations, followed by prioritized recommendations. It also describes its testers teaching classes on attacking LLM applications and releasing open-source training tools.
Scope
Publicly lists prompt injection, evasion, model extraction, model inversion, membership inference, data poisoning and unintended data leakage, across AI models from concept to deployment, training infrastructure and data pipelines, and application programming interfaces (APIs), user interfaces and input and output channels. It cites OWASP guidance for machine learning and LLMs, MITRE ATLAS, and the AI Risk Management Framework from the National Institute of Standards and Technology (NIST).
Buyer fit
Specific industries are not stated on its AI assessment page.
What to confirm
— Whether agents and tool calling inside your own application are in scope
— Whether testing covers the model, the application around it, or both
— Whether retesting of fixed findings is included
Red Cell takeWorth shortlisting for teams that want model-level attack classes tested alongside the interfaces that reach them.
Publicly describes LLM and AI application penetration testing with validated findings and proof of exploitation, delivered with a report, fix recommendations, a shared chat channel, free re-testing, an attestation letter and a technical presentation.
Scope
Publicly lists prompt injection and jailbreak simulation, context manipulation and instruction override, model poisoning and adversarial attacks, data extraction, training data exposure, sensitive information leakage, hallucination and manipulation risks, and model governance and access controls. It cites OWASP guidance for LLMs, the NIST AI Risk Management Framework and MITRE ATLAS.
Buyer fit
Specific industries are not stated on its AI testing page; it lists compliance frameworks its reports can support.
What to confirm
— Whether agents, tool use and retrieval in your application are exercised
— Who performs the testing and their AI-specific background
— How long free re-testing remains available
Red Cell takeWorth a look for buyers who want a packaged LLM application test with re-testing and an attestation letter.
Publicly describes AI and machine learning specialists who deliver tailored assessments and strategic recommendations, within a broader Securing AI offering that also covers governance and adoption. It states it has dedicated AI testing infrastructure for testing models, applications and ecosystems.
Scope
Publicly lists AI and machine learning threat modelling, evaluation of bias and toxicity to prevent misuse, identification of security risks in AI systems and models against industry standards, and support for a secure development lifecycle.
Buyer fit
Large organizations that want AI testing alongside AI governance and regulatory guidance from the same provider.
What to confirm
— Which LLM attack classes, such as prompt injection or tool misuse, are in scope
— Whether testing covers deployed agents and integrations or the model alone
— Whether retesting is included
Red Cell takeSuited to organizations that want AI security testing inside a wider governance and assurance program.
Publicly describes AI and LLM testing inside its penetration testing as a service (PTaaS) offering, as custom evaluations, LLM web application tests, benchmarking and jailbreaking, and continuous AI penetration testing with human validation and findings delivered through its platform. It also offers validation of findings produced by other AI security tools.
Scope
Publicly lists LLM-specific vulnerabilities in web applications, jailbreaking, data leakage, adversarial attacks, content moderation, response bias and data drift, direct prompt injection, and custom evaluations covering training data review, model extraction, inference, inversion and evasion. It states testing works across model providers and frameworks.
Buyer fit
Organizations that fine-tune, build or embed models and want AI testing run as a continuing program.
What to confirm
— Which of its AI services a proposal covers
— How work divides between platform and human testers, and who signs off findings
— Whether agents and tool use are tested, beyond the web application layer
Red Cell takeA strong fit for organizations that want AI testing folded into an existing platform-managed testing program.
Publicly describes AI penetration testing and LLM red team assessments in stages: discovery and threat modeling, attack surface mapping, manual and automated adversarial prompt campaigns, data and pipeline testing, application and infrastructure testing, and impact analysis with follow-up verification.
Scope
Publicly lists prompt injection and jailbreak resistance, retrieval pipeline evaluation, training and fine-tune poisoning, model backdoor checks, model extraction, inversion and membership inference, function-calling and tool misuse, insecure output handling, and supply chain scenarios. It cites MITRE ATLAS, OWASP guidance for LLMs and the NIST AI Risk Management Framework.
Buyer fit
Specific industries are not stated on its AI testing page.
What to confirm
— Which stages are manual and which automated
— Whether AI testing is sold standalone or within its platform subscription
— What follow-up verification covers and for how long
Red Cell takeA fit for teams that want a staged methodology running from threat model to pipeline and infrastructure testing.
Publicly describes AI red teaming that brings AI tooling and tradecraft to simulate AI-enabled adversaries, and AI security assessments across the system lifecycle, alongside application security assessments that include AI systems and models.
Scope
Publicly lists evaluations across design and development, infrastructure assessments from the perspective of a compromised user, and offensive testing of production AI systems, including model behavior and agentic workflows, with multi-role and multi-tenant testing in application assessments.
Buyer fit
Publicly lists financial services, public sector and healthcare among its industries.
What to confirm
— Whether a proposal is an AI red team exercise, an AI security assessment or an application assessment
— Whether agents are tested with live tool access
— What the report contains and whether retesting is included
Red Cell takeA natural shortlist entry where AI systems sit inside a wider identity and adversary-simulation concern.
Publicly describes end-to-end AI and machine learning assessments staffed across machine learning, application security, systems and cryptography, following a staged flow from system boundary to root cause, with a live readout, patch re-testing and open publication of generalizable findings.
Scope
Publicly lists training data and pipeline review, model artifacts and serialization, inference hardware, deployed agent loops, prompt injection, data exfiltration, inference manipulation, capability misuse, threat modeling and capability benchmarking against expert baselines. Deliverables include static-analysis rules and evaluation harnesses.
Buyer fit
Teams building or serving models who want pipeline, hardware and code-level review as well as testing of the deployed model and agent.
What to confirm
— The publication policy for your engagement and whether you can opt out
— Which harnesses and rules are handed over
— How long patch re-testing is available
Red Cell takeThe specialist choice for deep review of the machine learning pipeline and serving stack, not only the chat interface.
Publicly describes a practitioner-led, adversarial approach to AI, with engagements scoped to the client's environment: AI security assessments, AI-focused red team testing, AI governance advisory and incident response for AI-related incidents.
Scope
Publicly lists structured evaluation of AI systems, pipelines and integrations, and testing through prompt injection, model manipulation, data extraction attempts and abuse-case simulation, plus review of internal AI use and secure adoption programs.
Buyer fit
Organizations that want AI testing, governance and incident response from the same consulting partner.
What to confirm
— Whether agents and tool calling are exercised, as well as the model interface
— Which frameworks the assessment maps findings to
— Whether retesting is included and within what window
Red Cell takeA good fit for teams that want AI testing tied to governance work and an incident response capability.
What the firms have in common
Most firms on this list name the same core attack classes: prompt injection, jailbreaks, data leakage and extraction, and poisoning. They differ on three points that matter more to a buyer:
Where the boundary sits. Some firms test the model and its interfaces; others test the application, the agents and tools, the cloud account and the training pipeline. Confirm that the system you actually run is the thing being tested.
How much is automated. Several firms combine automated prompt campaigns or platforms with human testers. Ask which findings come from which, and who validates them before they reach the report.
What happens after the report. Retest terms, handed-over test cases and regression harnesses vary widely. They decide whether the engagement keeps paying off after the fix.
If you want to see how Red Cell scopes this work, the AI red teaming page sets out the attack classes and deliverables, our approach describes the engagement documents, and pricing explains how a quote is built. When you are ready, request an engagement.
Sources
Every statement about a firm other than Red Cell comes from the pages below, accessed on the evidence cutoff date.
Each firm publicly offers AI or LLM security testing on its own website and describes enough of that work to set out its attack classes, frameworks and delivery. The list is a starting shortlist, not a complete survey of the market.
Why is Red Cell first?
Red Cell publishes this article and places itself first, as the publisher disclosure at the top states. That placement is the publisher's point of view. The remaining nine firms are listed alphabetically and are not scored or ranked against each other.
What is the difference between testing a model and testing an AI application?
Model testing looks at how the model itself behaves under adversarial input, such as jailbreaks, extraction or poisoning. Application testing looks at the system around the model, including its tools, permissions, retrieval sources and the identities it runs as. Several firms below offer both, so confirm which one a proposal actually covers.
Which frameworks do these firms reference?
The most commonly cited on the firms' own pages are the OWASP guidance for LLM applications, MITRE ATLAS and the NIST AI Risk Management Framework. A framework is a shared vocabulary and a completeness check; it does not replace a threat model of your own system.
How current is this information?
Every statement about another firm comes from that firm's own website as of the evidence cutoff date shown at the top. AI security services change quickly, so confirm anything that matters to your decision directly with the firm.