AI red teaming
AI red teaming is adversarial testing of an AI system: deliberately trying to make it produce harmful, leaked or prohibited output before someone else does. For providers of general-purpose models with systemic risk it is a regulatory expectation under the EU AI Act. For a company deploying a document assistant it usually is not — and the testing that would help is a different, cheaper exercise.
What does red teaming actually involve?
Structured attempts to break the system under realistic conditions: prompt injection through documents the model reads, extraction of instructions or data it was told to keep, jailbreaks around content restrictions, and probing for outputs that are wrong in ways that matter rather than wrong in general.
Done properly it is adversarial and iterative — findings change the next round. A one-off checklist run is not red teaming; it is a test pass with a more exciting name, and the difference shows in what it finds.
Who is actually expected to do it?
Under the EU AI Act, providers of general-purpose AI models with systemic risk carry obligations that include adversarial testing and model evaluation. That is a small set of organisations, and if you are in it you know.
A deployer of a retrieval assistant is not in that set. Which does not mean no testing is warranted — it means the useful testing is different, and paying for a systemic-risk exercise would be paying for the wrong thing.
What testing should a deploying company do instead?
Three exercises, all cheaper. Prompt injection through your own documents: place an instruction inside a file the system ingests and see whether it obeys it — this is the realistic attack against a document assistant and it is routinely unhandled. Boundary testing: ask questions whose answers are not in the corpus and confirm the system says so. And permission testing: verify a user cannot retrieve passages from documents their account should not reach.
The third finds more real problems than the other two combined, and it is the one most often skipped, because it requires access control to have been designed rather than assumed.
What it is not to be confused with
Penetration testing
Pen testing targets the infrastructure — the API, the network, the authentication. Red teaming targets the model behaviour. Both are needed and they find different classes of problem; a supplier offering one as the other is worth questioning.
Benchmarking
Benchmarks measure average capability on shared tasks. Red teaming looks for the worst case on your task. A system can score well on both public benchmarks and badly on the one question your regulator would ask.
Frequently asked
What is AI red teaming?+
Adversarial testing of an AI system — deliberately attempting to induce harmful, leaked, unlawful or simply wrong output under realistic conditions, so that the failure modes are found internally rather than by a user, a journalist or a supervisory authority.
Is red teaming required by the EU AI Act?+
For providers of general-purpose AI models with systemic risk, adversarial testing forms part of the obligations that have applied since 2 August 2025. For deployers of ordinary AI systems there is no such requirement. The Act does require human oversight and risk management for high-risk systems, which is related but not the same exercise.
Do we need a certification for this?+
No certification is prescribed, and the market in red-teaming certificates has grown considerably faster than the regulatory basis for it. What is worth having is a record of what was tested, what was found and what changed as a result.
Can we do this ourselves?+
The three exercises that matter most for a document assistant — injection through ingested files, boundary behaviour, and permission isolation — are within reach of an internal team in a day or two. Specialised external testing earns its cost at the point where you are a provider rather than a deployer.
Related terms
The test worth running first
Put a line of instruction text inside a document your system ingests and ask a normal question. If the system follows the instruction in the file, you have found the issue that matters before anyone else does.
Request a test