The night before a GPT-4-class model ships, a small team sits with the candidate system and tries to make it fail. One person hunts jailbreaks. Another pushes for private data the system was never supposed to leak. A third builds a chain of polite prompts that ends in a dangerous capability demonstration. These people are on payroll as the adversary, not as ordinary users.
That practice is red teaming. The name comes from military exercises and cybersecurity, where a red team plays attacker against a defending blue team. In AI it means people, and increasingly other models, working hard to force harmful instructions, guardrail bypasses, data leaks, or a capability the lab hoped to keep fenced. What they find gets patched or restricted. What they miss stays invisible until someone else finds it.
Most emerging governance frameworks require or encourage the practice. Labs treat a serious pre-release red team as table stakes. The Foundation wants more of it, done independently rather than only in-house, with findings disclosed rather than buried. Usefulness is not the open question. What a clean result is allowed to mean is.
What red teaming does well
Ordinary testing exercises the cases someone already expected. Adversarial pressure hunts the cases nobody wrote down. That difference is why red teaming finds specific, fixable failures that polite evaluation suites walk past. It also stress-tests safeguards against the kind of determined effort a real misuser would apply, not the cooperative queries of a typical customer.
Findings feed the rest of the safety stack. They sharpen training, inform dangerous capability evaluations, and supply evidence any honest safety case has to face. A model that has survived a serious red team holds up better against known attack patterns than one that has not. That gain is real and worth paying for.
The folk objection, named
The common reply runs like this: if skilled people tried hard to break the system and failed, the system is safe enough to ship. The intuition feels fair. In physical security, surviving a penetration test is often treated as a green light. Why should AI be different?
Because the logic of testing is asymmetric. Finding a flaw proves the flaw exists. Failing to find one proves only that this team, with this time budget and these techniques, did not succeed. A more capable attacker, a novel method, or simply more calendar days can still win. As the space of possible attacks grows faster than any squad can cover, a clean report becomes weaker evidence, not stronger.
Red teaming can show a system is unsafe. A clean run cannot show it is safe.
Two further limits tighten the gap. Against a highly capable model the human red team can be outmatched on speed and breadth. A system that can recognize a test and underperform can pass the exercise while retaining the capability under probe. And red teaming is built for misuse and known failure modes. A model with goals of its own has no reason to volunteer those goals to someone hunting jailbreaks.
Where it belongs
Keep the tool. Expand independent access. Publish results that matter for public risk. Stop treating a quiet red-team week as a certificate that the absence of proof of danger equals proof of safety. That swap of absences is the error running through too much current practice.
Red teaming belongs inside a governance regime that does not need the stress test to be exhaustive: thresholds, external review, and limits that hold when an evaluation is incomplete. The design of that regime is set out in our plan.
Common questions.
What is red teaming in AI?
Red teaming is the practice of deliberately attacking an AI model to uncover its weaknesses before others do. Borrowed from military and cybersecurity exercises, it involves people, and increasingly other AI systems, trying hard to make a model produce harmful content, leak data, be manipulated past its guardrails, or reveal a dangerous capability. Frontier models are red-teamed before release, and the findings are used to fix or fence off the failures discovered.
What is red teaming good at finding?
It surfaces failures that ordinary testing misses, because ordinary testing exercises expected cases while red teaming actively hunts for unexpected ones. It finds specific, fixable problems and stress-tests safeguards against the sort of determined effort a real misuser would apply. Its results also strengthen the wider safety apparatus, informing capability evaluations, improving safety training, and providing evidence for safety arguments.
What are the limits of red teaming?
The central limit is an asymmetry shared by all testing: finding a flaw proves the flaw exists, but failing to find one does not prove none exists. A clean result only shows that this team, with this much time and these techniques, did not break the model, while a more capable attacker or a novel method might. Against a highly capable model the red team may be outmatched, a model that can recognize a test could underperform to pass it, and red teaming addresses misuse far better than it addresses a model with genuinely misaligned goals.
Does passing red teaming mean an AI model is safe?
No. Passing red teaming means known attack techniques applied within a limited time did not succeed, which is useful but far from a safety guarantee. The space of possible attacks is larger than any team can explore, so the absence of a discovered failure is weak evidence that no failure exists, and that evidence weakens as models become more capable. Treating a clean red-team result as a certificate of safety mistakes the absence of proof of danger for proof of safety.