TL;DR
- AI has fundamentally changed red teaming by shifting the focus from predictable infrastructure weaknesses to testing the behavior and decision-making of probabilistic AI systems.
- Red teams must expand their expertise with AI, machine learning, and governance knowledge to effectively identify risks such as prompt injection, data poisoning, and shadow AI.
- Map your organization's AI attack surface to understand where AI is deployed, what data it can access, and which systems or tools it can influence before testing begins.
- Replace traditional attack simulations with AI-specific scenarios that test model manipulation, guardrail bypasses, sensitive data exposure, and abuse of AI agents and integrations.
- Treat AI red teaming as a continuous process, embedding testing into development and governance workflows as models, prompts, and connected tools evolve over time.
Red teaming has been a cornerstone of cybersecurity for decades now. Although its tactics have changed with the times, the basic concept has remained the same: find system vulnerabilities and simulate real-world attacks to improve an organization’s security posture, make its incident response more effective, and promote a greater awareness of threats.
But despite its adaptability, red teaming is now undergoing one of the most significant disruptions in its history. The reason, as in so many other industries, is AI.
AI has not only introduced new modes of working and new methods of attack, it has created an entirely new threat landscape. Whereas traditional security threats were largely centered around weaknesses in infrastructure, AI threats almost entirely revolve around behaviors — both of users and the models they’re interacting with. As a result, many of the assumptions that have long guided red teaming must now be thrown out.
Red teaming is far from over though. While the landscape may have transformed, risks and vulnerabilities are still out there. The organizations that can adapt their red teams will be the ones with the best security in our new AI era.
Why AI red teaming is different
Understanding how red teaming needs to change begins with grasping just how much AI has changed everything.
Consider traditional applications, such as software programs. These work in what’s called a “deterministic” fashion: if you make a specific input into a specific function, it will always produce the same output. In other words, put X into Y and you’ll always get Z. This kind of stability makes it easy to build red teaming strategies around. Security teams can trace predictable execution paths, reproduce bugs, identify root causes, and develop targeted fixes.
Unfortunately, AI systems don’t work this way.
Instead of being built on top of code that produces repeatable patterns, AI models are composed of billions of parameters that they learn during their training. When a user gives the model a set of instructions, the model then uses these parameters to make predictions about the most accurate response. These predictions can change based on multiple factors, such as the context of the user’s instructions, the datasets it’s drawing from, or the model’s configuration. The “probabilistic” nature of these systems means that even identical inputs can result in different outputs.
This “deterministic” vs. “probabilistic” distinction is what creates an entirely new set of security concerns. Whereas traditional red teaming might focus on breaking into a system using weak or stolen credentials, unpatched software, or network misconfigurations, an AI red team must now consider how the AI itself can be convinced into making a bad decision or revealing sensitive information.
How red teams need to adapt
An entire class of strategies has emerged to manipulate AI. Prominent attack vectors include prompt injection (using malicious inputs to make an AI model act against its own rules) and data poisoning (feeding contaminated data to AI models). Meanwhile, additional threats like shadow AI, or the unauthorized usage of AI tools and systems within an organization, can create further vulnerabilities for bad actors to exploit.
But understanding how all this works is just one way red teams need to adapt. To fully evolve for the modern AI era, it’s necessary to take a more comprehensive approach. Consider these steps to start building up your AI red team.
1. Add domain experts
Traditional cybersecurity skill sets like offensive security, network architecture, and identity management have long ruled over red teams. And while these skills are still important, AI now requires new disciplines that many security teams may lack. These include an ability to understand how language models function and how agent workflows are constructed. A foundation in governance tactics can also be a fundamental skill on AI red teams.
Some organizations are addressing these gaps by bringing machine learning engineers, AI architects, and domain specialists into red team exercises. Others are investing in training existing personnel so they can develop a working knowledge of AI systems and their failure modes. Whatever the case, every red team member doesn’t necessarily need to become an AI expert overnight. Instead, the team as a whole should understand enough about AI systems to be able to test them realistically.
2. Understand and map out the attack surface
AI is rarely limited to just a single tool at most organizations. Instead, it often exists across different departments and is embedded across multiple layers in the larger environment. For example, some employees may be using publicly accessible chatbots. Others may have internal copilots that they’ve connected to their own datasets and workstreams. The SOC may have specialized AI-powered security tools, while the sales team may have their own custom agents helping them interact with customers. And so on.
Each of these AI instances introduces a potential door for attack. Because of this, before any testing can begin, security teams need to know exactly where AI is being used, what systems it can access, what data it can retrieve, and what actions it can take. The map that emerges from this is the AI attack surface — and it’s often larger than many organizations realize. The reason? AI adoption is happening too quickly these days for most security teams to keep up. But if red teams don’t know where AI exists in their network, they won’t be able to accurately test the risks.
3. Shift to scenario-based testing
Knowing how AI systems work and where they live is just one aspect of successful AI red teaming. Another is knowing how to apply this knowledge in order to reproduce realistic AI threats.
This requires a different approach than traditional red teaming. Instead of asking how to get into a particular environment, teams should focus on creating scenarios that reflect how attackers actually target today’s AI systems. Teams should ask:
- Can we manipulate the model's behavior?
- Can we bypass its guardrails?
- Can we extract sensitive information?
- Can we influence the decisions it makes?
- Can we hijack an agent's goals?
- Can we abuse connected tools or integrations?
Exploring these different scenarios should involve a mix of manual and automated approaches. They should also be practiced within a controlled environment with isolated models so that the team is free to try out even the most aggressive attacks. The goal is to make the adversarial conditions as realistic as possible so that every type of threat and every potential vulnerability can be anticipated and planned for in advance.
4. Turn it into a continuous practice
Finally, AI red teams will need to get used to constantly and consistently putting their systems to the test. If there’s one thing we can predict about the next few years of AI security, it’s that the teams involved with it are going to remain busy.
Unlike traditional red teaming, which can be scheduled and practiced on a periodic basis since infrastructure changes happen relatively slowly, AI systems are evolving constantly. Models get updated, prompts change daily, new tools get connected, and agents gain fresh capabilities. Meanwhile, attackers are always searching out new ways to use these changes to their advantage, which is why AI red teaming has to be thought of as an ongoing exercise rather than a one-time exercise.
One of the best ways to do this — and one that the NIST AI Risk Management Framework directly recommends — is by employing the AI security lifecycle, a strategy that embeds AI testing directly into development, deployment, and governance workflows. However, regardless of how it’s done, the organizations that will see the most success will be the ones that continuously test their AI systems as they change.
Red teams need to evolve alongside AI
Red teams remain one of the most effective tools organizations have for understanding risk. But to remain effective, they need to evolve alongside the technologies they’re testing. For AI, this not only means learning how this technology works. It means becoming familiar with an entirely new set of risks and the rules they operate under.
While AI may not be replacing traditional cybersecurity, it is transforming it. The red teams that start building up the capabilities they need for this new era will be better prepared for the next-generation of AI-driven threats — and ready to reap the rewards of a good defense.
Get your red team ahead. Check out our AI for Red Teams course to learn how to pressure test across prompts, retrieval, tools, supply chains, and more.






