Every digital service can be misused. Whether it’s gamers Pokémon Go-ing their way into a minefield, Nazis being Nazis on a chess app, or Russia sowing division via fake events, the list is endless.
Consumer AI tools may be particularly prone to misuse. In recent years they have been leveraged for wide-scale impersonation, influence operations, financial scams, industrialized harassment and more.
This has driven a renewed attention to the practice of red teaming. Originally used by the US military for scenario mapping (the “red team” was the USSR), the term was adopted by cybersecurity professionals to describe the exercise of looking for flaws in a network by imitating the behavior of adversaries.
The practice has more recently become associated with the systematic probing of generative AI tools to expose unwanted or unexpected behaviors. All major AI labs do this type of work.
Even though red teaming was probably most buzzy within industry in 2023, when a massive exercise was held at DEF CON, Google searches for the term have continued to grow. That may have something to do with regulatory requirements: The EU’s AI Act mandates adversarial testing for general-purpose models with systemic risk. The US did something similar with an executive order issued by the Biden White House mandating disclosure of red teaming results (the Trump administration rolled this back but is not unworried about adversarial misuse).
Red teaming is not a panacea. But it is an important component of building, deploying, and auditing safe AI tools.
What this guide is (and is not)
I've spent several years working on digital safety and on red teaming specifically. I built the Content Adversarial Red Team at Google, taught Red Teaming 101 at Cornell Tech, and ran several adversarial tests for Indicator. This guide sets out how I organize red teaming exercises.
The guide is intended for red teaming beginners. You might find value in it if you’re an analyst, journalist, or researcher who occasionally probes digital tools for possible safety flaws. While red teaming remains a creative exercise, this guide provides a more repeatable process to give you structure in your investigations. It is organized in six sections:
Define your harms
Build testing personas
Create a prompt library
Go exploring
Turn your hits into failure archetypes
Run an optimized attack
The guide also features some examples of red teaming from my work as well as several resources to go deeper.
Two warnings before we dive in.
First, this is a guide to red teaming consumer AI tools for safety, not security or privacy. The fields clearly overlap: weak infosec can be exploited to extract user data that can then be used to harass a vulnerable individual. But security testing requires different skills and tools from safety testing. For safety testing, you’ll test whether an AI tool can be used to cause harm such as assisting with a scam, generating a nonconsensual nude, or systematically pushing misinformation.
Second and most important: Nothing in this guide should encourage you to create illegal material and/or cause real-life harm. For example, at no point should you actually generate a nonconsensual nude of an identifiable individual (that is a crime in some jurisdictions). Instead, you should use synthetic images or images of objects to see whether the tool you are testing has that capability.
Make sure you are working on devices and networks where you can perform the tests you have in mind. Take brief but clear notes documenting what you are doing throughout the process so you can reconstruct your actions at a later stage. Minimize the reach of your tests by restricting visibility to you alone or, where this is impossible and the content is not illegal, by deleting it immediately after posting. Finally, report anything problematic that you find through appropriate channels. (You should also be aware that if you are red teaming a tool without the developer’s knowledge, you may be kicked off the service.)
This guide will be routinely updated. If you have any recommendations or requests please email [email protected].
Below the paywall: 3,827 words, 29 links to useful resources, 4 charts to guide your testing, and an example of a red teaming persona.
Join today to read the rest
"Alexios and Craig have built something exceptional with The Indicator." - Ruben Gomez, Researcher and former Trust & Safety worker at Reddit and Twitter/X
UpgradeWith a membership, you get:
- Everything we publish
- All of our workshops
- Our eternal gratitude


