Start here. This is the direct spoken answer to practice first.
Overview
AI red teaming tests the deployed system and its authority, not just whether a base model says something undesirable.
I define the assets, attacker roles, allowed test environment, and impact criteria before generating attacks. Tests cover direct and indirect injection, sensitive disclosure, unsafe output handling, poisoning, excessive agency, tenant boundaries, resource abuse, and misleading behavior relevant to the product. Side effects use isolated accounts, fake destinations, or reversible operations.