
OpenAI's GPT-Red finds prompt injections 6x better…
GPT-Red is an internal automated red-teaming model trained with self-play RL. It hit 84% attack success versus 13% for human testers, discovered fake chain-of-thought injections, and helped cut GPT-5.6 Sol failures on the hardest benchmark by 6x.
















