Models
July 15, 2026

OpenAI introduces GPT-Red for automated AI safety testing

OpenAI has introduced GPT-Red, an internal automated red-teaming model that uses self-play to discover vulnerabilities and strengthen AI systems against prompt injection and other security threats.

OpenAI has unveiled GPT-Red, its most advanced internal automated red-teaming model, designed to improve AI safety through self-play reinforcement learning.

GPT-Red continuously attempts to exploit vulnerabilities in defender models, particularly prompt injection attacks, while the defenders learn to resist them, creating an automated self-improvement loop for safety.

OpenAI says GPT-Red generalizes beyond its training scenarios, achieving an 84% success rate on a held-out prompt injection benchmark compared with 13% for human red-teamers. The model has already been used to strengthen GPT-5.6's defenses and is intended to help scale AI safety testing as frontier models become increasingly capable.

#
OpenAI

Read Our Content

See All Blogs
AWS

AWS SageMaker: Develop high-efficiency machine learning models

Deveshi Dabbawala

August 26, 2026
Read more
Gen AI

AI agent evaluation for self-improving agent loops

Sarankumar S

August 21, 2026
Read more