Models
July 15, 2026

OpenAI introduces GPT-Red for automated AI safety testing

OpenAI has introduced GPT-Red, an internal automated red-teaming model that uses self-play to discover vulnerabilities and strengthen AI systems against prompt injection and other security threats.

OpenAI has unveiled GPT-Red, its most advanced internal automated red-teaming model, designed to improve AI safety through self-play reinforcement learning.

GPT-Red continuously attempts to exploit vulnerabilities in defender models, particularly prompt injection attacks, while the defenders learn to resist them, creating an automated self-improvement loop for safety.

OpenAI says GPT-Red generalizes beyond its training scenarios, achieving an 84% success rate on a held-out prompt injection benchmark compared with 13% for human red-teamers. The model has already been used to strengthen GPT-5.6's defenses and is intended to help scale AI safety testing as frontier models become increasingly capable.

#
OpenAI

Read Our Content

See All Blogs
Gen AI

Top Anthropic consulting partners for Claude AI development in 2026

Deveshi Dabbawala

August 4, 2026
Read more
AI safety

Enterprise AI security: How GoML builds prompt injection-resistant applications

Paushigaa S

July 21, 2026
Read more