OpenAI Built an AI That Attacks Its Own Models to Harden Them
OpenAI says its new internal system, GPT-Red, found prompt-injection exploits far faster than human red-teamers and used them to make GPT-5.6 six times more resistant to attacks.
Read more →