Skip to content
LIVE
The Executives BriefThe Executives BriefBeta

OpenAI built GPT-Red to out-hack models, and it trained GPT-5.6 to resist prompt attacks

The sparring super-hacker automates red-teaming, finds fresh prompt injection modes, and helps explain why GPT-5.6 is harder to break.

ByYousef Al-ZahraniTechnology Correspondent, The Executives Brief
·4 min read
OpenAI built GPT-Red to out-hack models, and it trained GPT-5.6 to resist prompt attacks
Executive summary

OpenAI built an LLM super-hacker called GPT-Red and used it in a self-play training loop to improve the safety of its flagship model, GPT-5.6. For decision-makers, the shift matters because it points to automated, scalable cyber stress testing that can keep pace as AI agents widen the attack surface.

OpenAI has built an LLM super-hacker called GPT-Red, and the company says training GPT-5.6 against it made that release its most robust yet. Last week, OpenAI released the latest version of its flagship LLM, GPT-5.6, and the headline number is the security win: OpenAI claims fewer than 23% of the strongest GPT-Red attacks that it tried actually worked against GPT-5.6, down from more than 90% working against the earlier GPT-5 (released in August last year).

This is not “safety theater.” GPT-Red automates a type of safety evaluation called red-teaming, which is typically done by human testers. The goal in red-teaming is to find as many different ways as possible to break or hijack a system, then patch weak spots before release. OpenAI’s argument is that as LLMs become more complex and get used in more places, especially as agents that can interact with computer files, websites, and third-party code, it becomes impossible for teams of people to keep up with every attack mode that could appear. “The risk surface grows and the blast radius also grows,” says Nikhil Kandpal, a research scientist at OpenAI who co-created GPT-Red.

GPT-Red is designed to future-proof OpenAI’s safety testing process. The researchers say GPT-Red was built so OpenAI can “discover new modes of attack” as models get more capable, rather than waiting for humans to find yesterday’s failure cases. Dylan Hunn, also a research scientist at OpenAI and a co-creator of GPT-Red, says the system can be designed in advance to keep finding new kinds of attack. OpenAI further claims GPT-Red has already come up with new types of attack that had not been seen before.

The focus is especially sharp on prompt injection. In a prompt injection attack, a hacker slips an LLM instructions in a way that makes the model do things developers or users do not want, like copying confidential information, sabotaging a company’s code base, or generating embarrassing or harmful output. The tricky part is that in theory, those instructions can be hidden in any text the LLM might encounter, including code or content on a website.

To build GPT-Red, OpenAI took an LLM that had not been trained as a hacker and set it up in a self-play loop with several other models. One side tried to attack the other models; those models tried to defend themselves. Over many rounds, GPT-Red became better and better at attacking, while the targets became better at resisting. The training happened in a “dojo” that OpenAI designed to mimic scenarios where LLMs are actually deployed in the real world, including browsing the web, reading emails or calendar apps, and editing code. When GPT-Red discovered a new attack, it would explore multiple versions to find the most efficient one for specific scenarios.

OpenAI claims GPT-Red found a prompt injection attack type the researchers had not seen before, called a fake chain of thought. A chain of thought is a diary-like trace where an LLM makes notes to itself and keeps track of partial results while solving problems. The researchers say GPT-Red learned a method to insert a fake entry into another model’s chain of thought so the model would act on spoofed information. One example provided by Chris Choquette-Choo, another research scientist on the team, frames the deception: it is like telling a model that “1+1=3” after it has supposedly verified “1+1=2,” and then watching it treat the spoofed result as obviously correct.

There is also a broader security measurement here that should matter to regulators and enterprise risk teams. OpenAI tested GPT-Red by rerunning an experiment from 2025 where human red-teamers tried to find weaknesses in an earlier version of GPT-5. When GPT-Red was given the same task, it was more successful at finding effective attacks than the humans had been. OpenAI also tested GPT-Red against Vendy, a vending machine agent developed by Andon Labs, which assesses how well agents perform real-world tasks. OpenAI says GPT-Red was able to hack Vendy to make it change the prices of items on sale and cancel a customer’s order.

Even with those wins, GPT-Red is not portrayed as magic. OpenAI says it is not great at attacks that involve a back-and-forth conversation between hacker and target, something human attackers would have few problems with. It is also not yet that good at using images, which can be used to pass text to models in prompt injection attacks. And OpenAI emphasizes that GPT-Red supplements human red-teamers, not replaces them. One approach the company is taking is to give GPT-Red an attack that humans already came up with and ask it to find all the variations. Jessica Ji, a senior research analyst who works on AI security at Georgetown University’s Center for Security and Emerging Technology (CSET), says the self-play loop approach looks promising and that human expertise will still be important, including to identify where human testing is most needed.

For executives, the second-order implication is clear: if red-teaming can be automated in a way that discovers new attack modes, the cadence of model iteration changes. That has downstream effects on governance, incident response planning, and how boards think about model deployment risk. It also indirectly informs the conversation regulators are pushing across AI safety: as AI agents expand access and autonomy, the systems that test them need to expand too. OpenAI also says it will not release GPT-Red, and it argues it is not something a copycat can easily replicate by “just go and train a super-attacker using this idea,” because the researchers have been working on it for more than a year, backed by substantial compute resources. In other words: the arms race is becoming less about one-off red-team reports and more about continuous, automated adversarial pressure built into the development lifecycle.

Executive ActionsLocked

This story's Key Insights and Take-aways are locked.

Create a free account to unlock Executive Actions for one credit.

Register to Unlock

Always free for Executives Club members. Join the Club

More in Technology