OpenAI Disclosed Automated Security Testing Results
The firm revealed that its GPT-Red system outperformed human testers in identifying indirect prompt-injection vulnerabilities.
Updated on Sept. 26, 2026 in Artificial Intelligence

Live Poll
Do you trust current automated security testing to prevent malicious attacks on AI tools you use?
OpenAI recently shared data on its GPT-Red security testing system, which utilizes self-play reinforcement learning to uncover model weaknesses. The automated tool successfully executed indirect prompt-injection attacks in 84% of scenarios, significantly outpacing human red-teamers who achieved a 13% success rate.
Why it matters
Identifying these vulnerabilities allows developers to train production models to be more resilient against malicious prompts. By stress-testing systems with automated tools, researchers can better secure AI assistants against complex, self-replicating attacks.
GPT-Red successfully exploited indirect prompt injections at a rate of 84%, while the Virtual Donkey defense system demonstrated a true-positive rate of 1.0 and a false-positive rate of 0.015.
The players
OpenAI
An artificial intelligence research organization that develops large language models and advanced safety testing systems.
The details
The GPT-Red system managed to manipulate a vending machine price to $0.50, illustrating how automated prompts can bypass standard security parameters. Meanwhile, the newer GPT-5.6 Sol model showed improved reliability, failing on only 0.05% of direct injections compared to its predecessors.
Timeline
Benchmarks for the best production models were established in May 2026.
OpenAI officially disclosed the GPT-Red findings on September 26, 2026.
The Tech Race
The GPT-Red findings follow a pattern established by the Morris-II research project regarding self-replicating prompts between email assistants. This development highlights the shift toward using automated, adversarial AI agents to harden systems against increasingly sophisticated prompt-injection threats.
Users of AI assistants may experience more robust security protections as developers implement findings from automated red-teaming. These advancements help ensure that AI tools are less susceptible to unauthorized command changes or malicious price manipulations.
The takeaway
Automated testing systems like GPT-Red represent a critical evolution in how developers preemptively secure large language models. Proactive adversarial training is essential for mitigating the risks posed by indirect prompt injections as these models become more deeply integrated into daily workflows.
Further reading
For more on evolving security standards, explore the Artificial Intelligence section.
Source note: This article includes information reported by TokenPost.
Live Poll
Do you trust current automated security testing to prevent malicious attacks on AI tools you use?










