OpenAI Disclosed Automated Security Testing Results

The firm revealed that its GPT-Red system outperformed human testers in identifying indirect prompt-injection vulnerabilities.

Updated on Sept. 26, 2026 in Artificial Intelligence

Isometric editorial illustration of a steel server lattice with a single glowing fiber-optic conduit, representing automated AI security testing.
OpenAI reported that its GPT-Red automated security system outperformed human testers, successfully identifying indirect prompt-injection vulnerabilities in 84% of scenarios. AI Illustration. Upload story photo >

Live Poll

Do you trust current automated security testing to prevent malicious attacks on AI tools you use?

OpenAI recently shared data on its GPT-Red security testing system, which utilizes self-play reinforcement learning to uncover model weaknesses. The automated tool successfully executed indirect prompt-injection attacks in 84% of scenarios, significantly outpacing human red-teamers who achieved a 13% success rate.

Why it matters

Identifying these vulnerabilities allows developers to train production models to be more resilient against malicious prompts. By stress-testing systems with automated tools, researchers can better secure AI assistants against complex, self-replicating attacks.

GPT-Red successfully exploited indirect prompt injections at a rate of 84%, while the Virtual Donkey defense system demonstrated a true-positive rate of 1.0 and a false-positive rate of 0.015.

The players

OpenAI

An artificial intelligence research organization that develops large language models and advanced safety testing systems.

The details

The GPT-Red system managed to manipulate a vending machine price to $0.50, illustrating how automated prompts can bypass standard security parameters. Meanwhile, the newer GPT-5.6 Sol model showed improved reliability, failing on only 0.05% of direct injections compared to its predecessors.

Timeline

  1. Benchmarks for the best production models were established in May 2026.

  2. OpenAI officially disclosed the GPT-Red findings on September 26, 2026.

The Tech Race

The GPT-Red findings follow a pattern established by the Morris-II research project regarding self-replicating prompts between email assistants. This development highlights the shift toward using automated, adversarial AI agents to harden systems against increasingly sophisticated prompt-injection threats.

Users of AI assistants may experience more robust security protections as developers implement findings from automated red-teaming. These advancements help ensure that AI tools are less susceptible to unauthorized command changes or malicious price manipulations.

The takeaway

Automated testing systems like GPT-Red represent a critical evolution in how developers preemptively secure large language models. Proactive adversarial training is essential for mitigating the risks posed by indirect prompt injections as these models become more deeply integrated into daily workflows.

Further reading

For more on evolving security standards, explore the Artificial Intelligence section.

Source note: This article includes information reported by TokenPost.

Live Poll

Do you trust current automated security testing to prevent malicious attacks on AI tools you use?

OpenAI Disclosed Automated Security Testing Results