GPT-6 Astra blocks 99.99% of direct injections but fails 8.5% of hidden document attacks
Gray Swan testing reported by The Decoder paints a split safety picture for OpenAI GPT-6 Astra: stronger resistance to obvious prompt attacks alongside vulnerability when instructions hide inside documents the model reads. Coverage states Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections, suggesting meaningful hardening against straightforward injection attempts. In indirect scenarios embedding attacks in readable documents, testers cracked Astra at least once in 8.5 percent of cases, compared with 4.8 percent for Claude Opus 5 under the same evaluation. Autonomous agents processing untrusted files should weigh those indirect failure rates more heavily than direct-block statistics. The packet does not detail Gray Swan test design, attack libraries, or sample sizes, so treat the percentages as comparative signals rather than standalone deployment guarantees.
GPT-6 Astra blocks 99.99% of direct injections but fails 8.5% of hidden document attacks
OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections. Gray Swan testing found Astra was cracked at least once in 8.5 percent of indirect injection scenarios, versus 4.8 percent for Claude Opus 5.
Key takeaway
Astra's near-total direct injection blocking does not carry over to hidden document attacks, where it still fails in a meaningful share of scenarios.
What happened
The Decoder reports that OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections, according to Gray Swan testing cited in the coverage.
When attacks were hidden inside documents the AI reads, Gray Swan testing found Astra was cracked at least once in 8.5 percent of indirect injection scenarios, versus 4.8 percent for Claude Opus 5 in the same evaluation.
Evidence
GPT-6 Astra blocks 99.99 percent of direct prompt injections.
The Decoder · attributed
OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections.
Astra was cracked in 8.5 percent of indirect injection scenarios involving hidden document attacks.
The Decoder · attributed
Gray Swan testing found Astra was cracked at least once in 8.5 percent of indirect injection scenarios, versus 4.8 percent for Claude Opus 5.
Claude Opus 5 had a lower indirect injection failure rate of 4.8 percent in the same testing.
The Decoder · attributed
Gray Swan testing found Astra was cracked at least once in 8.5 percent of indirect injection scenarios, versus 4.8 percent for Claude Opus 5.
GPT-6 Astra hallucinates less than its predecessor.
The Decoder · attributed
OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections.
Why it matters
For autonomous AI agents handling real data, indirect injection failure rates may be a more relevant safety signal than headline direct-injection block rates.
Limits and uncertainties
The packet does not provide Gray Swan test methodology, attack libraries, or sample sizes behind the reported percentages.
Several source excerpts in the packet are truncated and do not include the full article text.
Practical implications
Teams deploying agents that read untrusted documents should evaluate indirect injection resistance, not only direct prompt injection benchmarks.
Model selection for document-handling agents should compare indirect failure rates, where Astra reportedly underperformed Claude Opus 5 in Gray Swan testing.
What to watch
Whether OpenAI publishes its own benchmarks or mitigations for document-borne indirect prompt injections in GPT-6 Astra.
Additional independent indirect injection test results beyond the Gray Swan figures cited by The Decoder.