Skip to main content
LLMgram · AI News · 2026-09-04

GPT-6 Astra blocks 99.99% of direct injections but fails 8.5% of hidden document attacks

GPT-6 Astra blocks 99.99% of direct injections but fails 8.5% of hidden document attacks

Gray Swan testing reported by The Decoder paints a split safety picture for OpenAI GPT-6 Astra: stronger resistance to obvious prompt attacks alongside vulnerability when instructions hide inside documents the model reads. Coverage states Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections, suggesting meaningful hardening against straightforward injection attempts. In indirect scenarios embedding attacks in readable documents, testers cracked Astra at least once in 8.5 percent of cases, compared with 4.8 percent for Claude Opus 5 under the same evaluation. Autonomous agents processing untrusted files should weigh those indirect failure rates more heavily than direct-block statistics. The packet does not detail Gray Swan test design, attack libraries, or sample sizes, so treat the percentages as comparative signals rather than standalone deployment guarantees.

Sources

GPT-6 Astra blocks 99.99% of direct injections but fails 8.5% of hidden document attacks

GPT-6 Astra blocks 99.99% of direct injections but fails 8.5% of hidden document attacks

OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections. Gray Swan testing found Astra was cracked at least once in 8.5 percent of indirect injection scenarios, versus 4.8 percent for Claude Opus 5.

Key takeaway

Astra's near-total direct injection blocking does not carry over to hidden document attacks, where it still fails in a meaningful share of scenarios.

What happened

The Decoder reports that OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections, according to Gray Swan testing cited in the coverage.

When attacks were hidden inside documents the AI reads, Gray Swan testing found Astra was cracked at least once in 8.5 percent of indirect injection scenarios, versus 4.8 percent for Claude Opus 5 in the same evaluation.

Evidence

  • GPT-6 Astra blocks 99.99 percent of direct prompt injections.

    The Decoder · attributed

    OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections.

  • Astra was cracked in 8.5 percent of indirect injection scenarios involving hidden document attacks.

    The Decoder · attributed

    Gray Swan testing found Astra was cracked at least once in 8.5 percent of indirect injection scenarios, versus 4.8 percent for Claude Opus 5.

  • Claude Opus 5 had a lower indirect injection failure rate of 4.8 percent in the same testing.

    The Decoder · attributed

    Gray Swan testing found Astra was cracked at least once in 8.5 percent of indirect injection scenarios, versus 4.8 percent for Claude Opus 5.

  • GPT-6 Astra hallucinates less than its predecessor.

    The Decoder · attributed

    OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections.

Why it matters

For autonomous AI agents handling real data, indirect injection failure rates may be a more relevant safety signal than headline direct-injection block rates.

Limits and uncertainties

The packet does not provide Gray Swan test methodology, attack libraries, or sample sizes behind the reported percentages.

Several source excerpts in the packet are truncated and do not include the full article text.

Practical implications

Teams deploying agents that read untrusted documents should evaluate indirect injection resistance, not only direct prompt injection benchmarks.

Model selection for document-handling agents should compare indirect failure rates, where Astra reportedly underperformed Claude Opus 5 in Gray Swan testing.

What to watch

Whether OpenAI publishes its own benchmarks or mitigations for document-borne indirect prompt injections in GPT-6 Astra.

Additional independent indirect injection test results beyond the Gray Swan figures cited by The Decoder.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections