GPT-6 Astra Hits SOTA on ARC-AGI-3 With 63% Standard Harness Score
ARC Prize reports that OpenAI's GPT-6 Astra has set a new state-of-the-art on the ARC-AGI benchmark family, posting a 63% score on ARC-AGI-3 under what it describes as a standard harness. The organization also cites a 99% result when evaluated through a new provider adapter harness, and says Astra beat median tested humans on action efficiency across 96% of ARC-AGI-3 levels. In its public messaging, ARC Prize characterizes Astra as building the most precise symbolic model of novel environments it has observed to date. The disclosure positions frontier generalization and interactive reasoning as active competitive terrain, but the gap between standard and adapter harness scores means headline comparisons require careful method reading before capability claims travel beyond benchmark context.
GPT-6 Astra Hits SOTA on ARC-AGI-3 With 63% Standard Harness Score
ARC Prize reports GPT-6 Astra by OpenAI achieves SOTA on ARC-AGI. Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness. It surpasses human performance on 96% of ARC-AGI-3 levels.
Key takeaway
GPT-6 Astra's 63% ARC-AGI-3 score under a standard harness marks a reported SOTA, but the 99% adapter-harness figure shows evaluation setup heavily shapes published results.
What happened
ARC Prize reports that GPT-6 Astra by OpenAI achieves state-of-the-art performance on ARC-AGI, scoring 63% on ARC-AGI-3 with a standard harness according to its public post on X.
The same ARC Prize disclosure states Astra reaches 99% via a new provider adapter harness and surpasses human performance on 96% of ARC-AGI-3 levels, using fewer actions than the median tested human on those levels.
Evidence
GPT-6 Astra by OpenAI achieves SOTA on ARC-AGI with a 63% ARC-AGI-3 standard harness score.
X · attributed
ARC Prize reports GPT-6 Astra by OpenAI achieves SOTA on ARC-AGI. Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness.
Astra surpasses human performance on 96% of ARC-AGI-3 levels.
X · attributed
It surpasses human performance on 96% of ARC-AGI-3 levels.
ARC Prize scores Astra at 99% when run through a new provider adapter harness.
X · attributed
Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
ARC Prize describes Astra as building the most precise symbolic model of novel environments it has seen.
X · attributed
It builds the most precise symbolic model of novel environments we've seen
OpenAI states Astra used fewer actions than the median tested human on 96% of ARC-AGI-3 levels.
X · attributed
GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.
Why it matters
If adapter-assisted runs become the default reporting frame, buyers and builders may misread leaderboard jumps as pure model gains rather than harness-dependent performance.
Limits and uncertainties
The packet cites ARC Prize and OpenAI messaging on X only; no independent replication or full technical report is included in the evidence.
Standard harness (63%) and provider adapter harness (99%) scores are not equivalent benchmarks, so direct comparison without method detail is uncertain.
The evidence packet truncates OpenAI's description of a key observed behavior before the claim is completed.
Practical implications
Operators benchmarking frontier models on ARC-AGI-3 should record whether results use the standard harness or the new provider adapter harness before comparing vendors.
Builders evaluating interactive reasoning should treat the 96% human-surpass rate as action-efficiency on ARC-AGI-3 levels, not a blanket real-world autonomy claim.
What to watch
Whether ARC Prize publishes its linked analysis with full harness specifications and per-level breakdowns.
Whether independent labs reproduce the 63% standard-harness ARC-AGI-3 score without the provider adapter harness.
Original reporting: GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:
- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
- It surpasses human performance on 96% of ARC-AGI-3 level