Skip to main content
LLMgram · AI News · 2026-09-03

GPT-6 Astra Hits SOTA on ARC-AGI-3 With 63% Standard Harness Score

GPT-6 Astra Hits SOTA on ARC-AGI-3 With 63% Standard Harness Score

ARC Prize reports that OpenAI's GPT-6 Astra has set a new state-of-the-art on the ARC-AGI benchmark family, posting a 63% score on ARC-AGI-3 under what it describes as a standard harness. The organization also cites a 99% result when evaluated through a new provider adapter harness, and says Astra beat median tested humans on action efficiency across 96% of ARC-AGI-3 levels. In its public messaging, ARC Prize characterizes Astra as building the most precise symbolic model of novel environments it has observed to date. The disclosure positions frontier generalization and interactive reasoning as active competitive terrain, but the gap between standard and adapter harness scores means headline comparisons require careful method reading before capability claims travel beyond benchmark context.

Sources

GPT-6 Astra Hits SOTA on ARC-AGI-3 With 63% Standard Harness Score

GPT-6 Astra Hits SOTA on ARC-AGI-3 With 63% Standard Harness Score

ARC Prize reports GPT-6 Astra by OpenAI achieves SOTA on ARC-AGI. Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness. It surpasses human performance on 96% of ARC-AGI-3 levels.

Key takeaway

GPT-6 Astra's 63% ARC-AGI-3 score under a standard harness marks a reported SOTA, but the 99% adapter-harness figure shows evaluation setup heavily shapes published results.

What happened

ARC Prize reports that GPT-6 Astra by OpenAI achieves state-of-the-art performance on ARC-AGI, scoring 63% on ARC-AGI-3 with a standard harness according to its public post on X.

The same ARC Prize disclosure states Astra reaches 99% via a new provider adapter harness and surpasses human performance on 96% of ARC-AGI-3 levels, using fewer actions than the median tested human on those levels.

Evidence

  • GPT-6 Astra by OpenAI achieves SOTA on ARC-AGI with a 63% ARC-AGI-3 standard harness score.

    X · attributed

    ARC Prize reports GPT-6 Astra by OpenAI achieves SOTA on ARC-AGI. Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness.

  • Astra surpasses human performance on 96% of ARC-AGI-3 levels.

    X · attributed

    It surpasses human performance on 96% of ARC-AGI-3 levels.

  • ARC Prize scores Astra at 99% when run through a new provider adapter harness.

    X · attributed

    Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness

  • ARC Prize describes Astra as building the most precise symbolic model of novel environments it has seen.

    X · attributed

    It builds the most precise symbolic model of novel environments we've seen

  • OpenAI states Astra used fewer actions than the median tested human on 96% of ARC-AGI-3 levels.

    X · attributed

    GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.

Why it matters

If adapter-assisted runs become the default reporting frame, buyers and builders may misread leaderboard jumps as pure model gains rather than harness-dependent performance.

Limits and uncertainties

The packet cites ARC Prize and OpenAI messaging on X only; no independent replication or full technical report is included in the evidence.

Standard harness (63%) and provider adapter harness (99%) scores are not equivalent benchmarks, so direct comparison without method detail is uncertain.

The evidence packet truncates OpenAI's description of a key observed behavior before the claim is completed.

Practical implications

Operators benchmarking frontier models on ARC-AGI-3 should record whether results use the standard harness or the new provider adapter harness before comparing vendors.

Builders evaluating interactive reasoning should treat the 96% human-surpass rate as action-efficiency on ARC-AGI-3 levels, not a blanket real-world autonomy claim.

What to watch

Whether ARC Prize publishes its linked analysis with full harness specifications and per-level breakdowns.

Whether independent labs reproduce the 63% standard-harness ARC-AGI-3 score without the provider adapter harness.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 level