Skip to main content
LLMgram · AI News · 2026-09-10

GPT-6 Astra tops ErdosBench as OpenAI deprioritizes targeted math research

GPT-6 Astra tops ErdosBench as OpenAI deprioritizes targeted math research

OpenAI's GPT-6 Astra has taken first place on ErdosBench, a benchmark for open mathematical problems, with a score of 3.23 and 106 of 226 problems addressed, including 43 fully solved. The result lands amid a wider flurry of Astra-related claims, from retailer simulations to Navier-Stokes headlines elsewhere in the packet. Chief scientist Jakub Pachocki says mathematics was deliberately not a research priority, with urgency instead on recursive self-improvement and automated alignment work. That framing suggests capability can surface in domains without targeted optimization, reinforcing a spiky development path with uneven strengths. Builders should treat the benchmark as attributed reporting, not independent verification, and note that top leaderboard performance does not by itself confirm how much generalization versus benchmark-specific tuning contributed.

Sources

GPT-6 Astra tops ErdosBench as OpenAI deprioritizes targeted math research

GPT-6 Astra tops ErdosBench as OpenAI deprioritizes targeted math research

OpenAI's GPT-6 Astra tops the ErdosBench for open math problems, scoring 3.23 and solving 106 of 226 open problems with 43 solved completely. Chief scientist Jakub Pachocki says OpenAI did not prioritize mathematics research because of urgency around recursive self-improvement and automated alignment work.

Key takeaway

Recursive self-improvement and alignment-focused R&D can produce top ErdosBench math performance even when OpenAI says mathematics was not a targeted priority.

What happened

According to The Decoder, OpenAI's GPT-6 Astra tops ErdosBench for open math problems, scoring 3.23 and addressing 106 of 226 open problems, with 43 solved completely.

Chief scientist Jakub Pachocki says OpenAI did not prioritize mathematics research because of urgency around recursive self-improvement and automated alignment work, framing the result as an outcome of design choices rather than targeted math optimization.

Evidence

  • GPT-6 Astra leads ErdosBench with score 3.23 and 106 of 226 problems addressed, 43 fully solved.

    The Decoder · attributed

    OpenAI's GPT-6 Astra tops the ErdosBench for open math problems, scoring 3.23 and solving 106 of 226 open problems with 43 solved completely.

  • OpenAI chief scientist Jakub Pachocki says math was deliberately not a research priority.

    The Decoder · attributed

    Chief scientist Jakub Pachocki says OpenAI did not prioritize mathematics research because of urgency around recursive self-improvement and automated alignment work.

  • OpenAI reportedly used about 10,000 agents and 130B tokens on a Navier-Stokes singularity search.

    Latent Space · attributed

    OpenAI reportedly used 10,000 agents and 130B tokens to find a singularity in Navier-Stokes equations, claiming a Millennium Prize contender.

  • Techmeme cites Axios reporting that an internal model more capable than GPT-6 Astra addressed Navier-Stokes with 10K agents over 88 hours.

    Techmeme · attributed

    OpenAI says an internal model "significantly more capable than GPT-6 Astra" solved the Navier-Stokes problem using 10K concurrent agents working for 88 hours (Madison Mills/Axios)

Why it matters

If alignment and recursive self-improvement work can spill into untargeted domains, capability roadmaps and deployment risk reviews that assume domain-specific training are required may need revision.

Limits and uncertainties

The packet does not provide independent verification of ErdosBench results or separate how much performance reflects generalization versus benchmark-specific optimization.

Navier-Stokes coverage in the packet frames Millennium Prize claims as hype that may conflate computational search with rigorous mathematical proof.

Practical implications

Treat ErdosBench and Navier-Stokes headlines as attributed signals until OpenAI or third parties publish reproducible methods, data, and peer review.

Update capability assessments to allow spiky emergent strengths from non-domain R&D rather than assuming flat domain-by-domain investment returns.

What to watch

Whether OpenAI publishes detailed ErdosBench methodology and whether independent groups replicate the 3.23 score and 43 complete solutions.

How OpenAI distinguishes Navier-Stokes computational findings from a formal Millennium Prize proof in official documentation.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design