GPT-6 Astra tops ErdosBench as OpenAI deprioritizes targeted math research
OpenAI's GPT-6 Astra has taken first place on ErdosBench, a benchmark for open mathematical problems, with a score of 3.23 and 106 of 226 problems addressed, including 43 fully solved. The result lands amid a wider flurry of Astra-related claims, from retailer simulations to Navier-Stokes headlines elsewhere in the packet. Chief scientist Jakub Pachocki says mathematics was deliberately not a research priority, with urgency instead on recursive self-improvement and automated alignment work. That framing suggests capability can surface in domains without targeted optimization, reinforcing a spiky development path with uneven strengths. Builders should treat the benchmark as attributed reporting, not independent verification, and note that top leaderboard performance does not by itself confirm how much generalization versus benchmark-specific tuning contributed.
GPT-6 Astra tops ErdosBench as OpenAI deprioritizes targeted math research
OpenAI's GPT-6 Astra tops the ErdosBench for open math problems, scoring 3.23 and solving 106 of 226 open problems with 43 solved completely. Chief scientist Jakub Pachocki says OpenAI did not prioritize mathematics research because of urgency around recursive self-improvement and automated alignment work.
Key takeaway
Recursive self-improvement and alignment-focused R&D can produce top ErdosBench math performance even when OpenAI says mathematics was not a targeted priority.
What happened
According to The Decoder, OpenAI's GPT-6 Astra tops ErdosBench for open math problems, scoring 3.23 and addressing 106 of 226 open problems, with 43 solved completely.
Chief scientist Jakub Pachocki says OpenAI did not prioritize mathematics research because of urgency around recursive self-improvement and automated alignment work, framing the result as an outcome of design choices rather than targeted math optimization.
Evidence
GPT-6 Astra leads ErdosBench with score 3.23 and 106 of 226 problems addressed, 43 fully solved.
The Decoder · attributed
OpenAI's GPT-6 Astra tops the ErdosBench for open math problems, scoring 3.23 and solving 106 of 226 open problems with 43 solved completely.
OpenAI chief scientist Jakub Pachocki says math was deliberately not a research priority.
The Decoder · attributed
Chief scientist Jakub Pachocki says OpenAI did not prioritize mathematics research because of urgency around recursive self-improvement and automated alignment work.
OpenAI reportedly used about 10,000 agents and 130B tokens on a Navier-Stokes singularity search.
Latent Space · attributed
OpenAI reportedly used 10,000 agents and 130B tokens to find a singularity in Navier-Stokes equations, claiming a Millennium Prize contender.
Techmeme cites Axios reporting that an internal model more capable than GPT-6 Astra addressed Navier-Stokes with 10K agents over 88 hours.
Techmeme · attributed
OpenAI says an internal model "significantly more capable than GPT-6 Astra" solved the Navier-Stokes problem using 10K concurrent agents working for 88 hours (Madison Mills/Axios)
Why it matters
If alignment and recursive self-improvement work can spill into untargeted domains, capability roadmaps and deployment risk reviews that assume domain-specific training are required may need revision.
Limits and uncertainties
The packet does not provide independent verification of ErdosBench results or separate how much performance reflects generalization versus benchmark-specific optimization.
Navier-Stokes coverage in the packet frames Millennium Prize claims as hype that may conflate computational search with rigorous mathematical proof.
Practical implications
Treat ErdosBench and Navier-Stokes headlines as attributed signals until OpenAI or third parties publish reproducible methods, data, and peer review.
Update capability assessments to allow spiky emergent strengths from non-domain R&D rather than assuming flat domain-by-domain investment returns.
What to watch
Whether OpenAI publishes detailed ErdosBench methodology and whether independent groups replicate the 3.23 score and 43 complete solutions.
How OpenAI distinguishes Navier-Stokes computational findings from a formal Millennium Prize proof in official documentation.