OpenAI Says Astra Produced Ten Advances in Mathematics

OpenAI said August 1 that an internal version of its unreleased Astra model produced ten results across mathematics and theoretical computer science, including three resolutions of Erdős problems. Humans prepared the arguments as manuscripts and the model formalized them in Lean, but independent expert review will determine the correctness, novelty and significance of each claimed advance.
OpenAI said on August 1 that an internal version of its unreleased Astra model generated ten results across mathematics and theoretical computer science. The company described each result as resolving or making substantial progress on a longstanding open problem, including three questions cataloged as Erdős problems.
The announcement covers high-dimensional sphere packing, binary and spherical codes, group theory, operator algebras, arithmetic circuit complexity, quantum complexity, lattice cryptography and extremal combinatorics. For the Erdős set, OpenAI says the work resolves problem 183 on multicolor Ramsey numbers and problems 146 and 180 in extremal graph theory.
How the reported workflow worked
OpenAI says Astra generated the mathematical arguments. Humans then used the same model to prepare manuscripts, after which the model formalized each argument as a Lean certificate. The company also released a collection of manuscripts and model-generated reasoning walkthroughs. It estimated that the tokens used to find the solutions would cost about $2,000 at Sol API rates.
Those details describe a more structured research pipeline than free-form answer generation: model exploration, human manuscript preparation and proof-assistant formalization. A Lean certificate can make a formal argument mechanically checkable against its encoded statement and assumptions. It does not by itself establish that a theorem was framed correctly, that a result is novel or that specialists agree with its importance.
A broader Erdős pattern
Quanta placed the announcement in the context of OpenAI's May unit-distance result, in which an internal model found a counterexample to an 80-year-old conjecture associated with Paul Erdős. Quanta reports that human mathematicians substantially improved that result within weeks and that the episode helped focus attention on why Erdős problems are proving useful test cases for AI-assisted mathematics.
Many such problems have concise statements and outcomes that can be tested through counterexamples, bounds or formal proofs. That makes them amenable to large-scale search and verification. It does not show that the same system can autonomously choose important research directions, formulate productive questions or operate across less structured scientific domains.
For ML practitioners, the central signal is the workflow rather than the headline count. Research systems become more credible when claims are tied to released manuscripts, exact model and prompt records, formal artifacts and independent domain review. Astra is not publicly available, and the retrieved sources do not provide a complete failure rate across attempted problems, so the ten selected successes should not be read as a general success rate for autonomous mathematical research.
Key Points
- 1OpenAI says an internal Astra model produced ten mathematical and theoretical-computer-science results, including three resolutions of Erdős problems.
- 2Humans prepared the arguments as manuscripts and Astra formalized them in Lean, creating checkable artifacts without replacing independent novelty and significance review.
- 3Astra is unreleased, and OpenAI did not publish a complete success rate across attempted problems, limiting broader capability claims.
Scoring Rationale
Ten released mathematical results with Lean artifacts would be a major advance in AI-assisted research if they withstand independent review. The impact remains bounded by Astra's unavailability, incomplete failure-rate reporting and the need for specialist validation of correctness, novelty and significance.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems
