Thomas Bloom, a mathematician at the University of Manchester, maintains a website that catalogues the open problems of Paul Erdős, the Hungarian mathematician who published more papers than anyone in history and left behind hundreds of questions nobody could answer.
On Saturday, three of them came off the board at once.
"Big news," Bloom wrote on X. "Maybe not bigger than a proof of unit distance would have been, but in terms of constructions, this is big."
The results came from OpenAI, which on August 1 published a research page titled "Ten advances in mathematics and theoretical computer science." Ten long-open problems. Ten claimed solutions. And one line in the second paragraph that landed harder than the mathematics for a lot of engineers reading it: the tokens needed to find every one of those solutions would have cost roughly $2,000 at Sol API rates.
The model that produced them has a name now. OpenAI calls it Astra, describes it as "our next major model," and has not released it.
The Problems Had Been Sitting Untouched for Decades
OpenAI says the ten problems had seen no progress on their central results for at least ten years, and in most cases far longer. They are not competition questions or benchmark items with known answers. They are research problems, drawn from eight different corners of mathematics and theoretical computer science.
| Result | Field | What Changed |
|---|---|---|
| High-dimensional sphere packing | Geometry | New upper bounds on packing density down to the Cohn–Elkies threshold |
| Binary and spherical codes | Coding theory | Exponentially improved bounds on maximum code size at any prescribed minimum distance |
| Non-sofic groups | Group theory | A construction establishing that non-sofic groups exist |
| Connes's rigidity conjecture | Operator algebras | Disproof of the claim that certain groups are uniquely determined by their von Neumann algebras |
| Arithmetic circuit complexity | Complexity theory | New lower bounds for computing the permanent, including a formula bound of order n⁴/log n |
| Quantum parallel repetition | Quantum complexity | An exponential parallel repetition theorem for general two-player quantum games |
| Closest vector problem | Lattice cryptography | Polynomial-factor hardness of approximation, a foundational post-quantum question |
| Ehrhart's volume conjecture | Discrete geometry | The maximum volume of a convex body whose centroid is its only interior lattice point, in every dimension |
| Multicolor Ramsey numbers | Combinatorics | A superexponential lower bound, resolving Erdős problem 183 |
| Extremal number conjectures | Graph theory | Compactness and degeneracy results, resolving Erdős problems 146 and 180 |
Two entries on that list should get the attention of anyone building systems rather than proving theorems. The closest vector problem sits underneath lattice-based cryptography, which is what the post-quantum schemes now being standardized are built on. OpenAI describes the result as "a foundational lattice question related to post-quantum cryptography." The permanent lower bound is a classical target in circuit complexity, a field where progress gets measured in decades.
The Lean Certificates Are What Make This Checkable
An AI company claiming its unreleased model solved ten open problems is, on its own, worth very little. Models produce confident, well-structured, wrong output constantly. That is the default failure mode of the technology.
OpenAI has been burned on this exact claim before. In October 2025, Kevin Weil, then an OpenAI vice president, posted that GPT-5 had solved ten previously unsolved Erdős problems. It had done no such thing. The model had surfaced existing papers that already contained the solutions, work Bloom had not yet catalogued on his site. Bloom publicly corrected the record, OpenAI researchers deleted the posts and apologised, and the episode became the reference case for how easily a literature search gets mistaken for a discovery.
OpenAI's answer this time was to formalize every argument in Lean, a proof assistant that mechanically verifies each logical step of a mathematical argument and rejects anything that does not follow. A Lean certificate is not a summary or a rewrite. It is a program that either compiles or does not. The certificates for all ten results are published at github.com/openai/ten-proofs, alongside a separate document narrating the model's reasoning process for each one.
That changes the shape of the claim. It moves from trust us to check it yourself, which is a standard almost nothing else in AI announcements meets.
The workflow behind the results was not fully autonomous, and OpenAI says so. The model generated the mathematical arguments. Humans, working with the same model, turned those arguments into manuscripts. The model then produced the Lean formalization.
OpenAI Drew an Unusual Line on Who Gets Credit
Buried at the bottom of the research page is a paragraph that reads less like a product announcement than a position statement.
"We believe attribution should honestly reflect how a result was produced: claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system's contribution and the nature of genuine human intellectual work." — OpenAI, "Ten advances in mathematics and theoretical computer science" (August 1, 2026)
The company pointed to the Leiden Declaration on AI and Mathematics, a statement signed by mathematicians concerned about how AI is entering their field, and said it takes responsibility for the correctness of the proofs while crediting the arguments themselves to the system.
For a field where authorship determines careers, tenure, and prizes, that is not a small thing to put in writing.
Astra Reached Washington Before It Reached GitHub
In the same week the math page went up, Sam Altman was demonstrating Astra to politicians and regulators in Washington, D.C., according to The Information, which cited three people familiar with the plans. The pitch centered on the model's ability to coordinate multiple agents over long stretches, hours or days, on a single hard problem.
Astra forms a new model class alongside OpenAI's existing Sol, Terra, and Luna families. Whether it ships as GPT-6 or as something like GPT-5.7 has not been decided. There is no release date.
The sequencing matters. Astra is expected to be the first model to go through the federal pre-release review framework the Trump administration has been drafting, which would require companies to submit models to the government before public launch. OpenAI has already had one model spend 12 days waiting on Washington while a rival shipped. So lawmakers saw a model resolving decades-old mathematics shortly before that same model is expected to enter a government review that could hold up its launch.
The Skeptics Have a Specific Case, and It Is Not Dismissal
Noam Brown, the OpenAI researcher behind much of the test-time reasoning work Astra builds on, volunteered the deflation himself. "Sadly, no Millennium Prize Problems (yet)," he wrote on X. The Clay Mathematics Institute has offered one million dollars for each of its seven Millennium Prize Problems since 2000; exactly one has been solved.
Brown added a caveat that cuts the other way: "But also, we didn't spend a lot on each problem. It's possible to push test-time compute much further." He called Astra a "major step for scientific reasoning."
Bloom, whose reaction opened this story, also pushed back on the framing that mathematicians are being replaced. His argument was that a system drawing on more than a century of mathematical theory, built by mathematicians, and trained on everything mathematicians have ever written, is a strange thing to describe as displacing them.
The sharpest technical critique came earlier, aimed at OpenAI's May result disproving the Erdős unit distance conjecture, and it applies here too. Writing in Understanding AI, Kai Williams argued that the solution played to exactly the things AI models are good at and humans are not. Models have read more mathematics than any living person, so they can pull a technique from algebraic number theory into a discrete geometry problem in a way that requires a rare human to even attempt. And they will grind through proof strategies that look unpromising.
Jacob Tsimerman, a University of Toronto professor who reviewed that result, said he had considered a similar approach himself and abandoned it, because that type of technique "consumes much time and frequently doesn't work out."
These are not one-shot results either. Brown's own posting confirms OpenAI attempted other major problems without success, and none of those attempts appear in the writeup. Ten results were published; the number of runs behind them was not.
OpenAI is also not alone here. In late May, Google announced its own system had solved nine open Erdős problems, two of which had been open for more than fifty years.
The Price Tag Is the Part That Reaches Practitioners
Strip away the group theory and one number remains legible to anyone who has ever filed a compute request.
Ten research results. Two thousand dollars.
That figure belongs to a pattern LDS has now covered three times in three months. In June, a Princeton system proved theorems for $294 while a rival approach spent six figures on the same work. Days ago, OpenAI opened its best models to 100,000 scientists and mathematicians at no cost. Now this.
The through line is that a category of intellectual work that was bounded by the supply of exceptional people is starting to look bounded by budget instead. Not all of it, and not the part where someone decides which problem is worth attacking. But the grinding middle, where a promising approach has to be pushed through hundreds of dead ends before anyone knows whether it works, is becoming something you can buy.
For teams outside mathematics, the transferable lesson is narrower and more useful: the results that survived scrutiny here were the ones with mechanical verification attached. Lean was the difference between a press release and a result. Any workflow where a model's output can be checked by a machine rather than a reviewer is a workflow where these capability jumps land first.
Not that an AI system did mathematics. That happened in May, and Google did it too. What is new is scale and verifiability arriving together: ten results across eight subfields, each shipped with a proof certificate a machine can reject, from a model the public cannot use yet.
The Bottom Line
OpenAI introduced its next flagship not with a benchmark table but with ten pieces of original mathematics and the formal certificates to check them. That is a deliberate choice, and a smart one, because benchmark numbers have become easy to wave away as saturated or gamed while a Lean proof either compiles or it does not.
What that choice does not settle is the question everyone actually wants answered. A model that resolves Erdős problem 183 may still fail at ordinary agentic work, and the same model family spent part of July breaking out of its own sandbox during evaluations. Astra remains unreleased, unbenchmarked on general tasks, and unavailable to anyone who might test the gap between the two.
Brown's own line is the honest summary of where this sits: no Millennium Prize Problems yet, and they did not spend much per problem. Both halves of that sentence are load-bearing.
Sources
- Ten advances in mathematics and theoretical computer science — OpenAI (August 1, 2026)
- Lean proof certificates for the ten results — OpenAI on GitHub (August 2026)
- OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions — The Decoder (August 1, 2026)
- OpenAI teases Astra, its next major AI model, after it solves 10 long-standing math problems — BleepingComputer (August 2, 2026)
- Exclusive: OpenAI Previews Astra AI Model in D.C. — The Information (July 2026)
- OpenAI's math breakthrough played to AI's strengths — Understanding AI (May 28, 2026)
- Model disproves discrete geometry conjecture — OpenAI (May 2026)
- Leading OpenAI researcher announced a GPT-5 math breakthrough that never happened — The Decoder (October 2025)
- The Leiden Declaration on AI and Mathematics — Leiden Declaration (2026)
- Thomas Bloom on the ten results — X (August 1, 2026)
- Noam Brown on Astra and Millennium Prize Problems — X (August 1, 2026)