Initializing portfolio

000

Aravind.
All articles
AI4 min read

OpenAI's Astra solved ten open maths problems — the receipts are the story

OpenAI says an unreleased model called Astra closed ten problems that had been open for a decade or more, for around $2,000 in compute. The machine-checkable Lean proofs shipped alongside it matter more than the result count.

AravindChief Technology Officer & Advisor — AI, Cloud & Cybersecurity

OpenAI says an internal build of its next major model, Astra, has produced ten new results in mathematics and theoretical computer science. Each of the problems had been open for at least a decade. There's a 249-page manuscript, and alongside it a machine-checkable Lean 4 certificate for every result, published on GitHub.

The certificates are the part I'd pay attention to. More on that below.

What was claimed

The headline result is the first explicit construction of a non-sofic group. Mikhail Gromov introduced soficity in 1999, and in the 27 years since, nobody had shown whether non-sofic groups exist at all. That's now answered.

The rest covers a lot of ground:

  • A disproof of Connes's rigidity conjecture on von Neumann algebras
  • A proof of Ehrhart's volume conjecture
  • Three problems from Paul Erdos's catalogue, including number 183 on multicoloured Ramsey numbers
  • The first improvement since 1978 to the general upper bound on high-dimensional sphere-packing density
  • A parallel repetition theorem for two-player quantum games
  • New lower bounds on the circuit complexity of computing the permanent

Sebastien Bubeck, who heads mathematics research at OpenAI, confirmed the results publicly and noted that each ships with its Lean certificate and a chain-of-thought walkthrough. OpenAI puts the total compute for all ten at roughly $2,000 at Sol API rates.

Two thousand dollars

That figure is doing more work than the proof count.

Ten results, each one open for somewhere between a decade and 27 years, for about the price of a laptop. If that number holds up, the cost of searching for this class of proof has stopped being the limiting factor. What's left is choosing which problems to point the thing at, and checking what comes back.

Why the Lean certificates matter more than the proofs

The objection to AI-generated proofs has been practical, not philosophical. They're hard to check. A long prose argument from a model arrives on a reviewer's desk as weeks of unpaid work with no guarantee at the end.

A Lean certificate changes that. Anyone with the compiler can validate the result without trusting OpenAI, the model, or the manuscript. Verification stops depending on who produced the thing.

That's a different situation from the usual "AI does science" announcement, and it's why this one is worth a closer read.

The part that isn't settled

How the results get accepted is another question.

In June, mathematicians published the Leiden Declaration, endorsed by the International Mathematical Union. It warns that AI companies are using published research without consent, going around peer review, and putting proof and attribution at risk. It named the practice of announcing results by press release instead of through journals.

OpenAI has been here before. In May it announced that the same long-horizon model family had disproved the Erdos unit distance conjecture, an 80-year-old problem in discrete geometry. Fields Medalist Tim Gowers said he'd recommend that proof for publication in Annals of Mathematics without hesitation, which is not the sort of thing he says lightly.

Thomas Bloom, who runs the erdosproblems site, called the new batch big news and rated it above the unit distance counterexample.

So the mathematics is landing with the people qualified to judge it. The process around it hasn't caught up. A blog post isn't a journal, and a Lean file settles correctness without settling credit, consent, or review.

What I take from this

Astra has no release date. OpenAI describes it only as its next major model, and the speculation that it's the GPT-6 series is speculation. The company is also giving 100,000 academic researchers free access to its frontier models through 2027, which deepens its ties to the research community and concentrates research infrastructure on one company's platform. Both of those are happening.

The transferable lesson here isn't about mathematics. It's that a result you can verify mechanically is worth far more than a result you have to take on trust. OpenAI shipped proofs with checkers attached, and that's why the claim is being examined rather than dismissed. Same principle applies to any agent system doing work nobody can personally audit — if you can't hand someone a way to check the output, you're asking them to believe you.

Source: The Next Web

#OpenAI#AI Research#Mathematics#Formal Verification#Frontier Models

Comments

Checking you're human…

Keep reading

Get the next essay first

Checking you're human…

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.