Skip to content

AI·News & analysis

OpenAI publishes 722 math papers written by an unreleased AI model

OpenAI put 722 manuscripts from an internal model on GitHub, with computer-checked proofs for many of them, after mathematicians asked AI labs to show their work.

By Dan Kost aka Poseidan8 min readX
Fuld Hall, a red brick building with a white clock tower, at the Institute for Advanced Study in Princeton, New Jersey, with cars parked in front
Photo: Famartin / Wikimedia Commons, CC BY-SA 4.0

Tide

Ripple

Sci-fi

9/10

Reality

Lab

An AI spends three hours each on problems mathematicians left open.How we rate

The Squeeze

OpenAI released 722 math manuscripts produced by an unreleased internal model, grouped into 372 result families, in a public GitHub repository.

Many come with Lean proofs that a computer can check, but not all of them. The model was posed about 4,000 problems, and an advisory group of mathematicians says the batch solves hundreds of open questions. Mathematicians will now review the papers, and OpenAI says it will record corrections as new versions.

What to know

  1. OpenAI released 722 math manuscripts in 372 result families, produced by an unreleased internal model, in a public GitHub repository.
  2. The model was given about 4,000 problems. The average result used the equivalent of about three hours of ChatGPT Pro thinking.
  3. Many papers come with Lean proofs a computer can check, but not all. OpenAI says some unformalized results could have issues.
  4. The release follows recommendations from an independent advisory group of mathematicians at the Institute for Advanced Study.

OpenAI announced on Tuesday that it is publishing 722 math manuscripts written by an unreleased AI model. The papers sit in a public GitHub repository, and an advisory group of mathematicians says they include solutions to "hundreds" of open questions.

That's the news. Now the interesting question: how do you check 722 papers that no human wrote?

What exactly did OpenAI publish?

The repository holds a lot more than a pile of PDFs. According to its README and Unite.AI's walkthrough of the release, here's what's inside:

  • Manuscripts: 722 papers, grouped into 372 result families.
  • Families: each one bundles related papers, such as a main result, companion arguments, consequences or alternative proofs. Every family is labeled by area of mathematics.
  • Preprints folder: PDFs, source files, and build and citation instructions for each paper.
  • Lean library: formal proofs plus a catalogue showing which paper each proof belongs to and how it was checked.
  • Reasoning summaries: 10 abridged write-ups of how the model reached certain results.
  • License: Apache-2.0, so anyone can read and reuse the material.

Each manuscript folder also carries a BibTeX block, the standard snippet researchers paste in when they cite a paper.

Where did all these papers come from?

OpenAI says they started as tests. As part of building its models, it evaluates them on open research problems, the kind nobody has solved yet. It expanded those tests after its models maxed out its existing math benchmarks.

By the numbers: the model was posed about 4,000 problems. OpenAI says the average result used the equivalent of about three hours of ChatGPT Pro thinking. It then kept only results that met its bar for significance and grouped them into families and papers.

The repository also says some outputs build on earlier results the models produced. So part of the batch is the model building on its own earlier work.

Did a human help write them?

Mostly no, OpenAI says. The vast majority of results came from the same fixed procedure with the same internal model.

There are two named exceptions:

  • The Riemann zeta function: work on a zero-free region, a stretch where the function provably has no zeros. The writeup for the region where Re(s) is greater than 11/12 was human edited for readability.
  • The Hodge Conjecture for CM abelian varieties: a proof covering one special family of shapes in geometry.

Both topics sit close to famous unsolved problems. The Riemann Hypothesis is about where the zeta function's zeros lie, and the full Hodge Conjecture is one of the seven Millennium Prize Problems. These results tackle narrower pieces, not the full prizes.

How do you check a proof nobody wrote?

This is where Lean comes in. Lean is a programming language for writing proofs so a computer can verify every single step. If a proof passes, it doesn't depend on a tired reviewer spotting a slip on page 40.

Think of it like a spell checker for logic, except it doesn't let anything slide, not even one "obviously true" step.

The catch: not every paper has a Lean proof yet. OpenAI's own README warns that some of the unformalized results could have issues. The company says it will fix problems quickly and add formal proofs as it gets them.

It also promises a paper trail. Corrections and revisions will be saved as new versions, and older versions will stay public.

In real life If you're a student or a math hobbyist, you can open the repository today, pick a family in your favorite area, and see whether it has a Lean proof attached. The ones that do are the ones a computer has already checked.

What did the model work on?

The 10 reasoning summaries give a sense of the range. They cover:

  • Number theory: two-point correlations of multiplicative functions, the irrationality exponent of π, and quasipolynomial bounds for arithmetic progressions.
  • Geometry: the symmetric and general Mahler conjectures.
  • Computer science: NP-hardness at the basic semidefinite threshold.
  • Algebra: Kaplansky's direct-finiteness conjecture in characteristic two.
  • Physics-flavored math: the Mézard–Parisi formula for diluted spin glasses, spontaneous magnetization in the quantum Heisenberg ferromagnet, and the three-dimensional relativistic Vlasov–Maxwell system.
  • Operator algebras: the isomorphism of free group factors.

You don't need to know what any of these mean. The point is the spread: the batch runs from prime-number puzzles to models of magnets.

Why is OpenAI releasing it this way?

Because mathematicians asked. The Verge reports that in September, OpenAI said its model had "resolved more than 100 long-standing open problems across most areas of mathematics," without saying which ones. Mathematicians had been waiting weeks for the details.

On September 29, the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), an independent group at the Institute for Advanced Study in Princeton, published recommendations for AI labs. Unite.AI reports the group drew on more than 600 replies from the math community. Its main asks:

  • Release where no lab is in control: scholarly repositories with permanent, citable identifiers.
  • Show the work: the model name, the prompts, a summarized chain of thought, the time taken and the estimated compute cost for each result.
  • Formalize: check proofs by computer as far as possible.
  • Count the misses: for big batches, say how many problems of similar difficulty the model tried and failed, and how problems were chosen.
  • Support people: fund workshops, summer schools and books that help humans understand AI-made math, with nonprofits deciding where the money goes.

The group also asked labs not to treat math results as marketing for their models. And it went further: "we do not endorse this practice," it wrote about testing hard problems on proprietary models, "and we ask them to stop."

OpenAI says it consulted the group and drew on its advice for this release. The compute estimates, the 4,000 attempted problems and the reasoning summaries line up with several of the asks. The model's name and a home outside GitHub don't, yet.

What does OpenAI say comes next?

OpenAI described its plan in a statement quoted by The Verge. It chose GitHub for now, "with protocols for paper revisions and citations," and is "continuing to explore other community-hosted alternatives" that meet the committee's guidelines. For future releases, it says it will improve the citations, the explanations and the presentation of results.

The big picture: The Verge notes the field is still digesting a fast-growing pile of AI math results from OpenAI and rivals such as Anthropic, including results connected to a Millennium Prize problem. It also reports a heated debate in the field over research practices, ethics and how labs credit the human mathematicians whose work their systems build on.

What's next: mathematicians now get to read, check and argue. With 722 papers, that will take a while, and the version history is where you'll see the outcome.

What it means for you

  • You can read it all: the repository is public and free, and the license lets anyone reuse the material.
  • You can't use the model: it's an internal, unreleased system, so ChatGPT won't solve these problems for you today.
  • Look for the Lean proofs: a paper with a formal proof has been checked by a computer. One without it is still waiting for human review.
  • Watch the corrections: OpenAI keeps old versions public, so any fixes will be easy to track.

The bottom line

An AI model produced 722 math papers from about 4,000 attempts, and OpenAI put them in the open with a lot of the evidence attached. The real test starts now, as human mathematicians check the work. For more on the company, see our OpenAI coverage.

Key facts

Manuscripts
722, grouped into 372 result families
Problems attempted
About 4,000
Compute per result
About 3 hours of ChatGPT Pro thinking, on average
Where
GitHub repository openai/math, Apache-2.0 license
Model
Unreleased internal OpenAI model, not named

Got questions?

Quick answers, plain words

What did OpenAI release?

722 mathematical manuscripts, grouped into 372 result families, plus Lean proof files and 10 abridged summaries of the model's reasoning. Everything is in a public GitHub repository called openai/math.

Which AI model wrote the papers?

An unreleased internal OpenAI model. The repository doesn't name it.

How many problems did the model try?

About 4,000, according to the repository. OpenAI kept the results that met its bar for significance and grouped them into families and manuscripts.

How much computing did each result take?

On average, the equivalent of about three hours of ChatGPT Pro thinking, OpenAI says.

Are the proofs correct?

Many manuscripts have Lean formalizations, which a computer can check line by line. Not all do, and OpenAI warns that some unformalized results could have issues. It says it will fix them quickly and record corrections as new versions.

What is Lean?

Lean is a programming language for writing mathematical proofs so a computer can verify every step. A proof that passes Lean's checker doesn't depend on a human reviewer catching mistakes.

Did humans help with any of the results?

OpenAI says the vast majority came from the same fixed procedure. The exceptions are work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties. The zeta function writeup was human edited for readability.

What is AGMAI?

The Advisory Group on Mathematics and Artificial Intelligence, an independent group of mathematicians at the Institute for Advanced Study. It published recommendations on September 29, 2026, on how AI labs should release math results.

Can I use the model that did this?

No. The model is unreleased. The papers, proofs and summaries are public on GitHub under an Apache-2.0 license.

Is this the first time OpenAI claimed math breakthroughs?

No. In September, OpenAI said its model had resolved more than 100 long-standing open problems across most areas of mathematics, The Verge reports. This release is the first to show which problems and how.

SourcesOpenAI, github.com
Topics and tagsOpenAI, openai, mathematics, ai research

The daily newsletter

Tech news you'll actually get.

One short email a day. Five minutes. Plain words. Free, every morning.

Free. One email a day. Unsubscribe anytime.

More in brief