Updated 22 hours ago
OpenAI says its model resolved 100+ math problems, but hasn’t released the list

AI and mathematics

OpenAI says its model resolved 100+ math problems, but hasn’t released the list

OpenAI reports that an internal model resolved more than 100 long‑standing math problems. Its new independent advisory group can publish recommendations, but it cannot control company decisions or compel the release of the underlying proofs.

OpenAI says an internal model it began training on August 28 has resolved more than 100 long‑standing open problems across most areas of mathematics. That is a larger research claim than the company’s recent public results, but the [September 21 announcement](https://openai.com/index/advisory‑group‑on‑mathematics‑and‑ai/) does not name those 100‑plus problems, provide a proof index for them, explain how OpenAI decided that each problem was resolved or say how many have been checked outside the company. The same announcement introduces an independent Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study. The group will advise OpenAI on how results should be reviewed, described and released. Its [own public charter](https://agmai.org/) makes the boundary clear: members can publish recommendations, but they have no decision‑making power at OpenAI or any other AI company. OpenAI also says the group will not advise it on the pace of internal mathematical progress. Those two facts belong together. OpenAI is reporting a large body of work whose supporting evidence is mostly absent from the announcement, while assigning mathematicians a role in organizing its disclosure. The advisers cannot require the company to release a proof, delay an announcement or adopt a particular validation standard. The immediate story is therefore not that 100 open problems have entered the accepted mathematical record. It is that OpenAI says it has accumulated a result set large enough to create a new release and review problem.

“Resolved” is a company claim until the work can be inspected

The 100‑plus figure should be read exactly as OpenAI presents it: a statement about what its internal model has done. The announcement supplies no inventory, so readers cannot tell which fields and problem types are included, whether the count contains complete proofs or disproofs, or how many items rely on formal verification. It also gives no per‑result record of prior human work, expert review or revisions made after review. That missing information matters because mathematical validation has several distinct stages. A laboratory can report that a model found a result. It can release a readable proof and the supporting code or formalization. Specialists can then reproduce the argument, identify gaps and decide whether the result contributes new ideas. Over a longer period, the result can gain broad acceptance. One stage does not automatically establish the next. OpenAI’s earlier [Navier–Stokes release](https://openai.com/index/navier‑stokes‑solution/) shows what a more inspectable package looks like. The company published a paper and a Lean formalization, and it says the proof establishes statements C and D in the Clay Mathematics Institute’s formulation. Those artifacts make scrutiny possible. They do not by themselves amount to outside acceptance, and OpenAI says it does not intend to claim the Millennium Prize. The distinction is unusually concrete for that problem. Under the [Clay Mathematics Institute’s published rules](https://www.claymath.org/millennium‑problems/rules/), a proposed solution must appear in a qualifying outlet, at least two years must pass, and the work must receive general acceptance in the global mathematics community before the institute will consider it. Those are the institute’s prize rules rather than a universal checklist for every theorem, but they show why publication and formal checking are not the same as settled recognition.

The advisers can make recommendations public, but cannot enforce them

The new group’s nine initial members are François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten and Melanie Matchett Wood. According to [the group’s statement](https://agmai.org/), OpenAI first approached some members about an external advisory board; they agreed with the company to form an independent group and invite others. The resulting structure gives the group several meaningful forms of independence. Members are unpaid for this work, can advise any AI company whose models may substantially affect mathematics, and plan to publish their recommendations. The group says its current task is to advise OpenAI on coordinating the release of a large number of significant results that the company reports its internal model has produced. Its authority stops at advice. The group says responsibility for company decisions remains with each company. OpenAI says members may challenge it publicly, offer unsolicited recommendations and change the group’s membership, while excluding decisions about how quickly its internal math research proceeds from the group’s remit. That combination can produce transparency if the recommendations are specific and published promptly. It cannot guarantee that OpenAI will follow them. The practical test will be whether the public can compare a recommendation with the action that followed. A dated recommendation about release order, attribution or review would create an auditable record. General consultation without a disclosed recommendation would make it harder to know where the advisers influenced the process and where the company chose a different course.

The dispute is about the research process, not only correct answers

OpenAI announced the group ten days after prominent mathematicians published [“A Severe Misalignment of AI in Mathematics”](https://mathandai.org/) on September 11. [AGMAI says](https://agmai.org/) the group formed after OpenAI approached some members about an external board; the mathematicians instead established an independent group in agreement with the company. The declaration’s current page lists 27 Fields medalists. It argues that a rush to generate answers to open problems can weaken the slower work that turns a proof into shared understanding: careful exposition, connection to prior work, discussion, simplification and teaching. The signatories also raise questions about attribution and plagiarism when results are announced before earlier work has been traced and credited. Their argument is broader than whether a proof is technically correct. A correct but poorly contextualized proof can still make it difficult to see which ideas are new, which came from existing literature, and which researchers supplied the framing or intermediate steps that made the result possible. This is where a batch of more than 100 claimed resolutions creates a capacity problem. Each item may require specialists in a different subfield, a literature search, a clean exposition and a record of how the model used prior work. A large result count does not expand the community’s review capacity at the same rate. Coordinating releases can reduce that pressure, but only if release decisions include the evidence needed for experts to assess each item.

The Navier–Stokes package shows the scale of generation, not the end of review

OpenAI’s account of the Navier–Stokes effort illustrates why those records should include computational provenance. The company reports that the effort used on the order of 10,000 concurrent agents. It says the agents reached the result about 88 hours after the work began, followed by 17 hours of Lean formalization and verification using GPT‑6 Astra. Across all attempted problems, OpenAI reports 4.9 million messages and about 300 billion output tokens; it attributes 2.7 million messages and roughly 130 billion tokens to the Navier–Stokes work. These are OpenAI’s operational figures, not independent measurements. They nevertheless show that a short headline such as “AI solved a math problem” can compress a large search and verification system into a single actor. The useful provenance record would identify the model version, prompts, tools, source corpus, agent structure, human interventions, failed paths, compute, formal checker and later corrections for each released result. That record would also help separate three questions that are often collapsed. Did the system produce a complete argument? Does a formal checker accept the encoded statement and proof? Does the formal statement match the mathematical claim specialists believe has been resolved? Lean can strengthen confidence in a correctly represented proof, but the public still needs to see what was formalized and whether it corresponds to the original problem.

A 100‑result release needs a result‑by‑result ledger

The advisory group can make the next disclosure substantially more useful by recommending a common ledger for every claimed resolution. Each entry should name the problem and its prior status; link a readable proof and any formal artifact; identify the exact statement proved or disproved; document the relevant literature and human contributions; name external reviewers with their consent; record open objections and revisions; and state whether the result is company‑reviewed, formally checked, independently reproduced or broadly accepted. The labels should remain separate. “Released” means the materials are public. “Formally checked” means a specified checker accepted a specified formalization. “Externally reviewed” should identify who reviewed what and when. “Accepted” requires evidence from the relevant mathematical community rather than a company announcement. If no proof has been released, “reported by OpenAI” is the most accurate status. OpenAI has not yet provided that result‑by‑result record for the more‑than‑100 claim. The advisory group has provided a mechanism for producing public recommendations, and its independence gives those recommendations potential weight. Whether the group changes the release process will become visible only when the recommendations and the first batch of underlying work appear. Until then, the defensible finding is narrower than the headline number: OpenAI reports a major increase in its internal model’s mathematical output, while the announcement does not provide the evidence needed to evaluate the aggregate claim.

Sources

Share this article

PostShare

Related News