Updated 2 days ago
Claude’s nine-loop result is public, but “verified” has four different meanings

AI News

Claude’s nine-loop result is public, but “verified” has four different meanings

Claude’s nine‑loop amplitude has unusually detailed checks. The public record shows where exact agreement ends and independent reproduction is still missing.

What the public record establishes

Anthropic’s Claude has produced a nine‑loop scattering‑amplitude result that researchers had previously reached only through eight loops. The data are public, the methods are known, and a leading amplitudes physicist spent two weeks checking much of the work. Those facts make the result substantially stronger than a company demo. They do not all prove the same thing. The published record contains four different layers of evidence: reproduction of the known eight‑loop benchmark, agreement between two computational routes for the nine‑loop symbol, public data files with recorded checks, and a full function that has been computed once without an independent second calculation. Keeping those layers separate gives a more useful answer than asking whether the result is simply “verified.” The distinction matters beyond this one physics problem. AI‑assisted research increasingly arrives with code, datasets, expert review and institutional claims mixed together. Here, the unusually detailed [public result page](https://smsharma.io/cosmic‑nine‑loops/) makes it possible to say exactly which claims can be checked and which still depend on the original team.

The result extends a real scientific benchmark

The calculation concerns a six‑particle amplitude in planar N=4 super Yang‑Mills theory. This is not a realistic model of nature, but it is a productive test bed for methods that calculate how particles scatter. Each additional loop increases the mathematical complexity, and the previous published frontier was an [eight‑loop amplitude computed by Lance Dixon and Andrew Liu in 2023](https://arxiv.org/abs/2308.08199). In an [Anthropic research post published September 25](https://www.anthropic.com/research/yes‑claude‑can‑do‑nine‑loops), physicist Matt von Hippel describes how Liam Fitzpatrick and Siddharth Mishra‑Sharma used Fable 5.1 inside Claude Science to apply two established routes to the nine‑loop problem. One was a direct bootstrap in the space of hexagon functions. The other first computed a related form factor and then used antipodal duality to map it to the amplitude. That is an execution result rather than a new physical principle. The researchers supplied the problem and a harness that repeatedly prompted Claude to continue, while the model developed the calculation and code. Anthropic’s account estimates each route would cost an end user roughly $1,000 to $2,000. The direct bootstrap’s conventional compute component was about $100, using 96 CPUs for a week. Those figures come from Anthropic’s commissioned guest post, not an independently audited cost study. The strongest baseline check is straightforward: the same programs reproduced the published eight‑loop symbol on 1,000 randomly selected words with nonzero coefficients. All 1,000 agreed modulo the first prime. That does not validate every path through the new calculation, but it shows that the machinery can recover a known result before being extended one loop further.

Two nine‑loop routes agree across 107,053 coefficients

The most substantial internal cross‑check is the comparison between the direct bootstrap and the form‑factor route. The result page says the two representations agree on all 107,053 nonzero word coefficients that determine the published septuple file. This is much stronger than a handful of spot checks because the comparison covers every nonzero coefficient in that representation. The exactness is slightly more qualified than the headline number suggests. Of those 107,053 coefficients, 3,401 were compared only modulo two primes. The remaining 103,652 were reconstructed exactly. OpenTools calculates that as 96.8% exact comparisons and 3.2% modulo‑only comparisons. Modular agreement is meaningful evidence, but it is not identical to reconstructing the corresponding rational number. A separate matrix on the form‑factor route shows a similar pattern. The page reports 1,014,476 of 1,018,297 nonzero coordinates reconstructed as certified rationals, or 99.62%; 3,821 coordinates did not reconstruct. That leaves about 0.38% represented through modular residues. These proportions do not imply that the unresolved entries are wrong. They identify where the public evidence has a different mathematical status. This layer is best described as independent‑route agreement inside one research effort. The routes use different constructions and converge on the same symbol‑level result, which is a strong error check. They also share conventions, target objects and parts of the published framework, so agreement is not the same as a second group rebuilding the calculation from scratch.

Public data improves auditability, but the programs are absent

The team released computer‑readable files for the symbol and full function, including large archives hosted on Zenodo, along with conventions, file formats and validation records. The result page links the full 107,053‑row comparison, a 2,000‑word check between representations, the eight‑loop control, symmetry tests and vanishing tests. That is enough for specialists to inspect the outputs and test many of the stated relationships. The page also says explicitly that the programs used for the computation are not distributed. A researcher can download the results, but cannot yet rerun the original end‑to‑end pipeline from the public package. This limits reproducibility even though the outputs are unusually transparent. Independent reporting by [Unite.AI](https://www.unite.ai/anthropic‑says‑claude‑computed‑a‑nine‑loop‑particle‑physics‑amplitude/) correctly notes both sides of that record: the two routes agree across the published coefficients, while the full function has only been computed once and the programs are unavailable. The article also identifies an important competing result. A team led by Song He posted a nine‑loop symbol dataset on Zenodo after reaching much of the result with GPT‑6 assistance. That makes the field less dependent on a single AI system, but the public accounts available here do not yet establish a complete coefficient‑by‑coefficient comparison between the two teams’ outputs. This is why “the files are public” and “the calculation is reproducible” should not be used interchangeably. The first is true. The second still requires a runnable implementation or a genuinely independent reconstruction with a documented match.

The symbol and the full function have different evidence

A scattering amplitude’s symbol captures a large part of its iterated‑integral structure, but it does not by itself contain every function‑level constant. The public package goes further and supplies what it describes as a complete function, including 786,352 beyond‑the‑symbol coefficients. That is scientifically useful—and the point at which the validation record becomes thinner. The result page states that the function was obtained separately, has been computed once and has no second independent computation. It also depends on an additional assumption: relations among the septuple components that hold at symbol level are taken to hold at function level, including some empirical higher‑level relations that need not follow automatically. Dixon’s review adds expert scrutiny without erasing that limitation. According to the Anthropic post, he spent roughly two weeks validating the work, mostly through the nine‑loop form factor. He is a co‑author of the prior eight‑loop result and helped establish the framework being extended. The same post discloses that Anthropic gave him Claude usage credits, and that Anthropic compensated von Hippel to write the guest post and commented on drafts. The accurate conclusion is therefore narrower than either “Claude solved physics” or “this is only a press release.” The nine‑loop symbol has multiple serious checks, including complete coefficient‑level agreement between two routes. The public full function is a concrete research output, but it has a single‑computation evidence boundary that the authors themselves acknowledge.

What would turn the result into an independently reproduced one

The clearest next step is a second end‑to‑end calculation of the full function. Publishing the original programs would let other researchers inspect implementation choices and rerun the pipeline, while a separate implementation would test whether the result survives different code, assumptions and failure modes. A coefficient‑level comparison with the Song He team’s dataset would also clarify how much independent convergence already exists at symbol level. Those checks would answer different questions. Code release would improve reproducibility of the original route. Independent implementation would reduce shared‑error risk. Cross‑team comparison would establish whether the public outputs agree rather than merely showing that two groups reached the same broad frontier. For now, the evidence is strong enough to treat the nine‑loop symbol as a serious scientific result and specific enough to map its remaining uncertainty. Its best feature is not that an AI system produced an enormous file. It is that the researchers exposed enough of the verification ledger to show where exact agreement ends, where modular evidence begins, and where the full function still waits for an independent calculation. *Figure: Four evidence layers in the Claude‑assisted nine‑loop result. Sources: [Anthropic Research](https://www.anthropic.com/research/yes‑claude‑can‑do‑nine‑loops), September 25, 2026; [Cosmic9 result page](https://smsharma.io/cosmic‑nine‑loops/), accessed September 27, 2026. Figure by OpenTools Team; no third‑party expressive material used.*

Share this article

PostShare

Related News