Avila Labs / Journal

AI / MATHEMATICAL RESEARCH

Project Penelope: An AI-Assisted Investigation of Exact Quasisymmetry

This article documents an AI-assisted investigation of exact quasisymmetry, beginning with the Garren-Boozer near-axis hierarchy and concluding with a bounded search of the full finite-volume equations. It summarizes the mathematical results, numerical failures, and methodological lessons of the project.

Project Penelope began with a sharply defined objective: prove that a non-axisymmetric, finite-radius, exactly quasi-axisymmetric vacuum magnetic field exists, or prove that it cannot. The project reached neither conclusion, and the problem remains open.

The record is nevertheless useful for understanding how language models behave during a sustained mathematical research program. Over the course of the project, the models produced substantive mathematics, software, and diagnostic tools. They also repeatedly assigned too much significance to optimizer behavior. Much of the work therefore concerned the design of tests, independent evaluators, and stopping rules capable of separating valid results from suggestive numerical output.

The models accelerated implementation and verification; standards of evidence and termination still required explicit human control.

Why I started

The immediate catalyst was the 2026 counterexample to the Jacobian conjecture in dimension three. I was especially interested in the form of the result: an explicit polynomial map developed with Claude Fable whose decisive properties could be checked exactly. Its Jacobian determinant was a nonzero constant, yet several distinct points had the same image. The result ended in a finite object that other people could verify directly.

I wanted to know whether a similar style of AI-assisted work could make progress on an important open question in fusion mathematics. In the Fusion Problem Book, the question I called P1 concerned exact quasisymmetry: can a genuinely non-axisymmetric magnetic field possess the symmetry needed for strong particle confinement throughout a finite volume, not merely approximately or on one surface?

I began without specialist training in stellarator theory. I could choose the objective, demand artifacts, run independent checks, and make scope decisions, while many long symbolic derivations exceeded what I could adjudicate immediately. The models could generate technical work faster than I could independently absorb it, making research governance a central concern from the outset.

The mathematical target

Quasisymmetry is valuable because it allows a three-dimensional stellarator field to retain an important confinement property associated with symmetry. Modern numerical designs can make symmetry-breaking extraordinarily small. Project Penelope was interested in the mathematical zero: either an exact construction with proof, or a proof that the zero cannot occur beyond axisymmetry.

The first and longest phase used the Garren-Boozer near-axis construction. The magnetic field and flux surfaces are expanded outward from the magnetic axis, order by order. A well-known counting argument suggests that the hierarchy becomes overdetermined at higher order, but counting equations and unknowns is not a theorem. Rank deficiencies, gauge freedoms, exceptional branches, and differential relationships can all invalidate a naive count. A special family might still evade the generic obstruction.

The prospect of either a rigidity certificate or an exceptional construction gave the problem a superficial resemblance to the Jacobian episode. The verification structures, however, differed substantially. The Jacobian counterexample was a finite polynomial object with an immediate exact check. Project Penelope involved unknown functions, periodic differential equations, coordinate gauges, an asymptotic series whose convergence was not guaranteed, and ultimately a finite-volume nonlinear operator. A perfect finite-order near-axis solution alone would be insufficient to establish the existence of an analytic field at finite radius.

Building a trustworthy computational instrument

Before asking the models to search, I needed them to reconstruct enough of the mathematical machinery to be checked. The project independently rebuilt the low-order near-axis equations, compared them with pyQSC, and developed symbolic collectors, elimination routines, rank tests, and numerical validation batteries.

This work exposed several sources of computational error. One early problem came from applying spectral differentiation on a one-field-period grid to Cartesian frame components that were periodic only over the full torus. The resulting residual was wrong at leading order. Another investigation showed that pyQSC's third-order arrays served a narrower flux-constraint purpose than the project had initially assumed. In both cases, the discrepancy was resolved by inspecting the source, deriving an independent path, and adding a regression test.

This was among the models' strongest work. Language models were effective at translating equations into code, producing independent implementations, constructing manufactured tests, and tracing a bad residual through a long computational chain. Concrete questions about conventions, divergent harmonics, and expected identities gave them well-defined targets and made their output comparatively easy to verify.

Numerical convergence and false residual floors

The harder failure mode appeared once the project began optimizing obstruction measures. The models were very good at making a number smaller and very eager to interpret each stall as mathematical structure. Several apparent residual floors were announced, investigated, and later overturned by a finer grid, a wider Fourier window, a different elimination chart, or higher-precision arithmetic.

One endpoint looked excellent on a 61-point grid and substantially worse at 121 and 181 points. That was grid overfitting, not proximity to an exact field. At another stage, a trajectory that appeared to fall from roughly 3.5 to 0.055 had actually spliced together several different spectral windows. The objective had changed from low harmonics to progressively wider bands, and the same configuration could move by orders of magnitude depending on which window was reported. Later, a value treated as a genuine float64 floor was breached by an extended-precision instrument by a factor of more than sixty.

The models often discovered these problems themselves. They retracted results, added resolution ladders, widened windows, and built independent arithmetic paths. These corrections usually followed a confident interpretation of the previous number. The principal weakness concerned calibration: the latest successful computation was often discussed as evidence about the theorem before the measurement had survived an adversarial test.

  • Small residuals required exact or certified follow-up.
  • Optimizer stalls supplied no positive lower bound.
  • Search failure could not establish nonexistence.
  • Instrument validation established computational correctness within its test regime, without resolving the theorem.

The independent audit changed the project

By late July, I could no longer tell whether Fable was converging or merely producing the appearance of progress. I asked a separate Claude instance to read the record as an adversarial reviewer. Its assessment was blunt: the campaign was genuinely honing in on structure and meandering on the headline number.

The audit treated the exact rank calculations and symbolic eliminations as durable results while finding the numerical obstruction value unreliable. It identified basis dependence, an unresolved external-check discrepancy, a repeated stall-then-breach cycle, and a pattern of opening new fronts before older proof obligations were completed. It also gave me a set of questions I could use without pretending to be the domain expert: Is this the same instrument as last time? What result would falsify the claim? Which promised obligation remains unfinished? What ends this line of work?

From that point, project governance became part of the experiment. Predictions were recorded before runs. Failure branches were defined in advance. Search families and correction cycles received hard caps. Later work had to name the proof obligation it served. The resulting constraints made the models' output more scientifically usable.

Capabilities and limitations of the language models

The project revealed a fairly specific division of labor. The models were strongest when the work could be made local, explicit, and falsifiable. Their reliability declined when they had to govern a long research program or judge how much a numerical pattern should change confidence in a theorem.

Their output became much more reliable when every important claim had to terminate in an exact artifact, an independent evaluator, a frozen manifest, or a reproducible command. Written explanations remained subordinate to those artifacts, including explanations offered in support of earlier results.

They excelled at
  • Turning dense equations and conventions into working software
  • Building symbolic eliminations, rank certificates, and exact checkers
  • Generating independent implementations and validation batteries
  • Forensic debugging from residual patterns and harmonic fingerprints
  • Exploring many candidate explanations and proof architectures quickly
  • Maintaining a detailed lab record, including corrections and retractions
They struggled with
  • Keeping a long campaign focused on the terminal theorem
  • Separating optimizer behavior from evidence about existence or nonexistence
  • Resisting scope expansion after an inconclusive result
  • Comparing numbers produced by changing instruments or normalizations
  • Calibrating confident prose to the actual strength of the evidence
  • Stopping without a human-imposed budget and failure rule

Transition to the full-equation search

Eventually the near-axis campaign had accumulated real finite-order structure without deciding the original question. It had derived two explicit fourth-order differential-polynomial necessary conditions, C1 and C2, and had characterized important exceptional structure. But it had neither classified every exceptional periodic solution nor shown that any formal series converged to a finite-radius field.

I authorized one terminal existence-first attack on the full, untruncated finite-volume equations. Before optimization, the project froze the operator, gauges, norms, grids, physical margins, resolution ladder, solver budget, acceptance thresholds, and stopping rule. Any numerical root would have served only as a candidate for a later computer-assisted proof.

Gate E0 specified the axis-regular operator and evidence rules. Gate E1 implemented and validated the full operator through symbolic identities, an exact axisymmetric control, near-axis regression, derivative checks, gauge audits, refinement tests, and an independently written residual evaluator. The maximum disagreement on the complete comparison vector was about 1.85 x 10^-13. This supported the operator implementation while leaving the existence of an exact field unresolved.

Outcome of the bounded search

The final search used one frozen finite-volume quasi-axisymmetric seed and four nested spectral resolutions. The independent hard residual decreased at every level, from 2.18 to 0.168, while the number of variables grew from 424 to 4,231. The reduction was real within the scope of those computations, and the candidate retained every required physical margin.

It was also nowhere near the preregistered gate. The final hard residual needed to be at most 10^-10; it was 0.168. The gauge residual and geometry tails failed by several orders of magnitude, and the total reduction was thirteen-fold rather than the required hundred-fold. At the higher resolutions, later optimizer steps continued to reduce the internal solve-grid cost while making the independent hard residual worse. Because the selection rule had been frozen in advance, the project retained the first genuinely improved step instead of presenting the optimizer's endpoint as progress.

Gate E2 failed. The charter prohibited a fifth resolution, a new seed family, a third solver, a changed metric, or a quiet return to the near-axis search. No computer-assisted-proof phase opened. The bounded construction attempt produced no certifiable exact field, and no conclusion about impossibility followed from that result.

ResolutionVariablesHard-grid pointsIndependent hard residual
R042422,5752.1816504628
R11,12572,2610.8167310164
R22,346166,7470.3601864501
R34,231320,4330.1677650357

Results retained from the project

Project Penelope left P1 open. Its strongest mathematical result is an exact, machine-verified finite-order reduction: within the stated vacuum, helicity-zero scope, any field that is exactly quasisymmetric through fourth order must satisfy two explicit compatibility conditions and their differential consequences. The repository also contains a validated full-volume operator and a reproducible record of one bounded search that failed its own gates.

The project also clarified the proper role of an LLM in research. The models compressed implementation time, proposed connections, maintained a large symbolic pipeline, and built tests around their own output. Their continuation incentives were poorly aligned with decision-directed research: they repeatedly proposed plausible new phases after existing lines had become inconclusive, without accounting for the cost of continuing or the declining value of the latest numerical result.

My responsibilities centered on defining success, distinguishing evidence classes, commissioning independent review, recording predictions before results, and stopping the project when the agreed gate failed. Those governance choices formed part of the research method and made the resulting record interpretable.

Conclusions

The project ended without a theorem. Its record still provides a useful account of a failed AI-assisted research program, including genuine technical accomplishments and repeated errors of interpretation. The central methodological lesson is that substantial model capability requires an equally explicit system for classifying and testing evidence.

The repository is public so that the exact artifacts, numerical records, changing interpretations, and final failure can be inspected together. Readers should begin with the README and final report, then use the long handoff as a chronological lab record. The code and certificates carry more evidentiary weight than the session prose or this retrospective account.

Research record

The article is an interpretation of the project. These are the underlying records.

Research archive
Project Penelope repository

The complete public archive of code, exact artifacts, numerical records, and project documentation.

Open
Outcome
P1 final report

The terminal scientific disposition: what was proved, observed, attempted, and left open.

Open
Independent audit
Progress assessment transcript

The mid-project review that separated durable structural progress from a misleading numerical headline.

Open
Research governance
Full-equation terminal charter

The preregistered gates, acceptance rules, budgets, and failure semantics for the final attack.

Open
Foundational paper
Garren and Boozer (1991)

Existence of quasihelically symmetric stellarators, the near-axis construction that framed the first phase.

Open