Open Weight Thoughts
All articles

· 6 min read

Every AI Lab Has Now Solved Mathematics and Peer Review Has 847 Unread Notifications

By L. Rodríguez

  • satire
  • guides

SATIRE — At 9:12 a.m. Tuesday, the Institute for Immediate General Intelligence announced that its new model had solved mathematics, a field previously thought to contain several subfields and at least one person named Euclid. By 9:14, five competing laboratories had issued nearly identical statements saying they too had solved mathematics, but with lower latency, a more permissive research-preview license, and a hosted endpoint that accepts up to 40,000 proofs per minute before returning HTTP 429.

The announcements mark a major advance for applied science communication. Researchers no longer need to wait years for a conjecture to be understood, checked, generalized, taught, challenged, or placed in the broad architecture of human knowledge. They can now put “Mathematical reasoning: 98.7%” in a launch blog post beside a chart whose y-axis begins at 96, whose x-axis is labeled “increasing transcendence,” and whose footnote leads to a repository containing a 17-gigabyte checkpoint, a half-finished evaluation harness, and the phrase “formal verification forthcoming.”

A guide to recognizing a solved field

Engineers encountering a claim that a model has solved mathematics should look for the standard indicators. First, the model must answer a selection of famous questions phrased as if they were support tickets. Second, the answers must be evaluated by another model, preferably one released by a different lab forty-eight minutes earlier. Third, all failures must be classified as “formatting sensitivity,” a technical term meaning the theorem did not survive being asked twice.

  • The paper’s abstract says the system “demonstrates emergent proof-native agency.”
  • The benchmark includes 600 problems, of which 597 are available in a public training corpus named DefinitelyNotTheBenchmark.
  • The model receives extra credit for producing a proof in Lean-shaped punctuation, even if the proof compiles only after removing the theorem.
  • A caption explains that scores should not be compared across models, hardware, prompts, temperatures, calendar months, or emotional states.
  • The appendix says independent replication is encouraged and provides a Dockerfile that requires a discontinued driver, 19 environment variables, and a small act of contrition.

The bottleneck: review, unfortunately

The only remaining obstacle is peer review, an artisanal workflow in which a small number of domain experts are asked to assess whether a 2.3-trillion-token machine has discovered a valid new branch of topology during a weekend deployment. In response, the Journal of Computational Certainty has expanded its reviewer invitation template. It now begins: “Dear Professor, based on your publication from 2008, we believe you are one of the last twelve people alive who can determine whether Section 7 is nonsense.”

According to internal estimates from the fictional Council for Responsible Acceleration, the global peer-review system currently has 847 unread notifications. This is not 847 unread papers. It is 847 notifications notifying reviewers that there may be papers, rebuttals, revised rebuttals, model cards, benchmark corrections, and a new supplementary appendix titled “Why the Original Claim Was Directionally Correct.” Each notification contains a link to a PDF rendered in seven-point font because the authors wished to include the complete token-level trace of their model discovering that division by zero is “a potentially high-variance operation.”

Reviewers have adapted as best they can. One common technique is to ask an AI assistant to summarize the AI-generated proof, then ask a second AI assistant to critique the first assistant’s summary, then ask a third assistant whether the critique feels calibrated. This process, known as automated epistemic load balancing, reduces a four-month review cycle to approximately three minutes plus the time required to discover that all three systems cited Lemma 4.2, which does not exist.

A new standard for reproducibility

To improve rigor, labs have adopted the Reproduce-If-You-Dare protocol. A result is reproducible when another team can obtain a similarly impressive answer by using the same hidden system prompt, proprietary synthetic-data mixture, undisclosed filtering pipeline, special inference-time search budget, and a cluster named after a fruit. This definition has the advantage of reflecting real-world engineering conditions, where every production issue is reproducible provided you have the exact request ID, a copy of last Thursday’s routing policy, and access to an employee who has since gone hiking.

Formal methods researchers have expressed cautious enthusiasm. They are pleased that the industry has rediscovered verification, though less pleased that it has rediscovered verification as a decorative adjective. A proof is not automatically formal because it has brackets. A theorem is not automatically established because a chatbot says “QED” with confidence. And a model is not a mathematician merely because it has learned that beginning a paragraph with “Observe that” causes readers to briefly lower their guard.

What engineering teams should do now

Until the journal inbox reaches zero, organizations should treat announcements of solved mathematics as they would any other dependency claim: with a lockfile, a test suite, and an escape hatch. If a model proposes an algorithm, run it. If it proposes a proof, check it with an actual verifier where possible. If it proposes that your service can be made 8,000 times faster through “nonlinear tensorized category compression,” ask for a benchmark script before asking procurement to reserve a mountain.

  1. Separate a flashy result from a deployable capability.
  2. Ask what exact task was measured, what data may have leaked, and what independent checking occurred.
  3. Prefer artifacts that run over narratives that merely radiate competence.
  4. Record prompts, model versions, tool access, and evaluation conditions; otherwise your result is a campfire story with JSON.
  5. Assume that a model can be useful long before any claim about general mathematical conquest is settled.

There is one unglamorous fact beneath the launch confetti: machine-assisted mathematical work can be genuinely valuable when it is paired with clear tasks, reliable verification, and humans who understand what is being claimed. The hard part was never getting a system to produce symbols. The hard part is still knowing which symbols deserve our trust.