Spare a thought for the mathematicians

Eight days after I bet that math was next, OpenAI posted 722 math papers written by an unreleased model, covering 372 results across 17 fields. I have mixed feelings.

Some of what it claims, with the year each question goes back to:

  • The quasi-Riemann hypothesis: the Riemann zeta function has no zeros with real part above 7/8. It’s a weaker cousin of the Riemann hypothesis (1859).
  • The Unique Games Conjecture (Subhash Khot, 2002), which means the classic approximation algorithms for problems like Max-Cut and Vertex Cover are already as good as they’ll ever get, unless P = NP.
  • Catalan’s constant is irrational: 1 - 1/9 + 1/25 - 1/49 + …, a number Eugène Catalan wrote about in 1865.
  • The Mahler conjecture (1939), about the smallest volume product of a convex shape and its “polar” dual.
  • Hilbert’s tenth problem over the rationals: no algorithm can tell whether any given polynomial equation has a rational solution. Hilbert asked about whole numbers in 1900, and that version was settled in 1970.

The first four come with a proof in Lean, a programming language where a computer checks every step, and 162 of the papers have their main result checked that way. Hilbert’s tenth isn’t checked yet. Each result took about 3 hours of ChatGPT Pro thinking on average, out of about 4,000 problems posed.

In that same post I called OpenAI’s Navier-Stokes proof a sign of acceleration. I didn’t write about the people. Diego Córdoba and Luis Martínez-Zoroa had developed “forcing”, the approach the model used, and Tristan Buckmaster and Levent Alpöge posted a blowup result for the related Euler equations the night before OpenAI’s announcement. “We’re a little bit in shock,” Córdoba told Scientific American.

I haven’t thought deeply about pure math in a long time. I’ve spent about 18 years building software, and I’ve always been more of a fan of math that does something in the real world. But I remember the joy of coming up with a new conjecture, or a proof nobody showed me first. In January 2005 I asked this blog for help because it took me 3 hours to solve 3 problems in my first proof-heavy course. The model averaged 3 hours per open problem.

I’ve been mostly positive about AI in software engineering, even with the short-term disruption, because of the Jevons paradox: when something gets cheaper, we use more of it, and cheaper code means a lot more software. Math might not work that way. A conjecture only gets proven once. If you’d spent 10 years on the Mahler conjecture, Tuesday took that problem away, and no amount of new demand gives it back.

I hope math as a form of human expression and creativity at the frontier keeps going, even if it looks more like cyborg exploration: people choosing which questions are worth asking and what the answers mean, with a model grinding through the cases. Maybe Jevons applies after all, and cheaper proofs mean more questions worth asking. Someone also has to read and understand 722 papers.

Back to writing