OpenAI's internal reasoning model disproves an 80-year-old Erdős conjecture
TL;DR: OpenAI announced May 20, 2026 that an internal general-purpose reasoning model disproved Paul Erdős’s 1946 planar unit-distance conjecture — finding constructions that beat the long-assumed-optimal square grid by a polynomial factor (n^(1+δ) pairs, δ ≈ 0.014). Verification: companion paper co-authored by four named external mathematicians — Fields medalist Tim Gowers, Noga Alon, Arul Shankar, and Jacob Tsimerman. Gowers called the result “a milestone in AI mathematics” and said he would recommend it for Annals of Mathematics. The credibility signal: one of the verifiers is Thomas Bloom — the researcher who exposed an earlier OpenAI false claim in October 2025. Status: not yet peer-reviewed in a journal. What’s distinctive: the model wasn’t math-specific — it was a general-purpose reasoning model, suggesting the math result emerges from broader capability gains, not from specialized fine-tuning.
What was announced
The reporting across OpenAI’s official announcement and external coverage confirms:
- Date: May 20, 2026
- Problem: Paul Erdős’s 1946 planar unit-distance conjecture — what is the maximum number of unit-distance pairs among n points in a plane?
- Result: The model produced constructions achieving n^(1+δ) unit-distance pairs with δ ≈ 0.014 — a polynomial improvement over the square-grid baseline that mathematicians had assumed was optimal for ~80 years
- Model: General-purpose internal reasoning model; specific name not publicly disclosed (likely an o-series successor)
- Verification paper: Co-authored by Tim Gowers, Noga Alon, Arul Shankar, and Jacob Tsimerman — all senior mathematicians in the relevant fields
- Peer-review status: Not yet published in a journal; verification by external mathematicians complete but full journal review pending
This is a substantive mathematics result, not a benchmark stunt. The Erdős unit-distance problem is one of the oldest open questions in discrete geometry, and the new construction settles a question that thousands of papers have tried to address since 1946.
Why the verification matters more than the result
The result itself is striking. The verification process is more important.
Thomas Bloom is on the companion paper. Bloom is the Cambridge mathematician who, in October 2025, publicly exposed an earlier OpenAI claim about solving Erdős problems as overstated — the original “solutions” turned out to be retrievals from existing literature rather than novel proofs. Bloom signing onto this companion paper is the strongest possible signal that this result is genuinely new mathematics, not a retrieval.
Tim Gowers is a Fields medalist. Gowers won the Fields Medal in 1998. His framing of the result — “a milestone in AI mathematics” and “I would recommend it for Annals of Mathematics without hesitation” — represents the highest possible endorsement from a working mathematician at the top of the field.
The companion paper does what AI-math papers usually don’t. It explicitly assesses whether the proof taught mathematicians something new about discrete geometry. The answer was yes — the constructions are interpretable, the technique is generalizable, and the result advances the field rather than just answering one specific question.
What’s distinctive technically
Three details from the OpenAI announcement that matter for understanding what’s actually new:
-
General-purpose, not math-specific. The model wasn’t fine-tuned on mathematics specifically. It was OpenAI’s internal general reasoning model — the same family that handles broader tasks. The math result emerges from general capability, not specialized training.
-
Autonomous, not human-guided. The model produced the constructions and the framing of the disproof without step-by-step human guidance. This is different from “the human gave the AI a clever hint and it filled in the calculation.”
-
Polynomial improvement, not constant. The δ ≈ 0.014 improvement is polynomial in n, meaning it doesn’t just shave a constant factor off the bound — it changes the asymptotic exponent. That’s the kind of result that opens a research direction, not closes one.
What it means for the AI research narrative
The story arc of 2024-2025 was “AI can do well on standardized benchmarks but struggles with genuinely novel reasoning.” The Erdős result is the clearest counter-data-point to that arc so far.
For ChatGPT users, this doesn’t change anything about what the consumer product can do today. The internal reasoning model that produced this result is not the same as the model behind ChatGPT, and the gap between “OpenAI’s internal frontier capability” and “what’s shipped in the product” remains roughly six to twelve months. But the trajectory is the relevant signal — what’s possible in OpenAI’s internal model today is approximately what could ship to consumers within a year.
For Claude users, the implication is that the research-frontier race is genuinely competitive. Anthropic’s hiring of Karpathy for Claude-accelerated pre-training and OpenAI’s substantive math result both point at the same underlying capability story: the frontier labs are now demonstrably producing novel research output, not just better benchmark scores.
How this pairs with the IPO narrative
The timing is unlikely to be coincidence. OpenAI filed its confidential S-1 with the SEC on May 22, 2026, two days after the Erdős announcement. A substantive research result two days before the IPO filing serves three audiences:
- Public investors: signals OpenAI has internal capabilities that compound over time, justifying valuation premium
- Talent: counters the Karpathy-to-Anthropic narrative that “the frontier is moving” — OpenAI’s internal frontier is clearly still producing
- Public discourse: shifts the conversation from “AI is a chatbot” to “AI is a research tool” right before the company asks the public markets for trillion-dollar pricing
Whether the Erdős result was scheduled specifically for this window is unknowable. The strategic effect is real either way.
The honest caveats
Two caveats this article needs to surface:
Not peer-reviewed yet. The proof has been checked by senior mathematicians, but it has not been formally published in a refereed journal. Journal peer review is the gold standard for mathematical results; checking by external mathematicians is the silver standard. The result could still be revised or have minor errors found in the review process.
The proof’s reproducibility is unclear. OpenAI’s announcement describes the result, but the model’s reasoning chain leading to the new constructions has not been published in a form that other researchers can independently reproduce step-by-step. Mathematicians have inspected the final proof and validated it, but the model’s discovery process is less transparent than a typical mathematical paper.
Neither caveat undermines the result. Both are reasons to wait for the full journal review before treating this as a settled milestone.
What to watch next
- Annals of Mathematics submission and review. If the paper proceeds smoothly through peer review (typically 6-18 months), this becomes a textbook result.
- OpenAI’s choice of next problems. If OpenAI announces additional Erdős-tier results in H2 2026, the pattern becomes harder to dismiss as a one-off.
- Anthropic’s response. Karpathy’s pre-training team specifically targets Claude-accelerated research. Expect a Claude research result of comparable scale within 6-12 months if Anthropic wants to match the narrative.
- DeepMind. Google DeepMind’s AlphaProof line is the most direct competitor in AI mathematics. Watch for a comparable result on a different Erdős-tier problem.
For broader context, see the ChatGPT review, the Claude review, and the head-to-head Claude vs ChatGPT comparison.
The framing
This is the most substantive AI research milestone of May 2026. The positioning week dominated by Anthropic’s enterprise consolidation had been narrating a one-sided story; the Erdős result re-balances it. OpenAI is filing for a trillion-dollar IPO and producing genuinely novel mathematics within the same week. Both stories are real. Neither cancels the other.
What changes from today is the burden of proof. The “AI can’t do real math” critique now has a counter-example with a Fields medalist’s endorsement. That doesn’t mean every future AI-math claim should be believed by default — but it does mean dismissive defaults need updating.
Sources
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.