Terence Tao Warns AI is Depleting Open Math Problems
Top mathematician Terence Tao says AI is non-renewably mining open math problems, signaling a massive data wall for reasoning models.
Research · Source: Hacker News
What happened
Top mathematician Terence Tao recently posted a stark warning on Mastodon. He stated that open math problems are being non-renewably mined by artificial intelligence. Full text and specific details from his post are currently limited. However, the core message is impossible to ignore. AI models are burning through our historical supply of unsolved mathematical proofs. These problems were meant to challenge human minds for centuries. Now they are being consumed by algorithms in mere months.
Math has long served as the ultimate benchmark for AI reasoning. Leading AI companies use these open problems to train and test their most advanced models. Tao suggests this is a strictly finite resource. We only have so many high-quality open problems available. Once an AI solves a problem, that problem becomes part of the training data. It immediately loses its value as an independent benchmark for future models. The well is running dry.
This creates a unique bottleneck for AI development. We cannot simply generate historically significant math problems on demand. Human mathematicians take decades to formulate the right questions. AI models are consuming them at an unprecedented rate. The industry is treating a precious scientific resource like cheap fuel. We are destroying the very tools we need to measure future progress.
Key facts
- Terence Tao — The mathematician who warned that AI is non-renewably mining open math problems.
Why it matters
This changes exactly how we evaluate AI reasoning capabilities. If models consume all existing open problems, we lose our best yardstick for measuring true intelligence. Builders rely heavily on these benchmarks to prove their models can actually think rather than just memorize patterns. Without fresh open problems, benchmark contamination becomes impossible to avoid. You will never know if a model solved a problem through logic or because it saw a similar solution in its massive training run. The trust in AI evaluation metrics will collapse.
The second-order effect is a massive forced shift in AI training strategies. Labs will have to stop relying on historical human genius for their training data. They will need to build entirely new synthetic environments. In these environments, AI must generate and verify its own novel problems from scratch. The companies that figure out how to create infinite verifiable reasoning loops will win the next era of AI. Everyone else will be stuck fine-tuning on depleted human data.
For builders
Build synthetic reasoning environments
The reliance on human-made math problems is ending fast. Founders who build systems that generate verifiable synthetic logic puzzles will find eager buyers among top AI labs. The big players will pay massive premiums for clean data.
Create new dynamic benchmarks
Static benchmarks are completely dead. Engineers need to build dynamic evaluation tools that generate novel problems on the fly. You win by helping labs prove their models are not just regurgitating memorized proofs.
Pivot from scraping to generation
Scraping the internet for reasoning data is a dying strategy. The future belongs to those who build self-play architectures. You need to focus on systems that create their own data exhaust to survive the coming drought.
My take
I have been warning founders about the data wall for months, and this proves it is here. When a mathematician like Terence Tao says AI is mining math non-renewably, you have to listen. Stop building wrappers around static datasets. If you want to survive in this space, you must build systems that generate their own verifiable truth.
Original reporting: Hacker News. This is my rewrite and opinion.