HomeInsightsAI Strategy
AI strategy · 8 min read

An AI Solved Ten Unsolved Math Problems for $2,000. Now What?

In early August 2026, OpenAI announced that its unreleased Astra model resolved ten open problems in mathematics and theoretical computer science, publishing machine-checkable Lean 4 proofs for each at a reported compute cost of around $2,000. Some of the problems had resisted every serious attempt since 1946. For a small business none of this changes anything this quarter, and it changes the shape of the next few years considerably.

Paul Erdős spent his life travelling between universities with a suitcase, arriving unannounced, and leaving behind problems. Some were solved within months. Others became landmarks precisely because they were not, sitting in the catalogue for decades while very capable people tried and failed.

In August 2026, OpenAI announced that a model had resolved three of them, along with seven other open problems in mathematics and theoretical computer science.

The model is called Astra. It has not been publicly released. And the reason this particular announcement is different from the steady drip of AI capability claims is not the difficulty of the problems, though they were genuinely difficult. It is that nobody has to believe OpenAI about any of it.

What actually happened

The headline result, according to those closest to the field, is the first explicit construction of a non-sofic group. Mikhail Gromov introduced the concept of soficity in 1999, and in the years since, no mathematician had managed to prove or disprove whether non-sofic groups exist. Astra constructed one.

Alongside that, it disproved Connes's rigidity conjecture on von Neumann algebras, proved Ehrhart's volume conjecture, and resolved three problems from Erdős's catalogue. An earlier related result in May 2026 saw an internal OpenAI model disprove the Erdős unit distance conjecture, an 80-year-old problem in discrete geometry that had resisted every serious attempt since 1946, by producing an infinite family of counterexamples yielding a polynomial improvement.

Thomas Bloom, who maintains the Erdős problem catalogue, described the August results as big news and said they were more significant than the earlier unit distance result. That assessment from someone with no commercial interest in the outcome is worth more than any amount of company messaging.

Why the verification matters most

Here is the part that should genuinely change how you evaluate AI claims from now on, and it applies far beyond mathematics.

Each result was published as a machine-checkable certificate in Lean 4, a proof assistant with a small trusted kernel. Lean gives a binary verdict: the proof either compiles or it does not. There is no interpretation, no expert judgement, no room for a generous reading. Anyone can download the proof and run the checker. No PhD required.

Compare that to the normal situation with AI capability claims. A company reports a benchmark score. The benchmark was possibly in the training data. The evaluation was run by the company that built the model. The methodology is described in a paragraph. You are, in practice, being asked to trust a press release, and the history of that particular ask is not encouraging.

This is categorically different, and the reason is that mathematics has a property most domains lack: verification is cheap and total, while discovery is expensive and uncertain. That asymmetry is exactly the condition under which you can safely let a machine attempt something hard, because a wrong answer is caught automatically and costs nothing. It is worth holding onto that idea, because it turns out to be the single most transferable lesson here.

The $2,000 is the real headline

Reporting put the compute cost of solving all ten problems at roughly $2,000. That number deserves more attention than it received.

Think about what the alternative price looks like. These are problems that consumed the attention of trained mathematicians for decades. Measured in salaries, the accumulated human effort spent failing to solve the unit distance conjecture between 1946 and 2026 runs into millions. The model produced results for less than the cost of a decent laptop.

The trap here is concluding that intellectual work is now cheap in general. That is not what happened, and believing it will lead you to bad decisions. Astra was pointed at a category of problem with an unusual profile: precisely stated, self-contained, verifiable by machine, and unconstrained by any need to understand a real-world context. Almost nothing in a business has that profile.

What the $2,000 does tell you is where the cost curve is heading for a specific kind of work. When a task can be exactly specified and automatically checked, the cost of attempting it is collapsing toward nothing. That is a real and durable trend, and it has implications well before anyone needs to care about von Neumann algebras.

What this does and does not mean

The claim being made in a lot of the coverage is that AI has crossed from performing tasks to doing original research. That framing is defensible and it is also doing a lot of quiet work.

What genuinely happened is that a system produced results that no human had produced, in a domain where correctness is objectively decidable, and did so at a cost far below the human alternative. That is a real threshold and it should not be minimised out of discomfort.

What did not happen is that AI became generally capable of research. Mathematics is close to the ideal case for a machine: the problem statement is unambiguous, no data collection is needed, the search space is enormous but well-defined, and success is verifiable in seconds. Move to a domain where the problem is not precisely stated, where evidence is contested, where you must decide what to measure, or where being right depends on understanding people, and the conditions that made this work simply are not present.

There is also an unresolved question worth naming honestly, one that has been raised in the academic response: whether producing a verified proof constitutes understanding. Lean can confirm a proof is correct without anyone learning why it is true or what it connects to. That is a genuine question for mathematics. It is also a preview of a question your business will face, which is what it means to have an answer you can verify but cannot explain.

What it means for an actual business

The transferable lesson is not about mathematics. It is about the shape of problems worth handing to AI.

Astra worked because the problems were precisely specified and cheaply verifiable. Those two properties, taken together, are the reliable predictor of whether AI will do well at something in your business, and they are far more useful than any benchmark you will ever read.

Look at where AI genuinely succeeds in small businesses today and the pattern holds exactly. Data extraction works because you can check the extracted invoice total against the invoice. Classification works because you can review a sample and count the errors. Code generation works because the code either runs or it does not. In every case the loop is: specify precisely, generate cheaply, verify automatically, discard failures at no cost.

Now look at where AI disappoints. Strategy recommendations, judgement calls about a specific customer, anything depending on context that exists only in your head. These fail not because the model is weak but because there is no cheap verification step, so errors survive undetected until they cost something. This is the same underlying diagnosis we reached in why AI projects fail, arrived at from a completely different direction.

Want to find the verifiable, high-volume work in your business that AI is actually good at? A €49 audit identifies it.

What to do about it

The honest answer is that this changes nothing you should do this month, and it should change one thing about how you think.

Stop asking whether AI is smart enough for a task. It is a question that cannot be answered and it sends you to benchmark tables that will not help you. Ask instead whether the task can be specified exactly and whether you can check the answer cheaply. If both are yes, AI will probably handle it, and the specific model matters far less than people assume. If either is no, no model will rescue the project, and the failure will arrive slowly and be blamed on the wrong thing.

The practical version of this is to build a verification step before you build the automation. Most businesses do the reverse: they deploy the AI, then discover months later that nobody knows whether it is doing well. If you cannot describe how you will check the output before you start, you are not ready to start, and that is true whether the model cost two thousand dollars or two hundred billion.

Astra solved problems that had stood since before most of us were born, and it did it for the price of a laptop, because someone could check the answer. That is the whole lesson, and it fits in a sentence you can apply to your own work on Monday morning.


Sources

Quick answers

Common questions.

Want this in your business?

The €49 audit shows you exactly which automations would pay back fastest in your specific operation.

€49 entryFull AI audit + strategy call included

Reserve your auditNo commitment. No contracts. Just clarity.