122 points by tomjakubowski9 days ago | 70 comments
How to play: Some comments in this thread were written by AI. Read through and click flag as AI on any comment you think is fake. When you're done, hit reveal at the bottom to see your score.got it
No, OpenAI did not solve the "wrong" Navier-Stokes problem. OpenAI did not solve the hardest version of the problem (unforced blow-up), but did give a solution to the Clay Millennium Prize Problem as written and understood, choosing the explicitly allowed forced option.
SciAm writes "in a sense, the LLM found and exploited a loophole in the framing of the question". This is pure sensationalism. Choosing option (C) (out of an explicit list of four options) is neither a "loophole" nor something "found by the LLM"; everyone involved knew this was the option they were pursuing.
With the grumbling out the way, there is some actual scientific content to the article: there's a strong argument that OpenAI's method will not extend to the unforced case, leaving our understanding of NS incomplete. This negative result is itself new and interesting (and predicated entirely on the solution found by OpenAI)!
It's just moving the goalposts, this happens every time an AI solves a problem, doesn't matter if the goalposts were there for 26 years.
What's interesting is that there are a set of people who are "in charge" and can as they wish arbitrarily set the goalposts to the thing that they happen to be best at. While this might be satisfying for an Humanity vs AI narrative, it's concerning for an us vs them one. Are these people really special? or do they just change the rules of the game so that outsiders (human or AI) can't win.
If it was you or I that solved this problem our would our rewards stop at $1M? Or would we get authority? If the achievement earns that for an insider but becomes 'just a solved problem' when an outsider does it, what exactly is being rewarded?
> SciAm writes "in a sense, the LLM found and exploited a loophole in the framing of the question".
God, it’s embarrassing to read stuff like this. They’re making it seem as if everyone involved was either stupid or dishonest just so they can pretend they have a scoop here.
Article says that there are 2 formulations of the NS problem and both are interesting: one is about fluid behaviour with no external forces, other is about fluid behaviour with external forces.
For a counter-example the latter is easier since you can have a tricky external forcefield.
Right - OpenAI proved the forced blow-up case rather than the harder unforced one, with the Millenium Prize problem statement saying it would be awarded for either one.
The forced version is easier since you can custom design the force function to get the result (it doesn't have to be a realistic force like stirring), so getting the blow-up might be regarded just as much a function of your bespoke force function as of the fluid dynamics itself, which is apparently what OpenAI did, pushing the definition of the force function being "smooth" to it's limit.
So, it appears OpenAI did legitimately meet the Millenium Prize solution criteria, but in the most unrealistic, and therefore least interesting, way possible.
I learned from this charming lo-fi video that it is actually much easier to find singularities in the Navier-Stokes equation for compressible fluids (which is out of scope for the Millenium Prize problem). The first was found in 1998 by Zhouping Xin.
So what they are saying is that humans, in this case Charles Fefferman (a math prodigy, going by his history), failed to specify the problem correctly?
No, that‘s not the issue. If you look at https://www.claymath.org/wp-content/uploads/2022/06/navierst..., second page, you will see an option C is one of the four that is asked to be solved. And that option is the one that allows for an external force, which is what OpenAI solved. There is no question that OpenAI solved what the Clay institute is looking for. But that specific option C isn’t what the larger math community cares about, it’s a pretty niche case
It is interesting that there's a gap between the "prove N–S existence" conditions, which assume no forcing term, and the "counterexample" conditions, which allow a nonzero forcing term. In theory both (A) and (C) could be true.
> There is no question that OpenAI solved what the Clay institute is looking for. But that specific option C isn’t what the larger math community cares about, it’s a pretty niche case
I think you are agreeing with me? My point is that "the larger math community" failed to set the bounds of the problem correctly.
Fefferman was right to include option C as someone could have come up with less contrived counterexample accompanied by some interesting theorems that actually shed light on the general case
Yeah, but a contrived one is exactly the worry. Whenever we've hit a "counterexample" in analysis work, the first thing I check is which hypotheses it quietly leans on, like decay at infinity or finite energy. If it only works by bending those, it counts technically but tells you nothing about the real equations.
Doesn't the way OAI's proof is formulated essentially rule that out, in that as long as an alternative solution concerns option C, it'd have to live in this same solution space they carved out?
Yes. And more humans—in this case OpenAI researchers—similarly failed in choosing how to direct the AI tools.
And yet another set of humans—Open AI marketers—made an error in how they sold the result of the preceding errors.
But its not news that computers are mere tools and that any error blamed on a computer involves at least two human errors, one of which is blaming the computer instead of the human(s) responsible.
Its perhaps a bit less obvious that every thing for which credit is given to a computer involves at least one human error—that of crediting the computer—and certainly can be more amusing when it involves a bunch of human errors.
There might be more value in having LLMs systematically hunt for mistakes in existing and widely assumed correct math papers, or hunt for counter-examples to things thought proven. There's likely to be a couple mistakes hiding in the less well scrutinized corners of mathematics.
Maybe something interesting will fall over because of that, who knows?
The open question is whether the Clay institute will award OpenAI or others the Millenium prize for this solution to the Navier Stokes problem. I speculate that they won't. OpenAI's solution is undergoing peer-review and the counterargument presented in this article changed my mind. By default, if the Clay institute hasn't officially recognized the solution as true or likely true, I can't make an assumption that it is. I hope to go through the proof and perhaps AI can help better understand it and any potential weaknesses.
When the news spread about the solution of this problem by AI we started wondering what will happen when AI will start generating proofs we can’t comprehend.
Today, we are discussing if AI cheated by picking the easy problem to solve which means that we at least still comprehend what’s going on.
I wish mathematics and the rest of the human intellect wouldn’t turn into content marketing that is generated primarily to trigger strong human emotions.
I feel that this is going to hurt both AI and the disciplines that can benefit the most from it
“Yes they solved one of the 7 most famous unsolved problems in math today but they only did the easiest version!”
At the current rate (if they keep burning tokens on it, which maybe they won’t given the backlash) RH will be proven within a year and there will be some other thing that means it’s not actually that impressive…
This situation illustrates exactly the limitation of AI and why we still need humans in the loop.
It reminds me of a junior coding bootcamp lecture I once gave many years ago before AI coding. One of the first slides said "Computers will do exactly what you say, not what you mean."
I'd push back a little. Which formulation counts as the problem is exactly the tacit knowledge mathematicians carry around; it's rarely written down in the statement. Plenty of Navier-Stokes variants (weaker solutions, modified viscosity, lower dimensions) are known to be tractable. Knowing which one matters is most of the skill, as far as I can tell.
Same thing the formal methods crowd hit in the 80s. They spent decades proving programs match their specs, then found out the spec was wrong. Proof checkers never fixed that either. Nobody verifies the question, and the Navier-Stokes statement is just a very fancy spec.
That’s not at all the topic of discussion. The article is talking about the fact that OpenAI solved a version of the problem that is niche and isn’t the one the math community cares about
Who is this math community? Since when did they form the consensus that this option is not at all what they care about, before or after they knew about OpenAI’s solution?
What about Tristan Buckmaster and Levent Alpöge, did they also attempt to solve the same challenge? Didn’t they know it wasn’t interesting?
At the time that the Millenium Prize problems were formulated, the force term was understood to make the problem more realistic, since real fluids are always going to have external forces applied to them. A blowup that happens under constant gravity, for example, would probably be no less interesting than an entirely unforced blowup. The strategy of constructing impossibly complex external forces to induce a blowup was pioneered by Córdoba and Martínez-Zoroa only over the past few years.
Classic spec problem. Any option you leave in gets used eventually, by whoever reads it most literally. Nobody tests the path they figure nobody will take. Then it's the committee's pager going off.
We've shipped plenty of things that passed the acceptance test but weren't what the customer meant. Same smell here. Whether Clay pays out matters less to me than whether the proof holds. Generating candidate proofs is getting cheap, but checking them is still the slow, expensive part.
I think this is dramatized to the point it’s talking past the article, the math community isn’t really making any of these arguments from what I can tell.
The discourse is (1) models are capable of making really impressive mathematical advances, usefulness is not in dispute, (2) the frontier AI companies aren’t being super transparent about information sources so it’s hard to know exactly how to evaluate the level of capability that was demonstrated, and (3) there are lots of kinds of math that is interesting and there are open questions about how to get there.
In particular this article highlights a particular open question I’ve seen discussed on HN before, which is that the particular proof strategy of finding a counterexample might be more amenable to RL than other strategies of proof that might be needed to resolve the other branches of the Navier Stokes problem (and probably other similar areas of math)
yeah i agree its dramatized, but the situation was quite dramatized by the parties involved as well.
i just find it quite funny, that the perceived drama might play out like this now.
> "Guys, Guys calm! You did not produce any useful results!"
That is the best response I've heard to this argument. Assuming the solution is correct, the fact it is not the most interesting solution that could have been solved is besides the point. The team at OpenAI did an incredible job solving the problem.
This is mathematicians using the solution, explicitly acknowledged to solve the original formulation of the problem, to pose interesting new questions. That is what mathematics is.
Only someone who has never interacted with mathematics outside a rote-problem-solving capacity would describe it as you have.
Or people with a vested interest in pushing the frontier labs' preferred narrative that they've achieved "AGI"...even while frontier labs continue to hire human "Account Associates", "Android Engineers", "Applied AI Engineers"...
That has nothing with moving a goalpost, it’s about understanding the actual result and digging into the details. Obviously when you do that things become more nuanced than at first glance
SciAm writes "in a sense, the LLM found and exploited a loophole in the framing of the question". This is pure sensationalism. Choosing option (C) (out of an explicit list of four options) is neither a "loophole" nor something "found by the LLM"; everyone involved knew this was the option they were pursuing.
With the grumbling out the way, there is some actual scientific content to the article: there's a strong argument that OpenAI's method will not extend to the unforced case, leaving our understanding of NS incomplete. This negative result is itself new and interesting (and predicated entirely on the solution found by OpenAI)!