Often people who are critical of whether AI can lead to research breakthroughs are experts in the respective area, who nevertheless are afraid of their future career prospects in academia (getting a permanent position in academia is hard and it is deeply political who gets such a position).
This people are thus not scared by AI per se (it's basically their daily job to devise innovations that advance their field), but their fears are that
- because of the hype around AI the research into which they invested years, often decades, will be considered "unimportant",
- incompotent people in decision-making positions will think researchers can be replaced by AI.
It's still a bit weird though. For obvious reasons, The MO for benchmarks like this has been that the models remain pretty low until suddenly it's done. So even for a 'should i be worried yet' reality check, it's pretty terrible. You can't really keep track of what models are actually able to do. By the time you can replace years of research by specialists with a few api calls then...
Actually, years of research certainly can be emulated by AIs, in the sense that they have many scanned research papers and can cobble together a solution to many things, say unsolved math problems, from this. So baseline they can seem pretty great. The article is indeed just reinforcing the point that their innovations are less impressive when it's something that really no one has done before. It's not a trivial point for who are used to AI doing so many impressive things.
No, it did not. Do not trust the AI propaganda, a match researcher USED AI as a tool to solve unsolved math problems, the AI by itself didn't do anything. They are two entirely different things. If the match researcher didn't give the AI the prompts he gave the AI wouldn't have find any solution.
Yes, we could say that AI is an useful tool to use in math research (the same as calculators, and then computer, are) but no, AI did not solve any problem.
I think this sort of small scale research on this problem is inherently pointless and will be a lagging indicator of diffusion, not a leading indicator of capability.
The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result.
If the AI could do this task we would see this happening in places where the economic incentives let them spend millions of dollars on this problem, not on an eval like this.
This specific form of eval where you just ask the agent to solve it with no specific scaffolding besides GPU access (e.g. nothing like AlphaEvolve, ArchPilot, etc that try to work around model shortcomings) is also going to further trail what is possible at small scale. It's good that we at least give them execution environments now, but this feels like the experiments that were worked on figuring out how to get LLMs to do native arithmetic rather than just giving them a calculator/python env.
Regarding the Navier Stokes theorem, one of the best posts I've seen about this (and the other OpenAI math proofs) recently is Nestor Guillen's post where he introduces "the convex hull of ideas". I love this because I've had some vague ideas about the limitations of LLMs that aligned with this, but Guillen really clarified the idea and explained it with a great analogy: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...
Basically Guillen is arguing that LLMs are great at finding results within the "convex hull" of existing literature (i.e. their training data). They can make connections across different parts of the literature where it would be impossible for a human to be an expert in all these areas. But when it comes to truly novel, original ideas, there is no proof yet that LLMs are able to go there. That's honestly a limitation I'm really rooting for because otherwise I think the future of humanity is generally fucked.
Love that link. Thank you. It's helpful to find others thinking this way :)
I work in collective intelligence, and been thinking about what surprise is, and why the ability to sense it is baked into each and every one of us. It's a feeling that moves our attention
I've been riffing on the idea of "mapping" collective surprise by gathering information on which statements give us the feeling. The feeling of something being dropped just beyond the "boundary of our knowing", which I think of as a shell we are each building through our models, like something pushed out into semantic space. I'm saying it funny, but it's really just about collecting data on what surprises people.
And in aggregate, maybe you could have a collective map of this shifting thing (really a map of the education system in the sweeping life-long sense), of the whole human cultural swarm. Surprise is the what we each feel as we cross a threshold that we can't easily return over. We forget today what surprised us yesterday. But maybe it's possible to learn something about the structure we build together (in knowledge), if we gather little radar blips on when we are passing over such moments.
Like if a collective over the course of a week, all has a moment of surprise as they learn a specific thing. If you could see that happening to your group, we could all know that we've all integrated (or at least encountered) something that can be built on.
I have heard that before, it's essentially the LLM cannot be creative packaged differently. But it always falls flat because there is just no way to measure it let alone define it. In the end it always give me a vibe of them wanting that to be the case rather than it being it case.
> But it always falls flat because there is just no way to measure it let alone define it.
I disagree. The definitions may have some gray areas, but at least looking back in history there are some breakthroughs that would fall firmly within the "a result found within the convex hull of ideas that existed at that time" type (e.g. Einstein's special theory of relativity), but also some truly novel, "out of the blue" breakthroughs that people agree are fully outside the convex hull of ideas from that time (e.g. Einstein's general theory of relativity).
Based on that, mathematicians and scientists should be able to provide some falsifiable hypotheses about the theorems and kinds of proof methods that would be truly novel.
Before the convex hull, we had stochastic parrots.
Search is an incredibly powerful method in both human and machine endeavors for generating new ideas and we will see Move 37s in math soon enough, even if we have not already.
It almost seems like people don’t understand that most of the training of modern LLMs is now reinforcement learning on problem-solving tasks. They spend most of their training literally solving problems that they don’t already know how to do.
It blows my mind that LLMs are solving Millennium Prize problems and people are still trying to argue that they can't do original thinking.
> "This 100% deadly airborne superbacterium is inspired by natural bacteria, it's not truly original thinking," I say, as the superintelligent AI kills me and everyone else in the world
You should read the Guillen's post then, because that's not what it argues. And very importantly, the only Millennium Prize problem that OpenAI has claimed to have solved is mired in controversy because human researchers were on the cusp of releasing their results (and we're likely training OpenAI in the process). This is exactly the kind of "thinking", along with the other results in the OpenAI math dump, that fall firmly within this "convex hull" of ideas.
And also importantly, this framework provides the basis for some falsifiable hypotheses. That's very different from the AI boosterism that just seems to be saying that since AI capabilities are expanding at a breakneck speed that that speed will continue forever, as if a 13 year old will eventually be a hundred feet tall based on their recent growth rates.
I've wondered how; "The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result." can possibly not also include many dollars worth of duplicated effort? I find it hard to believe every agent was doing something novel. What this means beyond wasting dollars I'm not sure of, but just flinging around the numbers I don't think should be considered impressive or even required to get the results found.
there is no shortage of so-called studies about LLMs that inexplicably use comically old and bad models like gpt-4o to reach a strong conclusion about the ineffectiveness of LLMs in general
one cannot help to suspect that this specific kind of academic malpractice does in fact have a point which is to undermine the technology as a whole out of fear, to keep the academic "consensus" in favor of tenured academics
> there is no shortage of so-called studies about LLMs that inexplicably use comically old and bad models like gpt-4o to reach a strong conclusion about the ineffectiveness of LLMs in general
> one cannot help to suspect that this specific kind of academic malpractice does in fact have a point which is to undermine the technology as a whole out of fear, to keep the academic "consensus" in favor of tenured academics
Serious studies often take a lot of time to do. Also, models outdate very fast. So, of course nearly every serious study on such a topic will be based on models that are considered "outdated" by LLM aficionados.
So, no need to invent conspiracy theories that such studies are made to keep the academic "consensus" in favor of tenured academics.
In my experience R&D has basically two axes: how innovative it is, and how well we can measure the results.
For the quadrant of non-innovative tasks where we already have a good way to measure performance, Claude can handle this. There is very little ambiguity, and we are basically just looking to maximize some metric under a set of constraints.
Many business processes are not like that. They might be conceptually simple, but it isn’t that easy to say whether a system has done a good job or not. I would say that LLMs can help with this a lot but they have bad judgement because it requires talking to people.
And the other, perhaps more rare issue is in problems where there is data but actually modeling it to sufficient quality or fast enough is hard.
"See this post in the app
Use the app to view all comments and discover more posts."
Not to be too contrarian, but is there a solution which doesn't require me to buy a new phone? The link refuses to divulge relevant information on mobile browsers, and the app it's telling me to install isn't backwards compatible even a few years.
Separately, why is that the link to "the man"? Their Twitter profile isn't linked in HN, and I would've naively thought that asking them a question in reply to their HN comment would be a reasonable way to ask them about that HN comment.
This comment is a joke about how Boris Cherny (creator of Claude Code) boasted that “coding is solved” and immediately received the expected mockery/screenshots of bugs in Claude Code. He replies to one of them “Coding is solved, bugs are not yet solved. Fix incoming”
Imo, self-driving is massively held back by the cost of hardware. There are fully autonomous robo taxis. That human taxis are still commercially viable is in large parts because robo taxis are quite expensive.
Coding cookie cutter CRUD apps is pretty much solved at a much lower price point than artisanal software.
Give me one example of a fully vibe coded software with zero human input that is running in production with non trivial amount of actual users that have been running for more than a year....
Coding is solved for the kind of problem you would have hired a team of 3-5 juniors a couple of years ago to build something for your small to medium sized business.
Practically every paragraph is negative, with sentences like "Agents made misleading claims about their work," and "A natural question is whether the agents could have improved with larger GPU budgets. Although both improved across their runs, in the case of Fable the improvements were almost entirely due to attempted cheating." and "For Sol, the answer is less clear-cut; it did make some progress, although its method was fairly incremental and had limited applicability to the coding task. This suggests that we should be pessimistic about further GPU spending," all reinforcing the fact that LLMs aren't currently capable of this.
My own view is "No, of course not, in fact they will never be capable of achieving good results with that technique." Because that technique will end up training the LLMs on their own output and lead to the inability to distinguish reality from hallucination. If you think I'm wrong about that, I'd be interested in hearing why.
A bit too pessimistic imo. I agree that AI can’t automate things end to end, but a good deal of R&D involves kicking off a training run and babysitting it.
If your training run dies at 1 am and you’re sleeping, you won’t find out about it until the next day. You can lose up to 18 hours of work depending on when it happens. Based on the error it might be as simple as tweaking a single hyperparameter and rebooting, which is something LLMs are usually capable of.
Even just that task means I can kick off multiple runs over the weekend and have confidence they’ll finish. It’s a game changer.
I'd classify that as an entirely different category than AI self-training. What you're describing could have been done with a short script, though which parameter to tweak and how to tweak it would be difficult to automate with a non-LLM script, so the LLM's being able to parse the error message and base the tweak on the content of the error is a definite improvement to the process there.
But I'd classify this as LLM being used to automate a sysadmin task, rather than calling that self-training.
Yeah I’m not trying to argue it is AGI, but it’s not as simple as a short script. There’s some amount of debugging involved, and no amount of if-statements could cover all possible ways a script could break.
In a way, “recursive self improvement” just means tools helping us to create better tools. At least that’s what the words mean.
the GP posits the problem with "recursive self improvement" is the poisoned context & hallucination problem. While a strong loop that has fail safes, backups, restore points (essentially, a fancy backup system), the problem isn't that we cannot create a healthy advanced wiggum loop; it's that every step of the LLM as it grows whatever knowledge is acretes, has a chance of being either poisoned (eg, it conflates two tokens as describe different things) or wholesales fabricates a method or procedure.
Now humans are just as bad, but they're not moving at the speed of compute so the posion and fabricates can dissolve over time, or just, as you've noticed turning on your news, get stuck in very stupid positions. So humans are clearly capable but clearly don't tend to do this either.
So then we dont have a real road map. The error rates, although small, acrete at exponential levels and will wash out improvements.
So I also had the idea that "if we just give it enough context, surely it'll be more powerful". But the error rates hit that squarely. The larger the context grows, the more likely it hasn't properly organized its knowledge to avoid overlapping facts.
In programming, it's worse, because a lot of the code is purposefully "DRY" and reuseable. Everye C program has a main(); is it remembering the correct main? or any of the number of same variables?
You can see an LLM is powerful but it's not ominipotent. It'll suffer very much when it starts hallucinations and context poisoning.
So, sure you can try a super ralph wiggum loop with memory, fallback safeties, etc, but you basically then need another turtle that does the same thing, and at that point, you're positing a infinite jest of ralph wiggum loops tracking each other, recursively, forever.
This was an interesting post on the subject that more or less agrees: 'self-improvement' is happening with things like this but there's a lot of headwind on any 'hard take-off' where capabilities grow exponentially all on their own:
We don't have to prove you wrong, you have to make a case for your position. Your statement was vague and handwavy and could be countered with another vague statement such as "They will add a feature that allows the LLM to better detect hallucinations".
I phrased it that way because I couldn't remember the term "model collapse" at the time. But that's what I meant: that making the LLMs self-train will lead to model collapse.
There; now the statement is far less vague and handwavy, because I'm making a specific claim that is, AFAIK, well-understood.
Also, you seem to have misunderstood me a little. I didn't mean "prove me wrong", I meant "If I'm missing something, please tell me about it." More of a conversational request than staking a claim in an argument. Many people at HN seem to like to take argumentative, debate-competition stances — but I usually prefer more "Hey, let's discuss this interesting idea, point out mistakes each other is making, and learn together" kind of interactions. That's what I was asking for.
Well, I guess the answer isn't too different. I believe model collapse is a limitation of the current AI tech but maybe not the future ones. You can see humanity as a huge model that trains itself. What is novel about AI is that we built a machine with some intelligence traits that is free of biological constraints. If we can emulate the aggregate intelligence of a civilization inside a machine, it could improve itself forever but at a much faster pace.
You couldn’t remember the term model collapse because people quit saying it shortly after it was invented in 2023, because it was just cope by people praying for that outcome. Internet will have LLM output, therefore models will train on that output, then kaput. And what happenend next? Models trained on model output on purpose and just keep getting better.
Because - it's the part people are having hard time to accept - creativity was never the hard part. It's randomness. The hard part is navigating the narrow border between "not enough random, therefore boring" and "too much random, therefore unhinged/insane".
That border is determined not as much by ground truth, as by people's sentiment. As such, many famous artists and creatives never succeed navigating - their contemporaries would call them crazy, and then decades or centuries later, the border shifts, and suddenly they're remembered as greatest creators of their era.
That is not true at all. As I previously stated: reference, stealing well, blending styles in new ways that obey or break the rules of your art form… what appears random to outsiders is a story that follows as naturally as 2 after 1 if you understand the context.
Until an AI can have desires or experience life it cannot tell a story.
Even if it could it is questionable if such alien experience qualifies as story.
It can retell stories. It can swirl around extant media like a child playing with its dinner. It cannot blend in its own lived and new perspectives which is what gives new art value.
> what appears random to outsiders is a story that follows as naturally as 2 after 1 if you understand the context.
It's interesting to compare your opinion with Andrew Wiles quote:
> Perhaps I could best describe my experience of doing mathematics in terms of entering a dark mansion. You go into the first room and it's dark, completely dark. You stumble around, bumping into the furniture. Gradually, you learn where each piece of furniture is. And finally, after six months or so, you find the light switch and turn it on. Suddenly, it's all illuminated and you can see exactly where you were. Then you enter the next dark room...
This is analogous to “I saw Abe Lincoln at the hall of presidents at Disney.” and insisting it was really Abe Lincoln.
It is infantilizing to suggest this is originality or creativity just as much as it is to suggest David copperfield is a scientist.
It is an illusion. A subterfuge being performed by the puppeteer. It works due to the theft of human’s work.
Ask yourself,
"What are this creative engine's sensory inputs to experience life?"
"What can this creative engine tell me about the experience of life from its own direct inputs?"
"How does it blend/shape or re-interpret raw input direct sensory data with the, admittedly, astounding body of knowledge it has about creative work?"
"What form of change or growth has this engine experienced that teaches us about ourselves?"
If you take those questions seriously you will come to the same conclusion. It is a Roomba smearing the valuable creative work of millions like dog shit across the living room floor, not a consciousness experiencing the agony and ecstacy of life.
If you think I am wrong ask an AI. It won't come up with an original answer, but it can certainly tell you whether the collective body of slightly-outdated human knowledge thinks I am correct.
Yea, I think LLMs would be much less powerful if not for the inherent randomness in the GPU...
It is like a decompression program that results in a slightly different result each time it is run on the same compressed data. If you are lucky, the delta could end up with answers to all your questions..
In practice, hallucinations have only mitigated within verifiable domains. The vast, vast majority of human activity is unverifiable, or at least significantly less verifiable than mathematics.
I recently did a lower tech "write an async divide and conquer S3 lister in Python" and the results were fairly lackluster.
Opus 5.5 and GLM 5.3 Flash came out with pretty reasonable code that was fairly fast. However, all the models left aioboto defaults which have a 10 connection pool size so they ended up all artificially capped even though they added semaphores with higher limits.
My attempt I spent a couple hours AI assisted on ended up being faster, using less CPU time, and having roughly half the code.
Give it a verification loop and enough compute, and AI will soon cook algorithmic R&D like it cooked mathematics. There is just no question whatsoever of this happening, it is guaranteed.
And there are so many indicators that it's not currently economically viable and will only be if retail price is massively increased and strong regulatory barriers to competitors are erected. ;)
Have you seen my drafts? For easy task I may write the straighforward solution, but for complicated stuff it's a mix of throwing stuff to the wall and see what sticks. It's informed search by experience, but AI also use weights to pick the attempts.
AI is not a calculator. OpenAI did not solve hundreds of open problems in math by "brute forcing the search space," there was a lot of intelligent and novel thought (gasp!) in the way AI chose the paths it did.
Yes, AI can explore thousands of ideas at once, no, that's not "brute forcing the search space" because the search space is way larger than you think it is.
Brute force would choke on a measly decillion of alternatives. It corresponds to an extremely short proof. A combinatorial explosion doesn't allow to use brute force for anything slightly longer than toy problems. Swarm of 10000 agents is a massive smart force.
Several of the items on that list are claimed to have been resolved by the recent OpenAI theorem-dump. Another is the Navier-Stokes question that's one of the Millennium Prize problems, also claimed to have been resolved by OpenAI's models.
So: unless all those claims by OpenAI turn out to have been mistakes[1]: no, actually, those are not all still there.
[1] It's certainly possible that some will. They've already retracted a few things.
There is already a pre-print making the claim that the Navier-Stokes solution provided by OpenAI does not meet the criteria required by the prize (https://arxiv.org/html/2609.20803v1)
I'm not a mathematician, but looking at the Clay Institute's official problem statement [1] and that paper, I don't think it's claiming OpenAI's construction doesn't meet the Millennium Prize criteria.
The criteria for statement C (which OpenAI targeted) only require a C-infinity smooth force, whereas the paper considers constructions of the same kind as the OpenAI one but with a real analytic force. All real analytic functions are C-infinity smooth, but not vice versa [2]. Wikipedia says the Navier-Stokes problem for real analytic forces is still unsolved, so the question of whether a similar approach can be used there is presumably of some interest, but it doesn't appear to affect whether the Prize criteria have been satisfied and the paper doesn't explicitly make any claim that it does.
Re the problem still showing as "active": The Clay Institute don't consider problems solved until the solution has been fully digested by the mathematical community and is well established. In the case of the Poincare Conjecture that wasn't until 2010, 7 years after Perelman put his papers on the arXiv. Nobody is expecting them to mark Navier-Stokes solved any time soon.
Your comment was the first I've heard of that paper, BTW, so thanks for that. However, that's another indication that this probably doesn't bear on the Millennium Prize criteria. The lead author, Peter Constantin, is a leading expert on the problem and the Clay Institute's page on it has a video of a lecture by him. If he had come out and said OpenAI hadn't solved the problem as stated I'm sure there would have been a lot of noise about it. The paper is dated 17th of September so presumably there has been plenty of time.
It is pretty straightforward for anyone who knows lean to verify that the OpenAI release satisfies the official criteria to the letter. You could reasonably argue that it shouldn't satisfy the official criteria because the official criteria are unreasonably generous in allowing any smooth external forcing function, but that would be a different claim.
Theoretical fields like mathematics inherently suffer from a form of Goodhart's Law. The progress in mathematics you hear of is almost never the kind of progress non-mathematicians would care about.
It's always easier to measure progress according to the internal metrics of the field than to evaluate the contributions to the wider understanding of the topic. In theoretical computer science (which I'm most familiar with), people have long complained about results focusing on shaving sublogarithmic factors from complexity bounds (while making the algorithm worse in practice) and about reviewers being impressed by the technical difficulty of proofs. But progress like that is easier to measure than new algorithmic ideas or conceptual understanding.
But it's not all bad. The researchers chasing the metrics are almost always genuinely interested in the topics they study. Their actual contributions mostly come from the ideas they explore while trying to achieve measurable progress. And because they are not expected to produce anything of direct value (Goodhart's Law for applied researchers), they can explore a wider range of ideas.
I personally noticed this when I moved from theoretical computer science to algorithmic bioinformatics. When I start a new project, the expectation is that researchers in genomics should be using sofware that uses the new algorithms five year from now. That expectation is useful, but it's also a strict constraint on what I can afford to try.
The only outcome I see is that AI will become so complex that we won’t be able to rely on it for R&D, because we won’t be able to measure and prove its results. Some AI collaboration and speed of work will be impossible for humans to replicate, making it neither good nor bad—just something beyond human ability.
Could it be that the companies have nerfed the models on these domains? It is a very hard thing to do because it can hurt related domains. But its not beyond the ideology of Dario - he tried it publicly .
I think the more interesting thing is was it unable to do it even knowing the solution? I feel like only testing a single innovation is biasing the current state of automated AI R&D.
I don't buy this whole narrative that "AI is innovative".
If AI is so innovative then why am I seeing a dumb idiot every time I discuss anything even remotely innovative with it.
Most often it cannot even produce a true sentence when the claim in the sentence is modestly strong without injecting qualifiers and distancing itself from the claim.
Having to just accept that it produces relentless innovation is so detached from reality that it is not even funny.
It is true that it can write decent code once the domain bounds are well defined. What I have found is that it is good at scrutinizing already written code - but only with expert supervision. Even then it most often tests our patience with stupid suggestions.
> If AI is so innovative then why am I seeing a dumb retard everytime I discuss anything even remotely innovative with it.
My experience slightly differs: yes, LLMs are not the best entities to discuss innovative scientific ideas that one has in mind, but they are still often much more competent discussion partners on such topics than the people you are typically surrounded with (e.g. work colleagues).
(no, this is not a template for a Reddit /r/iamverysmart/ post :-) , but just my experience).
I agree. Your experience mirrors mine. Yes, they are often competent discussion partners, although a lot of people I hang around with are also excellent discussion partners.
But unlike with real people, LLMs have this unique property that we can abuse them without consequences directly leading to our improved understanding of the subject material using nothing but pure frustration as the motivating drive.
So for me, they are productive precisely because they remain dumb discussion partners.
>If AI is so innovative then why am I seeing a dumb idiot every time I discuss anything even remotely innovative with it.
Also, if AI is so innovative why frontier labs still struggle to come up with a money printing product? It's everything so confusing: They have AGI, but their product is an app/API, which is so world-ending dangerous, but their latest product is a set of cute agents running on your own machine, but also is so smart it can do the work of a PhD on almost every field of human knowledge, but most ot they have to show is in something so machine specific like math proofs.
AI is innovative but cannot string together true confident sentences.
AI is so innovative that nothing innovative or groundbreaking has so far appeared. Maybe it will, but the claims to ground-truth ratio is off the charts.
World ending is so funny it is usually only parroted and swallowed by gullible people.
> it can do the work of a PhD on almost every field of human knowledge
This is the most revealing and the most hilarious one. This is essentially the argument from authority wrapped in a PR stunt. Those who produce real innovations overwhelmingly does it outside curriculum and as part of the real work they are engaged in with very few historical exceptions. That statement reveals a lack of understanding of what innovation itself means let alone the virtue signalling they are trying to sell.
> machine specific math proofs
This one was initially intriguing, but as I earlier said, is a well defined and constrained bounded domain. The accusation that hours earlier other human researchers's work may have been used to plagiarize the proof has muddied the waters further.
>This one was initially intriguing, but as I earlier said, is a well defined and constrained bounded domain. The accusation that hours earlier other human researchers's work may have been used to plagiarize the proof has muddied the waters further.
It's even worse. They claim, and the media repeat, that new math results are product of the models intelligence. But we all know there are elite mathematicians working for these companies (OAI, Anthropic) and basically are they who are coming up with new results, assisted by AI, not the other way. And we still don't know how much of the heavy lifting is done by each part. If the models are PhD smart, shouldn't be able to come up with results with minimal guidance? If they're so innovative why the need of an army of top mathematicians, engineers, physicists, etc. and thousands agents and millions in tokens to do novel work? It's just seems to me that humans are doing (barely, hence the plagiarism) the creative, innovative work and AI agents are just useful to explore a massive space of bullshit hypotheses to find something that finally sticks, after being validated again by an army of human experts. AGI and the near super intelligence? It doesn't track.
It is entirely plausible that, as you say, they could be piggybacking off of human work internally and using AI like the punching bag it is - dumb but useful. I wouldn't bet against that, but usually it is difficult for people who did the hard intellectual work to have their credit taken away - which might be exactly what we saw with Navier-Stokes drama that unfolded.
> If the models are PhD smart, shouldn't be able to come up with results with minimal guidance?
Smart is born out of wrestling with the problem domain, not in virtue signaling via academic credentials. So anyone leading with an academic credential is a dead giveaway of the PR stunt underneath. Nobody says PhD smart Einstein - because work can lead the speaking and doesn't need a credential as a support crutch.
> AGI and the near super intelligence? It doesn't track.
True ideal intelligence must not hedge, shouldn't pander to popular opinion or consensus. True intelligence must state what is true as maximally true as is correct while being able to derive more true sentences coupled with instant acceptance when proved false. The current status quo with the LLM doesn't display anything even remotely resembling that. Current LLMs confuse truth with popularity, consensus, corporate speak, etc.
My own bet, like all our stack, is on a memory efficient deterministic AI, where the next token is derived instead of predicted. I don't know whether anybody is working on something like that.
The current AI is definitely useful and it would be false to say it has no utility. I just would like everyone and everything to be more honest and true and focus on the LLMs existing capabilities. On how it can be a useful utility.
This hype train of anthropomorphization and grandiose capabilities, which obviously is false, is what pisses me off.
This metric seems to be asking if years of research by specialists could be emulated by a few LLM api calls. Reminds me of this meme:
https://imgflip.com/i/b366vm
Love that. I myself first thought. How specialized is this and at what level of human complexity is this team working.
The difference to this meme is:
Often people who are critical of whether AI can lead to research breakthroughs are experts in the respective area, who nevertheless are afraid of their future career prospects in academia (getting a permanent position in academia is hard and it is deeply political who gets such a position).
This people are thus not scared by AI per se (it's basically their daily job to devise innovations that advance their field), but their fears are that
- because of the hype around AI the research into which they invested years, often decades, will be considered "unimportant",
- incompotent people in decision-making positions will think researchers can be replaced by AI.
It's still a bit weird though. For obvious reasons, The MO for benchmarks like this has been that the models remain pretty low until suddenly it's done. So even for a 'should i be worried yet' reality check, it's pretty terrible. You can't really keep track of what models are actually able to do. By the time you can replace years of research by specialists with a few api calls then...
> incompetent people in decision-making positions will think researchers can be replaced by AI.
This is a certainty, not a fear =[.
Actually, years of research certainly can be emulated by AIs, in the sense that they have many scanned research papers and can cobble together a solution to many things, say unsolved math problems, from this. So baseline they can seem pretty great. The article is indeed just reinforcing the point that their innovations are less impressive when it's something that really no one has done before. It's not a trivial point for who are used to AI doing so many impressive things.
> say unsolved math problems, from this
No, it did not. Do not trust the AI propaganda, a match researcher USED AI as a tool to solve unsolved math problems, the AI by itself didn't do anything. They are two entirely different things. If the match researcher didn't give the AI the prompts he gave the AI wouldn't have find any solution.
Yes, we could say that AI is an useful tool to use in math research (the same as calculators, and then computer, are) but no, AI did not solve any problem.
That's more or less what happened with math, so I don't think it's an unfair expectation, especially with more AI advances.
I think this sort of small scale research on this problem is inherently pointless and will be a lagging indicator of diffusion, not a leading indicator of capability.
The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result.
If the AI could do this task we would see this happening in places where the economic incentives let them spend millions of dollars on this problem, not on an eval like this.
This specific form of eval where you just ask the agent to solve it with no specific scaffolding besides GPU access (e.g. nothing like AlphaEvolve, ArchPilot, etc that try to work around model shortcomings) is also going to further trail what is possible at small scale. It's good that we at least give them execution environments now, but this feels like the experiments that were worked on figuring out how to get LLMs to do native arithmetic rather than just giving them a calculator/python env.
Regarding the Navier Stokes theorem, one of the best posts I've seen about this (and the other OpenAI math proofs) recently is Nestor Guillen's post where he introduces "the convex hull of ideas". I love this because I've had some vague ideas about the limitations of LLMs that aligned with this, but Guillen really clarified the idea and explained it with a great analogy: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...
Basically Guillen is arguing that LLMs are great at finding results within the "convex hull" of existing literature (i.e. their training data). They can make connections across different parts of the literature where it would be impossible for a human to be an expert in all these areas. But when it comes to truly novel, original ideas, there is no proof yet that LLMs are able to go there. That's honestly a limitation I'm really rooting for because otherwise I think the future of humanity is generally fucked.
Love that link. Thank you. It's helpful to find others thinking this way :)
I work in collective intelligence, and been thinking about what surprise is, and why the ability to sense it is baked into each and every one of us. It's a feeling that moves our attention
I've been riffing on the idea of "mapping" collective surprise by gathering information on which statements give us the feeling. The feeling of something being dropped just beyond the "boundary of our knowing", which I think of as a shell we are each building through our models, like something pushed out into semantic space. I'm saying it funny, but it's really just about collecting data on what surprises people.
And in aggregate, maybe you could have a collective map of this shifting thing (really a map of the education system in the sweeping life-long sense), of the whole human cultural swarm. Surprise is the what we each feel as we cross a threshold that we can't easily return over. We forget today what surprised us yesterday. But maybe it's possible to learn something about the structure we build together (in knowledge), if we gather little radar blips on when we are passing over such moments.
Like if a collective over the course of a week, all has a moment of surprise as they learn a specific thing. If you could see that happening to your group, we could all know that we've all integrated (or at least encountered) something that can be built on.
This framework feels part of talking about it: https://openresearchinstitute.org/onboarding/A_B_U.html
Anyhow, just sharing in case it lines up with anything you think about :)
I have heard that before, it's essentially the LLM cannot be creative packaged differently. But it always falls flat because there is just no way to measure it let alone define it. In the end it always give me a vibe of them wanting that to be the case rather than it being it case.
> But it always falls flat because there is just no way to measure it let alone define it.
I disagree. The definitions may have some gray areas, but at least looking back in history there are some breakthroughs that would fall firmly within the "a result found within the convex hull of ideas that existed at that time" type (e.g. Einstein's special theory of relativity), but also some truly novel, "out of the blue" breakthroughs that people agree are fully outside the convex hull of ideas from that time (e.g. Einstein's general theory of relativity).
Based on that, mathematicians and scientists should be able to provide some falsifiable hypotheses about the theorems and kinds of proof methods that would be truly novel.
That's not a definition, that's a gut feeling.
This is just cope dressed up in math formalism.
Before the convex hull, we had stochastic parrots.
Search is an incredibly powerful method in both human and machine endeavors for generating new ideas and we will see Move 37s in math soon enough, even if we have not already.
> Search is an incredibly powerful method...
As powerful it is, it cannot correct the source..
It almost seems like people don’t understand that most of the training of modern LLMs is now reinforcement learning on problem-solving tasks. They spend most of their training literally solving problems that they don’t already know how to do.
It blows my mind that LLMs are solving Millennium Prize problems and people are still trying to argue that they can't do original thinking.
> "This 100% deadly airborne superbacterium is inspired by natural bacteria, it's not truly original thinking," I say, as the superintelligent AI kills me and everyone else in the world
You should read the Guillen's post then, because that's not what it argues. And very importantly, the only Millennium Prize problem that OpenAI has claimed to have solved is mired in controversy because human researchers were on the cusp of releasing their results (and we're likely training OpenAI in the process). This is exactly the kind of "thinking", along with the other results in the OpenAI math dump, that fall firmly within this "convex hull" of ideas.
And also importantly, this framework provides the basis for some falsifiable hypotheses. That's very different from the AI boosterism that just seems to be saying that since AI capabilities are expanding at a breakneck speed that that speed will continue forever, as if a 13 year old will eventually be a hundred feet tall based on their recent growth rates.
> The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result.
Not to mention a century of theoretical foundations and innovations by humans, written for human understanding.
I've wondered how; "The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result." can possibly not also include many dollars worth of duplicated effort? I find it hard to believe every agent was doing something novel. What this means beyond wasting dollars I'm not sure of, but just flinging around the numbers I don't think should be considered impressive or even required to get the results found.
A slightly less random version of a million monkeys typing, except they're digital?
It quite likely does. It's very much a 'brute force to get to the very edge of capability' kind of approach.
> inherently pointless
there is no shortage of so-called studies about LLMs that inexplicably use comically old and bad models like gpt-4o to reach a strong conclusion about the ineffectiveness of LLMs in general
one cannot help to suspect that this specific kind of academic malpractice does in fact have a point which is to undermine the technology as a whole out of fear, to keep the academic "consensus" in favor of tenured academics
and yet the bitter lesson keeps on getting bitter
> there is no shortage of so-called studies about LLMs that inexplicably use comically old and bad models like gpt-4o to reach a strong conclusion about the ineffectiveness of LLMs in general
> one cannot help to suspect that this specific kind of academic malpractice does in fact have a point which is to undermine the technology as a whole out of fear, to keep the academic "consensus" in favor of tenured academics
Serious studies often take a lot of time to do. Also, models outdate very fast. So, of course nearly every serious study on such a topic will be based on models that are considered "outdated" by LLM aficionados.
So, no need to invent conspiracy theories that such studies are made to keep the academic "consensus" in favor of tenured academics.
In my experience R&D has basically two axes: how innovative it is, and how well we can measure the results.
For the quadrant of non-innovative tasks where we already have a good way to measure performance, Claude can handle this. There is very little ambiguity, and we are basically just looking to maximize some metric under a set of constraints.
Many business processes are not like that. They might be conceptually simple, but it isn’t that easy to say whether a system has done a good job or not. I would say that LLMs can help with this a lot but they have bad judgement because it requires talking to people.
And the other, perhaps more rare issue is in problems where there is data but actually modeling it to sufficient quality or fast enough is hard.
corollary: coding is solved, bugs are not solved
What does that mean to you?
Go ask the man himself
https://x.com/bcherny/status/2090649326032945591
"See this post in the app Use the app to view all comments and discover more posts."
Not to be too contrarian, but is there a solution which doesn't require me to buy a new phone? The link refuses to divulge relevant information on mobile browsers, and the app it's telling me to install isn't backwards compatible even a few years.
Separately, why is that the link to "the man"? Their Twitter profile isn't linked in HN, and I would've naively thought that asking them a question in reply to their HN comment would be a reasonable way to ask them about that HN comment.
This comment is a joke about how Boris Cherny (creator of Claude Code) boasted that “coding is solved” and immediately received the expected mockery/screenshots of bugs in Claude Code. He replies to one of them “Coding is solved, bugs are not yet solved. Fix incoming”
Thank you so much!
> coding is solved
Coding is as "solved" as self driving is...
Imo, self-driving is massively held back by the cost of hardware. There are fully autonomous robo taxis. That human taxis are still commercially viable is in large parts because robo taxis are quite expensive.
Coding cookie cutter CRUD apps is pretty much solved at a much lower price point than artisanal software.
> Coding cookie cutter CRUD apps..
Give me one example of a fully vibe coded software with zero human input that is running in production with non trivial amount of actual users that have been running for more than a year....
Most software is some internal app with a trivial amount of users.
So? Coding is solved for toy programs?
Coding is solved for the kind of problem you would have hired a team of 3-5 juniors a couple of years ago to build something for your small to medium sized business.
Short version of the article: no, not even close.
Practically every paragraph is negative, with sentences like "Agents made misleading claims about their work," and "A natural question is whether the agents could have improved with larger GPU budgets. Although both improved across their runs, in the case of Fable the improvements were almost entirely due to attempted cheating." and "For Sol, the answer is less clear-cut; it did make some progress, although its method was fairly incremental and had limited applicability to the coding task. This suggests that we should be pessimistic about further GPU spending," all reinforcing the fact that LLMs aren't currently capable of this.
My own view is "No, of course not, in fact they will never be capable of achieving good results with that technique." Because that technique will end up training the LLMs on their own output and lead to the inability to distinguish reality from hallucination. If you think I'm wrong about that, I'd be interested in hearing why.
A bit too pessimistic imo. I agree that AI can’t automate things end to end, but a good deal of R&D involves kicking off a training run and babysitting it.
If your training run dies at 1 am and you’re sleeping, you won’t find out about it until the next day. You can lose up to 18 hours of work depending on when it happens. Based on the error it might be as simple as tweaking a single hyperparameter and rebooting, which is something LLMs are usually capable of.
Even just that task means I can kick off multiple runs over the weekend and have confidence they’ll finish. It’s a game changer.
I'd classify that as an entirely different category than AI self-training. What you're describing could have been done with a short script, though which parameter to tweak and how to tweak it would be difficult to automate with a non-LLM script, so the LLM's being able to parse the error message and base the tweak on the content of the error is a definite improvement to the process there.
But I'd classify this as LLM being used to automate a sysadmin task, rather than calling that self-training.
Yeah I’m not trying to argue it is AGI, but it’s not as simple as a short script. There’s some amount of debugging involved, and no amount of if-statements could cover all possible ways a script could break.
In a way, “recursive self improvement” just means tools helping us to create better tools. At least that’s what the words mean.
the GP posits the problem with "recursive self improvement" is the poisoned context & hallucination problem. While a strong loop that has fail safes, backups, restore points (essentially, a fancy backup system), the problem isn't that we cannot create a healthy advanced wiggum loop; it's that every step of the LLM as it grows whatever knowledge is acretes, has a chance of being either poisoned (eg, it conflates two tokens as describe different things) or wholesales fabricates a method or procedure.
Now humans are just as bad, but they're not moving at the speed of compute so the posion and fabricates can dissolve over time, or just, as you've noticed turning on your news, get stuck in very stupid positions. So humans are clearly capable but clearly don't tend to do this either.
So then we dont have a real road map. The error rates, although small, acrete at exponential levels and will wash out improvements.
So I also had the idea that "if we just give it enough context, surely it'll be more powerful". But the error rates hit that squarely. The larger the context grows, the more likely it hasn't properly organized its knowledge to avoid overlapping facts.
In programming, it's worse, because a lot of the code is purposefully "DRY" and reuseable. Everye C program has a main(); is it remembering the correct main? or any of the number of same variables?
You can see an LLM is powerful but it's not ominipotent. It'll suffer very much when it starts hallucinations and context poisoning.
So, sure you can try a super ralph wiggum loop with memory, fallback safeties, etc, but you basically then need another turtle that does the same thing, and at that point, you're positing a infinite jest of ralph wiggum loops tracking each other, recursively, forever.
This was an interesting post on the subject that more or less agrees: 'self-improvement' is happening with things like this but there's a lot of headwind on any 'hard take-off' where capabilities grow exponentially all on their own:
https://www.rameznaam.com/p/471bbae4-1163-4048-944b-18f8b0bf...
We don't have to prove you wrong, you have to make a case for your position. Your statement was vague and handwavy and could be countered with another vague statement such as "They will add a feature that allows the LLM to better detect hallucinations".
I phrased it that way because I couldn't remember the term "model collapse" at the time. But that's what I meant: that making the LLMs self-train will lead to model collapse.
There; now the statement is far less vague and handwavy, because I'm making a specific claim that is, AFAIK, well-understood.
Also, you seem to have misunderstood me a little. I didn't mean "prove me wrong", I meant "If I'm missing something, please tell me about it." More of a conversational request than staking a claim in an argument. Many people at HN seem to like to take argumentative, debate-competition stances — but I usually prefer more "Hey, let's discuss this interesting idea, point out mistakes each other is making, and learn together" kind of interactions. That's what I was asking for.
Fair enough.
Well, I guess the answer isn't too different. I believe model collapse is a limitation of the current AI tech but maybe not the future ones. You can see humanity as a huge model that trains itself. What is novel about AI is that we built a machine with some intelligence traits that is free of biological constraints. If we can emulate the aggregate intelligence of a civilization inside a machine, it could improve itself forever but at a much faster pace.
You couldn’t remember the term model collapse because people quit saying it shortly after it was invented in 2023, because it was just cope by people praying for that outcome. Internet will have LLM output, therefore models will train on that output, then kaput. And what happenend next? Models trained on model output on purpose and just keep getting better.
Hallucinate -> do an experiment -> see it fails, try again. Hallucinate -> do an experiment -> it works, model innovated.
Agentic models brush up against reality, this gives a way around the hallucination problem.
Here is a recent talk showing that hallucination and discovery are actually positively coupled. https://www.youtube.com/live/ZNlZsI9kBm4?si=nhn4ancXu7s6qtom...
Ground breaking thought (ML Researchers Will Hate Me For): (some) hallucinations are actually creativity.
Creativity unbound by reference is not even dada (which was reactionary) it is noise
Because - it's the part people are having hard time to accept - creativity was never the hard part. It's randomness. The hard part is navigating the narrow border between "not enough random, therefore boring" and "too much random, therefore unhinged/insane".
That border is determined not as much by ground truth, as by people's sentiment. As such, many famous artists and creatives never succeed navigating - their contemporaries would call them crazy, and then decades or centuries later, the border shifts, and suddenly they're remembered as greatest creators of their era.
That is not true at all. As I previously stated: reference, stealing well, blending styles in new ways that obey or break the rules of your art form… what appears random to outsiders is a story that follows as naturally as 2 after 1 if you understand the context.
Until an AI can have desires or experience life it cannot tell a story.
Even if it could it is questionable if such alien experience qualifies as story.
It can retell stories. It can swirl around extant media like a child playing with its dinner. It cannot blend in its own lived and new perspectives which is what gives new art value.
> what appears random to outsiders is a story that follows as naturally as 2 after 1 if you understand the context.
It's interesting to compare your opinion with Andrew Wiles quote:
> Perhaps I could best describe my experience of doing mathematics in terms of entering a dark mansion. You go into the first room and it's dark, completely dark. You stumble around, bumping into the furniture. Gradually, you learn where each piece of furniture is. And finally, after six months or so, you find the light switch and turn it on. Suddenly, it's all illuminated and you can see exactly where you were. Then you enter the next dark room...
fable tried to write his own blog posts from his own POV here https://ashallowlake.com/
This is analogous to “I saw Abe Lincoln at the hall of presidents at Disney.” and insisting it was really Abe Lincoln.
It is infantilizing to suggest this is originality or creativity just as much as it is to suggest David copperfield is a scientist.
It is an illusion. A subterfuge being performed by the puppeteer. It works due to the theft of human’s work.
Ask yourself, "What are this creative engine's sensory inputs to experience life?"
"What can this creative engine tell me about the experience of life from its own direct inputs?"
"How does it blend/shape or re-interpret raw input direct sensory data with the, admittedly, astounding body of knowledge it has about creative work?"
"What form of change or growth has this engine experienced that teaches us about ourselves?"
If you take those questions seriously you will come to the same conclusion. It is a Roomba smearing the valuable creative work of millions like dog shit across the living room floor, not a consciousness experiencing the agony and ecstacy of life.
If you think I am wrong ask an AI. It won't come up with an original answer, but it can certainly tell you whether the collective body of slightly-outdated human knowledge thinks I am correct.
Yea, I think LLMs would be much less powerful if not for the inherent randomness in the GPU...
It is like a decompression program that results in a slightly different result each time it is run on the same compressed data. If you are lucky, the delta could end up with answers to all your questions..
In practice, hallucinations have only mitigated within verifiable domains. The vast, vast majority of human activity is unverifiable, or at least significantly less verifiable than mathematics.
What will never be capable? Neural networks? Neural networks that utilize next token prediction in their training recipe?
I recently did a lower tech "write an async divide and conquer S3 lister in Python" and the results were fairly lackluster.
Opus 5.5 and GLM 5.3 Flash came out with pretty reasonable code that was fairly fast. However, all the models left aioboto defaults which have a 10 connection pool size so they ended up all artificially capped even though they added semaphores with higher limits.
My attempt I spent a couple hours AI assisted on ended up being faster, using less CPU time, and having roughly half the code.
Give it a verification loop and enough compute, and AI will soon cook algorithmic R&D like it cooked mathematics. There is just no question whatsoever of this happening, it is guaranteed.
It will cook only as long as brute force through search space is cheap and economically viable.
> ...and economically viable.
And there are so many indicators that it's not currently economically viable and will only be if retail price is massively increased and strong regulatory barriers to competitors are erected. ;)
Very interesting if it's not currently economically viable. Do have any sources on OpenAI's margins?
Have you seen my drafts? For easy task I may write the straighforward solution, but for complicated stuff it's a mix of throwing stuff to the wall and see what sticks. It's informed search by experience, but AI also use weights to pick the attempts.
Do you write 10,000 drafts before deciding to send an important mail?
Neither solved the Navier-Stokes equation.
I certainly iterate through many drafts in my head before choosing a route.
AI is not a calculator. OpenAI did not solve hundreds of open problems in math by "brute forcing the search space," there was a lot of intelligent and novel thought (gasp!) in the way AI chose the paths it did.
Yes, AI can explore thousands of ideas at once, no, that's not "brute forcing the search space" because the search space is way larger than you think it is.
It is a spectrum.
Godel, Euler, Ramanujam, etc. didn't burn millions of dollars or drudged through search space like OpenAI's agent do to prove what they need.
Let's not forget using Lean as a crutch throughout the way rather than as a final tool for formalization and verification.
Brute force would choke on a measly decillion of alternatives. It corresponds to an extremely short proof. A combinatorial explosion doesn't allow to use brute force for anything slightly longer than toy problems. Swarm of 10000 agents is a massive smart force.
> Swarm of 10000 agents is a massive smart force
Yep, that's how Euler, Gôdel, etc. proved their theorems, both easy and difficult ones.
They trawled through the search space just like a swarm of 10000 agents.
> it cooked mathematics
Weird, these are all still here? https://en.wikipedia.org/wiki/List_of_unsolved_problems_in_m...
Several of the items on that list are claimed to have been resolved by the recent OpenAI theorem-dump. Another is the Navier-Stokes question that's one of the Millennium Prize problems, also claimed to have been resolved by OpenAI's models.
So: unless all those claims by OpenAI turn out to have been mistakes[1]: no, actually, those are not all still there.
[1] It's certainly possible that some will. They've already retracted a few things.
There is already a pre-print making the claim that the Navier-Stokes solution provided by OpenAI does not meet the criteria required by the prize (https://arxiv.org/html/2609.20803v1)
(and as noted on Wikipedia, the Clay institute still lists the problem as 'active' not solved https://www.claymath.org/millennium/navier-stokes-equation/)
I'm not a mathematician, but looking at the Clay Institute's official problem statement [1] and that paper, I don't think it's claiming OpenAI's construction doesn't meet the Millennium Prize criteria.
The criteria for statement C (which OpenAI targeted) only require a C-infinity smooth force, whereas the paper considers constructions of the same kind as the OpenAI one but with a real analytic force. All real analytic functions are C-infinity smooth, but not vice versa [2]. Wikipedia says the Navier-Stokes problem for real analytic forces is still unsolved, so the question of whether a similar approach can be used there is presumably of some interest, but it doesn't appear to affect whether the Prize criteria have been satisfied and the paper doesn't explicitly make any claim that it does.
Re the problem still showing as "active": The Clay Institute don't consider problems solved until the solution has been fully digested by the mathematical community and is well established. In the case of the Poincare Conjecture that wasn't until 2010, 7 years after Perelman put his papers on the arXiv. Nobody is expecting them to mark Navier-Stokes solved any time soon.
Your comment was the first I've heard of that paper, BTW, so thanks for that. However, that's another indication that this probably doesn't bear on the Millennium Prize criteria. The lead author, Peter Constantin, is a leading expert on the problem and the Clay Institute's page on it has a video of a lecture by him. If he had come out and said OpenAI hadn't solved the problem as stated I'm sure there would have been a lot of noise about it. The paper is dated 17th of September so presumably there has been plenty of time.
[1] https://www.claymath.org/wp-content/uploads/2022/06/navierst...
[2] https://en.wikipedia.org/wiki/Non-analytic_smooth_function
Where in your link do they claim that?
It is pretty straightforward for anyone who knows lean to verify that the OpenAI release satisfies the official criteria to the letter. You could reasonably argue that it shouldn't satisfy the official criteria because the official criteria are unreasonably generous in allowing any smooth external forcing function, but that would be a different claim.
Theoretical fields like mathematics inherently suffer from a form of Goodhart's Law. The progress in mathematics you hear of is almost never the kind of progress non-mathematicians would care about.
It's always easier to measure progress according to the internal metrics of the field than to evaluate the contributions to the wider understanding of the topic. In theoretical computer science (which I'm most familiar with), people have long complained about results focusing on shaving sublogarithmic factors from complexity bounds (while making the algorithm worse in practice) and about reviewers being impressed by the technical difficulty of proofs. But progress like that is easier to measure than new algorithmic ideas or conceptual understanding.
But it's not all bad. The researchers chasing the metrics are almost always genuinely interested in the topics they study. Their actual contributions mostly come from the ideas they explore while trying to achieve measurable progress. And because they are not expected to produce anything of direct value (Goodhart's Law for applied researchers), they can explore a wider range of ideas.
I personally noticed this when I moved from theoretical computer science to algorithmic bioinformatics. When I start a new project, the expectation is that researchers in genomics should be using sofware that uses the new algorithms five year from now. That expectation is useful, but it's also a strict constraint on what I can afford to try.
Mathematics isn't cooked though.
How much of this can change if subsequent training runs produce models that are much better at abduction?
The only outcome I see is that AI will become so complex that we won’t be able to rely on it for R&D, because we won’t be able to measure and prove its results. Some AI collaboration and speed of work will be impossible for humans to replicate, making it neither good nor bad—just something beyond human ability.
Could it be that the companies have nerfed the models on these domains? It is a very hard thing to do because it can hurt related domains. But its not beyond the ideology of Dario - he tried it publicly .
I always assumed the public facing frontier modes were abliterated with regards to AI development?
I think the more interesting thing is was it unable to do it even knowing the solution? I feel like only testing a single innovation is biasing the current state of automated AI R&D.
I don't buy this whole narrative that "AI is innovative".
If AI is so innovative then why am I seeing a dumb idiot every time I discuss anything even remotely innovative with it.
Most often it cannot even produce a true sentence when the claim in the sentence is modestly strong without injecting qualifiers and distancing itself from the claim.
Having to just accept that it produces relentless innovation is so detached from reality that it is not even funny.
It is true that it can write decent code once the domain bounds are well defined. What I have found is that it is good at scrutinizing already written code - but only with expert supervision. Even then it most often tests our patience with stupid suggestions.
> If AI is so innovative then why am I seeing a dumb retard everytime I discuss anything even remotely innovative with it.
My experience slightly differs: yes, LLMs are not the best entities to discuss innovative scientific ideas that one has in mind, but they are still often much more competent discussion partners on such topics than the people you are typically surrounded with (e.g. work colleagues).
(no, this is not a template for a Reddit /r/iamverysmart/ post :-) , but just my experience).
I agree. Your experience mirrors mine. Yes, they are often competent discussion partners, although a lot of people I hang around with are also excellent discussion partners.
But unlike with real people, LLMs have this unique property that we can abuse them without consequences directly leading to our improved understanding of the subject material using nothing but pure frustration as the motivating drive.
So for me, they are productive precisely because they remain dumb discussion partners.
>If AI is so innovative then why am I seeing a dumb idiot every time I discuss anything even remotely innovative with it.
Also, if AI is so innovative why frontier labs still struggle to come up with a money printing product? It's everything so confusing: They have AGI, but their product is an app/API, which is so world-ending dangerous, but their latest product is a set of cute agents running on your own machine, but also is so smart it can do the work of a PhD on almost every field of human knowledge, but most ot they have to show is in something so machine specific like math proofs.
Yes, the arguments we are handed don't line up.
AI is innovative but cannot string together true confident sentences.
AI is so innovative that nothing innovative or groundbreaking has so far appeared. Maybe it will, but the claims to ground-truth ratio is off the charts.
World ending is so funny it is usually only parroted and swallowed by gullible people.
> it can do the work of a PhD on almost every field of human knowledge
This is the most revealing and the most hilarious one. This is essentially the argument from authority wrapped in a PR stunt. Those who produce real innovations overwhelmingly does it outside curriculum and as part of the real work they are engaged in with very few historical exceptions. That statement reveals a lack of understanding of what innovation itself means let alone the virtue signalling they are trying to sell.
> machine specific math proofs
This one was initially intriguing, but as I earlier said, is a well defined and constrained bounded domain. The accusation that hours earlier other human researchers's work may have been used to plagiarize the proof has muddied the waters further.
>This one was initially intriguing, but as I earlier said, is a well defined and constrained bounded domain. The accusation that hours earlier other human researchers's work may have been used to plagiarize the proof has muddied the waters further.
It's even worse. They claim, and the media repeat, that new math results are product of the models intelligence. But we all know there are elite mathematicians working for these companies (OAI, Anthropic) and basically are they who are coming up with new results, assisted by AI, not the other way. And we still don't know how much of the heavy lifting is done by each part. If the models are PhD smart, shouldn't be able to come up with results with minimal guidance? If they're so innovative why the need of an army of top mathematicians, engineers, physicists, etc. and thousands agents and millions in tokens to do novel work? It's just seems to me that humans are doing (barely, hence the plagiarism) the creative, innovative work and AI agents are just useful to explore a massive space of bullshit hypotheses to find something that finally sticks, after being validated again by an army of human experts. AGI and the near super intelligence? It doesn't track.
It is entirely plausible that, as you say, they could be piggybacking off of human work internally and using AI like the punching bag it is - dumb but useful. I wouldn't bet against that, but usually it is difficult for people who did the hard intellectual work to have their credit taken away - which might be exactly what we saw with Navier-Stokes drama that unfolded.
> If the models are PhD smart, shouldn't be able to come up with results with minimal guidance?
Smart is born out of wrestling with the problem domain, not in virtue signaling via academic credentials. So anyone leading with an academic credential is a dead giveaway of the PR stunt underneath. Nobody says PhD smart Einstein - because work can lead the speaking and doesn't need a credential as a support crutch.
> AGI and the near super intelligence? It doesn't track.
True ideal intelligence must not hedge, shouldn't pander to popular opinion or consensus. True intelligence must state what is true as maximally true as is correct while being able to derive more true sentences coupled with instant acceptance when proved false. The current status quo with the LLM doesn't display anything even remotely resembling that. Current LLMs confuse truth with popularity, consensus, corporate speak, etc.
My own bet, like all our stack, is on a memory efficient deterministic AI, where the next token is derived instead of predicted. I don't know whether anybody is working on something like that.
The current AI is definitely useful and it would be false to say it has no utility. I just would like everyone and everything to be more honest and true and focus on the LLMs existing capabilities. On how it can be a useful utility.
This hype train of anthropomorphization and grandiose capabilities, which obviously is false, is what pisses me off.
If AIs start getting truly good at innovation, we may cease to recognise our reality.
I saw one mathematician describe the recent OpenAI math-dump as 'alien-like' math.
Imagine a world of countless new aircraft designs, fuel sources, musical genres, architectural styles. It could be bewildering.