postalcoder 1 day ago

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

  • enraged_camel 1 day ago

    Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.

    • nullbio 1 day ago

      Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle.

      • thereitgoes456 1 day ago

        The Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less.

        While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others.

        Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.

        • selectodude 1 day ago

          The cursor acquisition just shows that there’s always a dumber shithead out there. Though vaporizing Elon Musks money is about as pure of a good as there is out there these days.

          • throwaway240403 1 day ago

            Improvements in models and products coming out of SpaceXAI since the acquisition would seem to disagree with you.

      • throwup238 1 day ago

        > Andreessen Horowitz is being played like a fiddle.

        Andreessen Horowitz is not being played like a fiddle here. This might be their only investment in a decade that isn’t entirely predicated on being a scam.

  • mediaman 1 day ago

    This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

    Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

    • felixgallo 1 day ago

      Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

      • letmevoteplease 1 day ago

        > Altman was caught in previous attempts trying to game benchmarks

        Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.

        > does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

        I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)

        And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.

        • vlovich123 1 day ago

          Yeah and Astra is much better still

        • kzrdude 1 day ago

          There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback.

          That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.

          Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.

          • xnickb 23 hours ago

            Sounds more like profitmaxxxing to me to be honest.

            • kzrdude 21 hours ago

              They have plausible deniability on that one: not making any profit.

    • general_reveal 1 day ago

      I take it to mean the benchmarks are a marketing line item, as in, to sell this fucking thing you have to go out there and lie and the way everyone is lying is by doing exactly that, lying. They build for benchmarks and build benchmarks for builds.

      You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world.

      Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase.

    • iLoveOncall 1 day ago

      > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

      Yes? Just like every single model from every single AI lab.

    • postalcoder 1 day ago

      > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

      A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?

      Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing.

      • Lucasoato 1 day ago

        > A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?

        Wait a second, are we taking into account the massive difference in terms of resources of these two companies?

        • solenoid0937 1 day ago

          Irrelevant when they say they're competitive with Fable and Astra. They don't get to then roll that back and then say "but we have less compute!"

          You're either competitive or not.

        • asdfsa32 21 hours ago

          You should use my model then, I spent about 30$ in electricity and used my existing RTX4090. It is not very good, but can you compare it with others really? You can use this service via a private API with a VPN, email me your credit card details for access.

    • nrmitchi 1 day ago

      > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

      Yes.

    • dpweb 1 day ago

      Fixed benchmarks will be debunked eventually I'd think. Better to use synthetic problems.

      Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.

      • gunalx 1 day ago

        Wouldn't work. A dynamic but verifiable problem. Is just a perfect target for a RL environment. If you don't have the verifiable part the benchmark is useless, or really expensive with human review. (Or just open ended)

        • dpweb 11 hours ago

          There doesn't need to be a single correct response. Generate "novel" problems with a set of acceptable solutions/outcomes, then verify that the response satisfies the criteria.

    • willcmcc 1 day ago

      "Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"

      Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound.

    • ben_w 1 day ago

      Benchmaxxing is the default case, and always has been.

      It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial.

      Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: reproductive success.

      • conmod278 1 day ago

        Why would I adapt myself to unseen environments unnecessarily?

        • Muromec 1 day ago

          Because then you could sit under the palm tree and enjoy your free bananas and relax.

        • NewJazz 1 day ago

          You can see possible existing environments even if you explicitly blind yourself to them, through reflections off of environments that you do not blind yourself to. Even with the benchmark excluded, the social zeitgeist that has considered the benchmark and included it or ideas from the benchmark either implicitly or explicitly in their code, documentation, et cetera is still part of your training data.

      • podocarp 23 hours ago

        You can't justify it being OK just because it's common. Here, this just makes benchmarks into a low signal and useless marketing number once people get numb to all the 99%s.

        Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.

        • injidup 22 hours ago

          Cancer is simply cells breaking free of the cooperative jail. Essentially the grey goo scenario of nano machines. Instead of cooperation they just do their own thing.

          Cells have a certain optimised DNA mutation rate kept in check by various machinery. Multi cellular life expects each of these little replication machine to co-operate in the grand scheme of running a body. But it's also required in the grand scheme for DNA to mutate a little bit to ensure population variance. So you could say that cancer is the tax paid for having cooperative yet flexible and adaptive nano machinery.

          So yes the propensity for cancer developed under evolutioniary pressure towards a non zero level.

          The population could have optimised for zero cancer but it would not have paid for itself in terms of overall population adaptability and survival.

        • ben_w 13 hours ago

          > Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts

          The host's survival is not the benchmark of the reproductive unit, which are the cancer's cells short-term reproduction.

          The distinction is the reason benchmaxxing is in fact not good: the benchmark is only approximately related to what people actually care about. You do want your cells to reproduce sucessfully, after all; you just also want some emergency stop buttons for when they go wrong, and those things failing is your body's benchmark rather than your cell's benchmark.

          (There's at least two examples of cancers that can be spread from host to host; lupine genital and taxmanian devil nasal, IIRC)

      • muddi900 14 hours ago

        Cancers being distinct living beings by itself is controversial, but you unnecessarily anthropomorphizing.

        The truth is much more mundane. It's just Goodhart's Law.

        • ben_w 13 hours ago

          If people already knew about Goodhart's Law, I wouldn't feel it necessay to give illustrated examples.

    • yieldcrv 1 day ago

      I mean, still benchmaxxed, I happen to not consider that a problem

      These firms are literally hiring professionals from all fields to teach procedure

      To teach processes that can subsequently be done agentically or in automated chains

      Its basically infinite permutations of tool calling, except the tools aren't external, they’re baked in upon birth

      So yeah still makes sense that the new benchmark has a low score and the older one has a high score. And sure, one day we wont have to debate it and a new model will ace everything. Do you actually want that day to be today?

    • ActionHank 15 hours ago

      TB4 has not been saturated yet.

      They are all gaming these benchmarks, it is perfectly reasonable not to trust any of them.

    • muddi900 14 hours ago

      > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

      Probably

  • eranation 1 day ago

    When a benchmark becomes a target, it's no longer a good benchmark...

    • tonychang430 1 day ago

      people are just fighting for numbers.. i don't fundamentally see the model being better

  • fallingbananna 1 day ago

    Those 27.3% are still in the ballpark of modern models:

    - Sonnet 5 - 12.4%

    - Luna - 17.3%

    - Grok 4.6 - 20.3%

    - Sol - 37.3%

    - GLM 5.3 - 41.8%

    - Opus 5 - 51.8%

    • p1esk 1 day ago

      Astra is 58%. The current title says it's "rivaling Astra"

      • thereitgoes456 1 day ago

        It is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.

    • nijave 1 day ago

      This explains a lot about Sonnet 5.

    • mokre 1 day ago

      GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...

      Also a lot of questions to benchmark because opus 5 is completely useless model right now.

      I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

      • gpt5 1 day ago

        Terminal Bench 4.0 did not introduce new questions. All tasks were public for a while. If you look at GLM 5.2, which is using the same base model as 5.3, but was released prior to most tasks, it does extremely horribly on terminal Bench 3.0 (4 to 8 times worse than every other model) - source: https://benchlm.ai/benchmarks/terminal-bench-3

      • didibus 23 hours ago

        Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.

        • simondotau 16 hours ago

          Opus 5 keeps making howlingly stupid errors, like one recently where a regexp would catch an invalid date because -\d{2} won’t match -00

          Seriously.

    • airstrike 1 day ago

      So, better than Sonnet and Luna? lol

  • throwatdem12311 1 day ago

    This is why I find benchmarks absolutely worthless.

    First, almost all models are within spitting distances of eachother.

    Second, it never translates to being better for my own workloads.

    You just need to make your own benchmarks.

  • walrus01 1 day ago

    For comparison Qwen 3.8-Flash-Next which runs in under 190GB of RAM locally scores 25.3% on terminalbench 4.0.

  • Readerium 1 day ago

    Yeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0

  • thefourthchime 1 day ago

    Came here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4.

    I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!

    Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.

    • dudeinhawaii 1 day ago

      Your post made me wonder if Artificial Analysis had finally moved to TB4 and lo and behold they have and Astra is tied with Fable 5.1 at 53.

      That then made me realize that they lower the bars of tied scores so on the site it looks like Astra in second place. Weird. Anyway, yes, so many of these composite benchmark sites are irrelevant if they're not trimming the fat and sticking to the most up-to-date variants.

  • muddi900 14 hours ago

    > benchmaxxed

    An aside: When did talking like incels became cool?

gruez 1 day ago

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked?

https://www.youtube.com/watch?v=tNmgmwEtoWE

As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.

  • notfromhere 1 day ago

    Well the models did get better but yeah their early product was godawful

  • fishtoaster 1 day ago

    A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

    • htrp 1 day ago

      a reminder that money does enable you to make mistakes and buys you the ability to correct from them

      • unshavedyak 1 day ago

        > and buys you the ability to correct from them

        Or at the very least, make more mistakes.

    • darkwizard42 1 day ago

      I think the evolution of the harness and ability to preserve loop context outside the context window has made running these kinds of agentic experiences easier.

      sorry so many buzzwords to say, the capabilities to do this kind of work are more accessible and easier to manage, so now it works!

      Good to see, and agree they were severely overhyping their product back then.

    • tyre 1 day ago

      When they launched Devin it was supposedly at the performance of an engineering intern. Friends who used it found the bad parts of an intern (tons of handholding, review required) but it didn’t learn from mistakes or add throughput.

      They seem to love a good overpromise.

    • wetpaste 1 day ago

      Not all that surprising. The original Devin really was just an early attempt at agentic coding before models were really even trained for it. Now that it's a well established pattern and we've figured out what works, I'm not surprised they've morphed into something reasonble.

    • arjie 1 day ago

      Their ads in SF are pretty funny. “Remember Devin? It’s good now”. Okay, very self-aware, Cognition.

    • klardotsh 1 day ago

      Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.

      • paimapi 1 day ago

        Had the same experience on it when it was going by Windsurf. Had to wire up a skill hooked to terminal runs or else it would hang or freeze and never finish literally every single time

        I see issues with other harnesses too but not with the regularity I was getting from this. And the one moat they had with the better UI for per-project multi-agent orch disappeared and now is standardized

      • fishtoaster 23 hours ago

        Can't say I've tried the CLI. I've mostly focused on the cloud agents, which I was explicitly looking for. I compared against cursor's cloud agents, ampcode, and hoplite, and came out surprisingly enjoying devin.

        I will say that the lack of parity between Devin cloud and Devin desktop is downright embarrassing. It's very clear that the latter is a thinly-reskinned Windsurf. A visually similar UI with vastly different capabilities. Definitely a black mark on the whole thing.

    • Saline9515 1 day ago

      Devin lacks many features and doesn't even have an exec mode. Other than they subsidize the subscription, the added value here is low.

      • princevegeta89 1 day ago

        I used Windsurf for quite long and they're basically dead for now after they became Devin.

        Everything feels dull and they're always several features behind while Cursor is just killing it every other week.

        I'm back to VsCode plus Copilot Pro.

        • airstrike 1 day ago

          Same, but I'm back to just Claude Code and the occasional vscode, though for my latest project I switched to Astra instead

  • esafak 1 day ago

    Previous versions were based on Kimi too. I'd consider if it I could access the model outside Devin. No lock in for me, thank you very much.

  • sterlind 1 day ago

    I'm surprised they can post-train a closed model off of K3. Is that the norm for open-weight licenses?

    • justincormack 1 day ago

      Yes. If they are really open, you can use them as you wish. Some licenses are less open though.

      • Ohentis 1 day ago

        I would be somewhat surprised if any license restriction actually holds up in court.

  • deet 1 day ago

    As others have mentioned, it's really matured a lot and at this point is one of the best cloud-hosted, team-managed coding agents, when factoring overall UX, testing and QA lifecycle via its sandboxes, and its ability to be controlled with an API. We use it quite heavily.

    It's coming from a different starting place than Claude Code or Codex are as individually controlled single-developer tools. Devin has been more persistent in pursuing the direction of something that operates more autonomously at the team level, as a peer. And while it might be slightly behind in raw harness ability (maybe?) it's probably ahead on the team-focus.

    • thereitgoes456 1 day ago

      Your startup is based around AI coworkers, so I expect you are biased towards overvaluing their usefulness.

      • deet 1 day ago

        Yes, perhaps fair. But my point was somewhat narrow. I wasn't saying that team-managed agents are good to go for all cases and that they're better than individual dev-managed ones. Just that they have gotten better and that of those Devin has some of the better UX.

        Our experience might also not be typical because we have built infrastructure around making Devin and similar agents work better. And for the record no ties to Devin/Cognition. Just pay them too much as a customer.

nullbio 1 day ago

Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1?

I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.

  • eru 1 day ago

    > If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

    Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

    • bayesianbot 1 day ago

      1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead.

      btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I really didn't expect it to be anywhere near this good so we'll see where it ends up. And it's really fun throwing crazy amount of tokens at the wall for ~free instead of watching the subscription limits tick closer while your agents churn away.

      • cmrdporcupine 1 day ago

        Yep. 4.1 Flash is good enough for most routine coding things, but it also makes up for a lot of weakness by being so fast (and cheap of course).

        I'm willing to tolerate babysitting things a lot more if I know I'll get almost instant results.

        • eru 1 day ago

          You could also have eg Sol do the babysitting.

          • cmrdporcupine 10 hours ago

            tbh i'm not really finding I need to. It's seriously quite impressive.

      • user43928 1 day ago

        That's 1/3 of GPT 5.6 Luna. It seems rather close to me.

        But great that we have a new leader in performance/price in that segment.

      • eru 1 day ago

        I did a lot of that kind of work with the older version of DeepSeek before they upped the prices.

        For example, it was quite good to get a decent Sashiko review. Sashiko is a Linux kernel review agent with interchangeable LLM driver. It's very good, but it eats tokens like crazy.

      • eru 1 day ago

        > 1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead.

        Losing most of your customers tends to sharpen the mind a bit. They could eg stop pushing out the absolute frontier for a while and focus on making what they have run cheaper. Or they go and do more lobbying against China. Or a million other little things that take more than 30 seconds to come up with when writing a HN comment, but less than a week for someone who's smart and paid to do this for a living.

    • colingauvin 1 day ago

      OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them.

      On the one hand you, if you bought a lot of compute a couple years ago (perceived demand, perceived shortage) you are in a good spot temporarily. But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be. I can almost, almost run DS4.1 Flash at home. 4 sparks can do it at 200+ tokens per second. I have two Sparks, so I am not in the club. Neither is your average laptop owner or gamer either. But your average HN software engineer can probably easily swing 2 sparks.

      • user43928 1 day ago

        I could buy four of them for ~20k.

        That's like four years of ChatGPT + Claude subscription.

        Eight years if only ChatGPT, or sixteen years of the Pro 5x subscription.

        • colingauvin 1 day ago

          It's not particularly good value if you are just comparing $ with no other context. Point is just that it's accessible and so now Anthropic and OpenAI need to make both a performance proposition and a value proposition.

      • eru 1 day ago

        > But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be.

        I'm not sure? If we have techniques to use the hardware even better, that will make the hardware even more valuable, won't it?

        • colingauvin 23 hours ago

          But they aren't really competitive for cost in that size class.

  • notfromhere 1 day ago

    This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI.

    Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own.

    Same reason Harvey is doing models now and basically every other provider

    • didibus 23 hours ago

      Couldn't they just grab and run an open weight model to save on API tokens?

      • onel 21 hours ago

        You get better performance if you also finetune it for your task

  • pizza234 1 day ago

    > we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

    DS 4 Flash requires large amounts of memory to run at reasonable quants (I think a system with 160 GB or so). DS 4.1 Flash is even larger, I think around 250 GB.

    Any DS version is dumb when compared (in realworld tasks) to Astra/Opus 5, which means, one would spend thousands of dollars, and still need to rely on cloud services to do jobs that are non trivial.

pkilgore 1 day ago

Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

  • AznHisoka 1 day ago

    I have not met a single person/company that uses Devin… does anyone here actually use it?

    • PolCPP 1 day ago

      I pay for it (mostly because they grandfathered me from the old prices)!

      I used to use windsurf as my main editor until they changed their pricing model. Now i use it just to burn my weekly tokens on fable/astra if i remember to that on a task and that's it.

    • chris_st 1 day ago

      I use it, and have been happy with it for the most part. Like sibling, I use it for GPT-5.6-Sol and Opus work, and use their free models (GLM 5.2 for the past few months, trying SWE-2 now).

    • dominotw 1 day ago

      my friends at infosys are being trained on it

      • 0l 1 day ago

        Deeply unserious company so not surprising

    • fschuett 1 day ago

      I only used their "DeepWiki" automatic docs, they are pretty decent at getting an overview of a large project and are relatively accurate, with diagrams and anything. Haven't tried out their coding agent stuff.

    • hightrix 1 day ago

      My company uses it. We have a bunch of seats in an enterprise plan and have been using it for 9 months or so.

      It's a great product compared to Copilot. It is also the first AI tool I used heavily outside of creating random images or one off questions.

      I'm now using all three, Devin, Claude, Codex. I'm finding Claude and Codex to be much better. One of my biggest gripes is that the web client and desktop client for Devin are two completely different harnesses, so the quality of responses varies greatly.

    • klardotsh 1 day ago

      Unfortunately yes. It's awful, other than that they at least support both open-weight and proprietary models to route to. So I guess their model router is "fine", but the harness? As I said in a sibling thread on this page:

      > Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.

  • SaltyBackendGuy 1 day ago

    > And yes, I tried again

    This made me laugh a bit. I was forced to do an evaluation of their shit product twice due to being backed by the same PE firm; "take a look at it again, it's much better now". It sucked the second time also...

  • dvfjsdhgfv 20 hours ago

    Well, where VC money is involved, the aim of these ads placed in SF is not exactly to convince any potential users.

TheJCDenton 1 day ago

> SWE-2 is post-trained from Kimi K3

On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

  • htrp 1 day ago

    why are all American AI models basically Kimi in a trench coat

    • ianm218 1 day ago

      No one in the US is going to fund pretty good open source with VC money.

      US has OpenAI/ Anthropic/ Google/ meta/ SpaceX atleast trying to make frontier foundation models 2 are using VC money + cash flow, last 3 are mainly cash flow + equity and debt.

      China has state banks and similar willing to fund lower margin open source labs.

      • Roark66 13 hours ago

        I'd be pretty surprised if someone told me few years ago communist China, state banks would become the main founders of open source compute and our last hope against monopolists like Musk, the whole bunch at OpenAI and so on.

        To be fair Zuck is releasing fairly capable open source models, but Chinese labs are way ahead.

    • nijave 1 day ago

      It's (one of) the best freely available?

  • mxmilkiib 17 hours ago

    I thought it would be GLM based, as Devin has had free GLM-5.2 for a while now

mydreamof 1 day ago

Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?

  • harmonic18374 1 day ago

    Probably, FrontierCode is made by Cognition itself. The model also seems worse in every way than DeepSeek v4.1 Flash, launched today.

    Also the submitter's account is very new which makes me suspicious of self-promotion.

bobtheborg 1 day ago

SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it.

Looking forward to 2 -- maybe it'll be usable

captainregex 1 day ago

I am skeptical. Lived experience is what matters and I don’t have anyone in my life (Devin shop) saying good things about SWE other than it’s free. Hope I’m wrong and it’s not so bad this time

CyLith 1 day ago

I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actually need a jack-of-all-trades model to back my coding agents.

  • skrhee 1 day ago

    I'm also in simulation software! Wondering which models you are finding helpful, the models I'm using for general SWE skills are horrible at our simulations and even basic physics/engineering calculation and intuition

    • CyLith 1 day ago

      I use Claude Opus 4.8 almost exclusively. I have had fairly good experiences with it. One time it derived an entirely novel simulation method different than anything in literature by combining its knowledge about how problems in other fields with similar underlying mathematical structure are solved. That was a bit of a Jacobian Conjecture moment for me.

    • mohamedkoubaa 1 day ago

      I'm in the same field and I find GPT to be better at understanding physics conceptually but Claude is better at writing numerical code. I use cursor so many of my sessions start in GPT and switch to Claude.

Take8435 1 day ago

Post made by account 2 days ago.

  • wy35 1 day ago

    Is there an implication here I'm missing?

palguna26 12 hours ago

At this point, i firmly believe all companies are building benchmaxed models, which perform well on older benchmarks but struggle on new one, terminal bench is the best example.

bluelightning2k 1 day ago

I like Cognition as a company and hope they succeed. Seemingly excellent engineering org.

I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.

andai 1 day ago

Their benchmark used to show other metrics, like output tokens and time, but now only shows cost:

https://cognition.com/frontiercode

Which is too bad, since all of the gains here appear to be from massively reduced output tokens?

The model SWE-2 is based on, Kimi K3, is cheaper per token than Sol, but costs more per task (ArtificialAnalysis) due to using way more tokens.

Whereas, based on the graphs, SWE-2 appears even more token-efficient than Sol! That might have been worth showing off, if true.

eyeris 1 day ago

Wonder if this was the model that drove factoring the rsa-260

The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve.

thimble_io 18 hours ago

92.8% on TB2.1 dropping to 27.3% on TB4 is the only number that matters. The rest is marketing.

pelorat 1 day ago

Unless it can do CAD via computer-use how can you say it rivals GPT-Astra?

sbseitz 1 day ago

Why doesn't clickbait trash like this get moderated ?

  • handoflixue 1 day ago

    Because it's not clickbait trash / doesn't violate the site rules?

    Dang is pretty good at enforcing stuff. There's a flag button and you can email reports if you're really bothered.

alansaber 1 day ago

Fair enough that they did a "propoganda and censorship" eval but not sure why i'd care about that in my highly juiced SWE kimi FT.

scronkfinkle 1 day ago

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

  • samyok 1 day ago

    SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI (https://docs.devin.ai/cli)

    :)

    Disclaimer: I work at Cognition, although was not involved in SWE-2

    • scronkfinkle 1 day ago

      But I don't want to use your CLI. I already have my own harnesses and workflows. The friction is too high to "just try out" a new model like this. It would be preferable if I can evaluate it over, say, open router like all the other models and then decide from there if it's worth downloading a bespoke tool chain for only 1 lab's models

      • jkelleyrtp 1 day ago

        It’s preferable to keep inference capacity available for users using main Devin products than openrouter atm. Might change in the future. Even OAI is cutting off new plan signups to keep up with demand.

    • randomblock1 1 day ago

      I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage. It does say 75% off though. Seems like for Pro subscribers SWE-1.7 is free, maybe SWE-2 is free for them?

      • samyok 1 day ago

        Did you use it via the CLI or Desktop? It's 75% off in cloud and free to run on your device.

        • randomblock1 18 hours ago

          Oh I tried in cloud. I'll give it another shot

    • wren6991 1 day ago

      Your own CLI? Not even a /v1/chat/completions API? Is your business model based on pretending LLMs are not an interchangeable commodity already?

      • anthonypasq 1 day ago

        they are an agent company not a model provider, is this that difficult to comprehend?

    • vopi 1 day ago

      Heads up: it doesn't appear to be available on the Devin CLI (for me as a free user).

    • breznev 1 day ago

      Hey man, how’s the Poke SOC2 audit going? Must be any day now that it’ll be finished, right?

  • CamperBob2 1 day ago

    If its weights are open, that covers a multitude of other sins. Sufficiently-strong performance on the part of the new model would justify adapting existing tools to work with it.

monkeydust 1 day ago

As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft. There is an English translation somewhere.

  • arrowleaf 1 day ago

    Pareto leaked out of the sociology/econ bubble a long time ago :) Pareto principle, Pareto efficiency, Pareto distribution have been in the pop-sci buzzwords for quite awhile, I probably encountered it first in the 4-Hour Workweek. I don't think you can read a self-help book without the author introducing it as a groundbreaking principle to live your life by.

  • ccapitalK 1 day ago

    Pareto frontiers are pretty commonly invoked to describe tradeoffs in computer science and have been for quite a while. I remember the term being used in one of my early algorithms courses to describe the tradeoff between data structures with fast writes, ones with fast reads and ones that tried to balance the two.

    IIRC cognition boasted about hiring a lot of competitive programmers and algorithms experts back when they released Devin, so it tracks that they'd use the term.

    • walrus01 1 day ago

      It's interesting watching people throw about pareto frontiers sort of like how RF nerds approach the shannon limit (in a practical real world sense of the term, like charting possible modulations/data rates on a two way satellite modem's manufacturer datasheet).

Tsarp 1 day ago

"SWE-2 is post-trained from Kimi K3"

  • airstrafer 1 day ago

    Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.

    Maybe still worth it if their "64% cheaper" figure holds.

    • Tsarp 1 day ago

      With the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2.

      • samyok 1 day ago

        SWE-2 is free for all subscribers on the CLI to try out for the next month :)

        • teddyX 1 day ago

          What about the gui/windsurf app? Same as cli?

    • Bolwin 1 day ago

      I don't think you know what distill means

      • airstrafer 1 day ago

        I guess I don't. Does post-training from another (larger) model not fall under the umbrella of distillation? I'd imagine it leads to the same spiky-ness issues...?

        • FergusArgyll 1 day ago

          Distilling you don't have the actual model weights of the teacher. All you have are the teachers answers to a lot of questions. You then teach your own smaller model to answer more similarly to the big teacher model.

          Fine tuning you have the actual model weights of the original model, you then train that model to answer in a different (or better) way.

          • Bolwin 1 day ago

            Distillation requires you to have the actual logits of each token from the teacher model, which in practice means having the model itself.

            What you're describing is just synthetic data.

            Note Anthropic misused the term in their post about Chinese model distillation, deliberately I assume.

  • xlbuttplug2 1 day ago

    I presume post training is significantly easier than the distillation/training the top Chinese labs are doing.

    I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.

llmslave 1 day ago

At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!).

I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.

Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job

  • hnedeotes 1 day ago

    that's why after 1 year of product development of these AI 20x maxxed speed, we reached AGI 'wizards', there's really no difference in output, outstanding bugs no longer get solved and sites still suck, even doing things that were just regular development 20 years ago. Are you sure they aren't only producing 2.5% of your output that you manage just by farting into your phone? Are you sure it's 25% really? Seems way to high, days when I have diarrhoea my AI agents move even faster

    • llmslave 1 day ago

      please keep thinking this so i can relax with my automated job

      • hnedeotes 1 day ago

        It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu

        • mogwire 1 day ago

          If you are too dumb or too lazy to figure out what this guy has, then you are the problem and will be looked at as a relic.

          Watching the AI slop my sales reps put in their emails is disgusting but the reply telling them how great of a job they are doing and how insightful their email was says differently.

          Many people are laughing to the bank while you are still running `--help` to figure out how to run a complex command.

          • hnedeotes 21 hours ago

            Maybe it's you that needs to learn how to run `--help` or ask an AI how to not cry about burnout on open source instead? I don't get it, you should just be cruising on auto-pilot now.

            The problem is retards that can only function on a cocktail of drugs, and as they were never good at anything other than anal retentive stuff built and continue to build these retarded systems. Those peddling RoR apps even when they couldn't serve more than 3 or 4 concurrent requests, JS backends to handle complex workflows that even after 2 years of dev. still have bugs and accrued a sprawl of crap to hide the issues of their own making, etc, and yet charge thousands of dollars, those that write shit software that's not even worth to clean your ass with, even though they have 20 years of experience, but then go give conferences and write books about their amazing architectural skills, those that write utils behind the "oh, it's open source, if you don't like it just fork it" and due to marketing get their crap everywhere, while making holes everywhere for their paycheques. Or the nepo babies that need their mexico border run to get their fix so they can have these "humanity changing" ideas? I bet they're the same that before would weasel a 2 week sprint to change the borders of a button. Or burn through 10k in meetings for irrelevant crap. Or get VC funding for a CSS styling company or a two prompt company. Or go on about the value of ideas, but then can't even get that going without outsourcing or an AI to help them have those same "ideas".

            Ultimately, you just need to turn into a little pig and party in the pigsty, it's not that difficult either, they say pigs are very close anatomically to humans.

            At least AI can help untangle the crap the anal retentive retards have built, and thank god, the pig-mor, this society can't even fuck to replacement levels (perhaps they'll manage now with AI).

yipinwong 1 day ago

Incest in human biology causes mutations that's bad in the long term.

Same for AI models trained on Kimi-3 or other models like Chinese models do. They suffer from the same issue.

ltsSmitty 1 day ago

Well written and good diagrams. No idea the verity of the TMBB (trust me bro benchmarks) but it was pleasing to look at

gigatexal 1 day ago

Not avail on open router?

microdrum 1 day ago

Do they have anything as good as Amp (which is able to use free models)? Amp has been in the lead for almost a year now and doesn't seem to be relinquishing it.

m3kw9 1 day ago

I'm using Codex, Gemini etc, they all have desktop apps and have a plan, how do i use SWE-2? Thats is a problem they have. I'm not about to switch out my workflow and plans with a shiny LLM that looks benchmaxxed and graph maxxed.

wqash71 1 day ago

The horrible website is made by Claude or Cognition is distilled. I'm so tired of it all.

_doctor_love 1 day ago

SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.

  • _doctor_love 20 hours ago

    Odd to get downvotes simply for sharing my experience. Like it or not, Cognition has a good frontier model and they are building serious products. They are working hard which is how you become successful. Sorry if that ruffles your feathers.

gexla 22 hours ago

If everything basically rivals Fable, then why is everything still using it for comparison?

  • MaxikCZ 22 hours ago

    Have you spent at least 10 seconds thinking about it or are you asking just out of spite?

    • gexla 21 hours ago

      Let me grab my calculator and add up the time I have spent reading about model releases since Fable has been released. It seems they all place themselves relative to Fable. I'm sure that time has added up to far greater than 10 seconds. At some point, it ceased to be a meaningful differentiation. This is especially true when I put the model through real usage.