simjnd 9 hours ago

Deepseek V4 Flash 0731 was such a massive jump in capability for such a small model (and price), that I'm a bit disappointed by this release.

I keep my agents on tight leashes, using them very interactively for bouncing off ideas, architecture, and then writing code (especially prototyping) and Flash has been crushing everything I ever needed it to do.

Maybe my ambitions are too tame compared to people needing Fable / Sol grade models, but I'm probably staying on Flash and not moving on to Pro for the foreseeable future.

  • srigi 6 hours ago

    People when OW LLM looks good in benchmarks: benchmaxxxed

    People when OW LLM looks mediocre in benchmarks: disappointed

simonw 1 day ago

Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

  • wolttam 1 day ago

    I think I saw a better overall composition out of Flash 0731

    Effort on this one?

    • simonw 1 day ago

      Default effort for OpenRouter. I'll try a grid of efforts...

      Wow, the low, medium, and high pelicans came out in surprisingly different styles: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

      • segmondy 1 day ago

        exciting. it's almost like 3 models in one. that variety would matter when trying to solve a creative problem.

      • p1necone 21 hours ago

        It's interesting that all three of those used roughly the same amount of tokens, and almost entirely output. Feels like the thinking level lever didn't alter cost at all for this specific task, even though it did change the output.

        • Lerc 20 hours ago

          That raises the question of what is it actually doing?

          If it isn't spending tokens on quality, is it the assumptions about the task difficulty that cause it to perform better? Or are their broader differences in the model being run.

        • ComputerGuru 18 hours ago

          I never trust OpenRouter to forward parameters correctly and would only ever conduct benchmarks with the official api, personally.

          • ljlolel 11 hours ago

            should use an open source one you can audit like my site TrustedRouter

  • delduca 23 hours ago

    You don’t need the best model in 99% of cases…

    • jug 23 hours ago

      This is true and is only becoming more important the more they improve. I am already moving to checking so they're at least somewhat following the status quo and otherwise prioritizing price and platform. I think this will be an emerging way of viewing AI in 2027 and the winner will probably be open models and China.

      • andy_ppp 21 hours ago

        I think this likely plateaus and we all just get the smartest intelligence humans need running locally…

        • f6v 11 hours ago

          Don’t think it’s happening any time soon for most people. My Mac has stayed 32gb for many years now. I don’t think I’m moving into 128gb territory any time soon with all the price hikes.

    • f6v 11 hours ago

      You don’t need Fable or Sol to execute tasks. However, you need them to supervise and plan. Like, Luna is cheap and is at DSV4F level, but it’s not capable of advanced reasoning.

  • khimaros 21 hours ago

    these links never work for me. always "Error: Enter a valid URL" when opening in Firefox. maybe a URL escape issue with Glider?

    • ticoombs 19 hours ago

      Can confirm. Also use glider which seems to double encode. I have to open the comment in Firefox/browser and then click on the link

    • anigbrowl 4 hours ago

      Is this really the right place to bring up an issue with a 3rd party frontend for HN?

  • nhecker 20 hours ago

    On either side of the front wheel is a perfectly reasonable place to carry cargo. I think I'd have taken more issue with the spokes, or at least that's what stood out to me. The chain is indeed nice, however.

  • maleldil 17 hours ago

    Your tool is giving "Error: Gist API returned 403"

  • polynomial 16 hours ago

    Honestly they should all use their respective pelicans as their logos. Or maybe a browser plugin to do do that on the Hugging Face and OpenRouter sites.

  • alpineman 15 hours ago

    I assume this isn't watermarked...

  • jeswin 15 hours ago

    For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.

    • schafberg 14 hours ago

      becuase most people don't care whether it's accurate, as long as it looks right and is funny...

      • m00dy 13 hours ago

        It doesn't look right at all.

    • zero0529 13 hours ago

      Well it is just a bit of fun I think. However, I also think an AGI or an extremely capable model approaching AGI would be able to paint a pelican on a bicycle fairly easily. So in that way it is a good metric.

      • rng-concern 8 hours ago

        I agree it's fun, no argument there.

        However, it's no longer a good metric, as "drawing svg pelicans" is now showing up too much in the training data, so is not proof of generalization.

    • pyaamb 13 hours ago

      I prefer the Browser OS test

    • amelius 10 hours ago

      At this point I want to see some human-drawn pelicans on bicycles. I suspect the LLMs aren't doing all that bad.

    • setsewerd 7 hours ago

      As others commenters said, it's amusing. But also the person you're replying to is the guy who created the pelican test in the first place and I appreciate the whimsy he brings to the discussion.

    • kanemcgrath 7 hours ago

      Because of all the svg rendering stuff, I added a draw_svg tool to my harness and it has been really nice to get a quick mock-up of ui changes. And conveniently, the new DeepSeek models are really good at knowing when to use it. So I do look at pelican rendering as a small metric of useful capability.

    • bean469 3 hours ago

      Because seeing a pelican on a bike is always a good time. Look at it go

  • crazybonkersai 14 hours ago

    Looks like a belt driven bicycle to me. :D

  • Melatonic 14 hours ago

    Not bad but the left foot is still in the wrong place

  • ionwake 13 hours ago

    we hitting singularity levels of bicycle chain here

  • jiscariot 10 hours ago

    Wonder if we'll ever see optimizations for pelican riding a bicycle svg, make it in to model training runs.

    • neuroticnews25 9 hours ago

      Wondering about this in simonw pelican threads is part of the tradition.

  • biztos 6 hours ago

    Basket? Fish? All I see is the model recursively running itself locally on an eye-pad, which for some reason beyond our understanding is obscuring the invisible fork.

  • Melatonic 4 hours ago

    Might be time to also have it try more three dimension pelican rendering. Or a short animated version (even just a few frames) still in standard SVG

freakynit 1 day ago

Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one...

Tested this model, and gpt-5.6-terra-high.

Results: this one had few issues. terra: none.

These results are consistent with my past observations with the latest flash version as well. What benchmarks say, vs what I've been observing are different.

They are good till the project is simple... not anymore.

  • shimman 1 day ago

    I've always wondered if I was using containers wrong because none of them I've ever had to create were complicated. Maybe it's because I choose tools that make local development easy (Go + sqlite + various CLTs) or maybe it's because I never hard to interact with this on the professional side outside of making images for our projects (which still weren't complicated for the reasons above).

    LLMs make containers in a pretty workable format for me (still hand tweak the env variables for a sanity check).

    How exactly does it struggle here and why does postgres need to be built? Were the needs beyond what you get in a base image?

    • freakynit 1 day ago

      This was the repo: https://github.com/amalshaji/portr

      And this was my gh issue: https://github.com/amalshaji/portr/issues/308

      And below was my prompt:

      """ give me single docker-compose file that i can run on my server to run current project... you can read README.md , and then, this relevant page: https://docs-custom-reverse-proxy.portr-docs.pages.dev/docs/... ... this was the result of me raising github issue: https://github.com/amalshaji/portr/issues/308 ... you can use gh cli to fetch the details and comments...

      i already have a caddy server running on my vps... and i will create wildcard certificates myself using certbot.. the domain name will be helloportr.xyz ... also, ports up to 9019 are already taken...

      ask me if anymore info is needed... """

      You can try yourself and let me know of what you got.

      • shimman 22 hours ago

        This is definitely beyond my capabilities lol but wow portr is a neat project. Never heard of it before, only the paid services from tailscale/cloudflare.

      • arch-choot 18 hours ago

        I've been using DS4F+Pi with great results, but I think one thing that helps is at the end of my prompt I'll tell it how to verify it, e.g. "Make sure the compose file works by running it locally (use self-signed certs if required)".

        The argument could be made that "the model should be smart enough to figure it out" , and maybe DS4 isn't. But with just a bit of steering you can get the correct result for like 1/10th the cost, or even cheaper.

  • npn 1 day ago

    wait for Deepseek Harness (yes it is the official name) release then try again.

    for your kind of task, harness tools matter.

    • freakynit 1 day ago

      I used pi

      • natrys 1 day ago

        For me, flash 0731 was much better in omp/opencode than in Pi.

        Anyway, it might be so that they are rolling out deployment. There haven't been an official announcement post yet (this submission is a link to openrouter). Some people have been saying they are getting results worse than GLM-5.1, that's obviously broken.

    • gkbrk 1 day ago

      If the model cannot figure out simple and ubiquitous tools, how is it supposed to figure out complex problems? All of the good models basically work with any harness, including giving them a single "shell command" tool. They can just figure things out.

      • hadlock 1 day ago

        When it comes to quality of outcome, since at least Feburary, the harness has almost equal, if not more weight than the model itself. It's no longer "which model is the best?" it's "which model + harness is the best?"

        I get drastically different tool call failure rates using Claude SDK vs OpenCode using Qwen 3.6 models

        • azinman2 1 day ago

          Which works better for you?

        • KronisLV 1 day ago

          > the harness has almost equal, if not more weight than the model itself

          This feels like a horrible failing of the models to generalize, then - both basic and intermediate tasks should be possible to do with Claude Code, OpenCode, Pi, ZCode, Kimi Code, Dirac and tbh any other mainstream or even slightly niche harness. Not doubting the claim itself, there's a reason why good benchmarks include the harness.

          • dominotw 1 day ago

            i think thats BS that harness has equal weight. most of intellegice is still coming from training data not from RL. so how is 'coevolved harness' equal weight.

            • disgruntledphd2 11 hours ago

              > most of intellegice is still coming from training data not from RL.

              For coding specifically, I'm not sure this is still true. Given the heavy use of RL to improve coding performance, I'd expect the harness to be important as it defines what tools the model is rewarded for using.

              • dominotw 9 hours ago

                i belive they looked at the traces and they were still legible ( RL traces should be gibberish)

        • davidlt 1 day ago

          I just wanted to emphasize this. Harness is a big part of how things perform thus usually it's harness + model co-design that's important.

        • HDBaseT 23 hours ago

          Yeah this is a complete lie.

          You can use effectively any harness and get good results. Harnesses are mostly placebo.

          • scrlk 22 hours ago

            With Opus 4.7, there's a 10 point improvement in the AA coding agent index when you swap out Claude Code for OpenCode: https://artificialanalysis.ai/agents/coding-agents#harness-c...

            To put that in perspective, the difference between GPT-5.6 Sol Max and 5.6 Luna Max is 8 points. That's a lot of extra performance that you can get for free just by using the best harness.

        • RideOnTime22 23 hours ago

          Every other week it's a new "X didn't matter, until Y date" without any hard quantitative claims.

          It's crazy how over the past years a field originating from math ends up succumbing to subjective feels.

      • sheeshkebab 1 day ago

        This. The same goes for “skills”, skill type “subagents” and other bullshit - powerful models don’t need any of that anymore I noticed.

      • npn 1 day ago

        I don't think so. there is a lot of tools with similar usage, some harness even bring their own internal tools for accurately manipulation.

        also, even if some models claim that they have full 1M context window, some only work effective with the head or tail of the window, a proper harness tool will know about the limitation of the model and act accordingly.

        then also the output format, the tool calling syntax, the quirks and gotchas of each model.

        it is not simple as just throwing everything at the model, especially when your project has hundred of files or so.

      • derefr 1 day ago

        Because complex problems can be decomposed (a skill in itself) into easy parts and hard parts; and the hard parts are almost always bottlenecked on understanding concepts and principles (i.e. things that are either in a model's weights, or not), not on having certain facts available. Models can solve complex problems insofar as they can decompose those problems, and have learned the concepts and principles relevant to approaching the hard parts of those problems.

        Whereas tool-use isn't a capability problem, but a context problem: the thing that makes models fail by default is that they have no idea, when first summoned out of the aether, what kind of conversation they're having, who it's with, what that person is trying to do, what tools they have available, and how those tools can be invoked.

        Think of the difference between how you'd respond to a casual programming question asked by a person sitting next to you on a flight, vs. a programming question asked of you by someone you're pair-programming with with your IDE open in front of you. Now imagine waking up blind and deaf and needing to discern which of the two situations you're in. LLMs know how to approach both of these problem-contexts (and more besides), but they need to be given context to know which problem-context they're in (and everything else about that problem-context: which IDE they're using, which OS it's installed on, what other tools are installed+accessible, etc.)

        And before you say "but why can't they just experiment to figure these things out" — if you think about it, knowing how to interface with a shell and an IDE are bootstrapping requirements for any kind of experimentation, in about the same way that "knowing how to open your eyes and move your head" is a bootstrap requirement for a human gaining information about the world around them. These capabilities are necessary to explore the world to "discover" and "probe" other capabilities.

        ---

        Also, a lot of the work LLMs do "needs" (i.e. is heavily improved by the use of) some kind of structured scratchpad, that they have been trained to manipulate and "look at" through tool-use. Even for a human who could accurately visualize a canvas based on a coordinate system, you still wouldn't expect said human to succeed at the pelican test if they had to write the SVG entirely in their head and then write it out sequentially with no rewinding to fix mistakes. You'd expect them to ask for at least a whiteboard, if not a text editor, to be able to write and rewrite the SVG XML.

        (Really, they'd ideally want to run the SVG and look at it to see how close it is, and optimize that way. I'm not sure if we're letting LLMs do that part in the classical pelican test. It feels like that would vaguely violate the "zero-shot"-ness of the test, though I'm not sure if we're currently considering a conversation to be "zero-shot" if it involves the model iteratively interacting with a third-party system [such that there are repeated model -> system -> model conversation turns] but holding off to responding to the user until they think they've fully solved the problem.)

        ---

        And also, on a lower level, all of these external capabilities are getting exposed to the LLM through MCP. Models can and do understand how to speak MCP itself. But there's no standard for how a given harness's capabilities (e.g. "execute command line in new shell session", "send patch edit command to active tab in IDE", etc) should be modelled to be exposed through MCP, either in their encoding or in their semantics. There's no MCP equivalent of WASM's WASI meta-standard, such that models could learn these specs and "assume by default" that things work like them until told otherwise; and nor are there even open harnesses that LLMs could learn about during training, and through them, learn some de-facto MCP-endpoint specs. Instead, there are mostly just proprietary harnesses, that hide all that info from public access, sharing it only with the LLM during inference, and even then, only at the moment the LLM needs it.

      • segmondy 1 day ago

        the single shell command is the terminal bench.

    • ghm2199 1 day ago

      I use pi harness with codex and all the tool calls are custom delegate extensions, I mean ALL(for security checks), i get consistently good results from sol on high and xhigh reasoning. I don't believe harness should matter because its at most just a way to abstract tool calls and maybe the system prompt. Training on the tool calls results should not(and in codex's case does not matter)

  • scrlk 1 day ago

    What harness are you using? DS V4 is harness sensitive.

  • derangedHorse 1 day ago

    Terra has not been able to do any of the technical tasks I've asked of it correctly. I'm surprised others get use out of it. Anything below Sol high tends to give me mostly unreliable results. I'm using codex as my main harness but maybe it performs better with a different one.

    • freakynit 1 day ago

      Depends on project complexity. For one of my more complex projects, I exclusively use sol-high ... nothing below that works correctly.

      For this however, a comparatively much simpler task, tarra-high works fine.

      • Foobar8568 1 day ago

        Right now, sol-xhigh is my favorite model. I feel that Opus 5 is dumber than 4.8. Fable is too expensive to do anything (limit of $50, started a prompt at $25, ended up at $75, is bullshit, but at least it's "free credits").

        DeepSeek is okay for random API-based stuff, as it's cheap.

        Local open models running on a 5090 are hit or miss. I feel that most GGUFs/quants are awful...

        • miohtama 1 day ago

          Opus 5 degrades to word salad.

          I wonder if it is because of watermarking.

          • SwellJoe 23 hours ago

            Opus 5 doesn't really even speak coherent English. I'm not sure what's going on, but it can't explain anything. It still does an excellent job with code and writing tests and code review and creating and completing a plan, and it seems to be able to understand English instructions, but it sure as hell can't explain what it did or how to use the code it wrote.

            That was true before they announced the watermarking, I'd already started to back off of using Opus as much because I like to understand what the model is doing and have it write documentation I can use to reproduce its results, but maybe watermarking was already in there unannounced.

            • sscaryterry 13 hours ago

              This 100%. The incoherent drivel has literally made me move away from CC completely.

          • mlrtime 10 hours ago

            Agreed, I'll get a terminal full of text and I just reply "I don't understand this"

            And it retypes it for a human, I'm doing this more and more lately.

          • pulkitsh1234 9 hours ago

            Ah, so it's not only me :-D

            I think it could be the watermarking, but at this point they might be deliberately complicating the prose so that we ask clarifying questions and that leads to more token spend.

        • ericfr11 23 hours ago

          I am still on Opus 4.8, with a custom built harness and it works very well even on multi-repos, across stack, deep changes. I also have a very solid test suite which is helping the coding agent a lot

    • cyanydeez 1 day ago

      the breadth and width of the universe of oneshot challenges are all arbitrary. It's unsurprising different workflows oneshot better than others.

      All the more reason to favor local models under your control, as once you find that sweet spot model, no one can change it, upgrade it, align it, take it down or otherwise harm the time investment you made it making it your own.

      I can't really believe no one understands, after decades, how valueable a rock solid development environment is.

      • Phemist 1 day ago

        Exactly! I am not opposed to cloud-based models, but I do only stick to open-weight models because I know I can move my whole stack to local (given enough hardware) and continue development without any of the LLM interaction contracts being broken.

        I would like to see some development where proof of authenticity certs are generated alongside the actual output of the model. Prove to me (or at least claim to me liable to breach of contract) that this output was generated by FP8 DeepSeek V4 Pro 0813. Not some cheaper quantization of the model.

      • f6v 10 hours ago

        But I don’t want to manage a local model…

        • cyanydeez 8 hours ago

          Likr sll infrastructure, once its setup and your flow is set, the only thingin the way is FOMO.

    • bob1029 1 day ago

      I am seeing essentially deterministic results with Terra running a custom browser automation agent across >100 interaction events.

      The harness is everything. If I just threw something like Codex at this and said "good luck" I wouldn't make it beyond 5-10 interactions. I tried that already. Carefully designing the views and tools over the environment is where you can go from 50% to 99.9999%.

      • tanishqkanc 1 day ago

        Curious about what you found. I agree harness for browser automation is vital - I work on https://libretto.sh

        • bob1029 22 hours ago

          Hand-crafted adapters that sit between playwright primitives and the agent loop are the secret sauce. The goal is to insulate the agent from the raw DOM without any loss in fidelity regarding the logical business information and available actions.

          • drewnick 17 hours ago

            +1 I have found extremely reliable systems require a mix of deterministic "adapters" is a good word for agents to actually get through the workflows I've created. I'm still amazed that it can work with both those and some "intuition" to bend the rules around the adapters if prompted.

    • mixedCase 1 day ago

      With Pi as a harness I've been using OpenAI models as a worker with an Opus 5 (in Claude Code) planner. I've only had a few issues with Terra High/Medium and absolutely none with Sol Medium+ on a fairly complex Rust project that targets Linux, Mac, Windows and Web, with plenty of nasty FFI, VMs, remotely debugging systems, among some other things within a monorepo.

      I think the key is to give them a nice assortment of self-verification tools, an AGENTS.md or reference document that they're encouraged to routinely check, and asking the planner to be thorough with the ACs but give the model some space.

      The planner routinely finds issues with the worker's output, but that's what it is for.

      • sejje 1 day ago

        Plan with sol-med, implement with luna-high. Rarely a problem.

        • ericfr11 23 hours ago

          Same for me, with Claude Opus/Sonnet. All the models are almost equivalent if well steered

      • ericfr11 23 hours ago

        Harness is the key. I built my own to "talk" our institutional knowledge and it's working great

    • zeven7 1 day ago

      I bounce between Sol high/medium and Luna max. I don't know why you'd use anything between Luna max and Sol medium. Luna is so extremely cheap and cranked up to max it does anything I'd want Terra to do for a fraction of the cost. What is Terra for?

      • Juvination 1 day ago

        One thing I've really noticed with Luna Max is its speed. I've got a review script setup on a custom Pi extension. Luna finds some issues/some false positives, while Sol finds issues but disregards false positives. The biggest thing is Sol finishes in about half the time.

    • Art9681 1 day ago

      Terra is great. It's wild how different our experiences are.

      Install the Superpowers plugin.

      Behold.

      • MuteXR 12 hours ago

        In my experience, Superpowers has begun to massively slow down the capable models at this point. The skills they add are incredibly bloated and just get you worse results nowadays, tbh.

        • agrippanux 10 hours ago

          I loved Superpowers and evangelized it heavily.

          This week I removed it because it now gets in the way of the frontier models.

  • apitman 1 day ago

    Wait people use terra?

    • miohtama 1 day ago

      I use mostly Terra. Much better than Opus 5. Much more token mileage.

      • apitman 23 hours ago

        But why? Luna Max is almost the same intelligence as Terra xhigh and way way cheaper. And Terra max is almost the same as Sol high. I just don't really see a place for Terra but slower.

        • miohtama 15 hours ago

          This is a good tip. I will try. Thank you!

    • smb06 23 hours ago

      My company pretty much exclusively uses Sol and Luna

  • ApolloFortyNine 1 day ago

    I use deepseek flash to do exactly this. Git repo (which I usually have it build from scratch) -> build docker image -> deploy to server with komodo/caddy-docker proxy.

    Works great, regularly one shot applications. I often make changes to the application after its deployed (to be fair, my prompts are usually quite laxidasical, just 'build x, use /deploy-to-komodo) but the deployment works great.

    I did make a skill, but if your doing anything repeatedly you should as well.

    Opencode, but any harness I'd think would work similar.

  • yassa9 1 day ago

    did you test kimi k3 or qwen 3.8 max on the same task ? or plan to test them ? I respect those genuine users tests other than those benchmarks that models are trained and overfitted to them

  • v3ss0n 1 day ago

    I do that kind of things all the time with Qwen 3.5 122B. It works well in one shot with Cline or Opencode.

    May be your harness problem?

  • amelius 1 day ago

    I didn't understand your use case, so it could also be the way you write your prompt, I suppose ...

  • celsoneto07 23 hours ago

    I've been doing pretty heavy stuff with DeepSeek with a good degree of success. The thing is: I don't trust it to go fully autonomous. I check the steps, I steer it. For the pricing, it's worthy. Let's how the price increase is going to change my behavior.

  • shunia_huang 19 hours ago

    The flash model will always use an outdated Treafik version that is not compatible with the newer docker engine, I tried to deploy some personal services with Traefik and everytime it uses this wrong version, and then fixes the version issue in the thinking chain.

    I was thinking to switch to Caddy but with your experience I'm gonna stay with Traefik and bare with the version issue...

  • acchow 15 hours ago

    I agree. The “frontier level” open weights models benchmark really well but fall behind in real world performance

monster_truck 22 hours ago

Have been letting it spin pretty hard (~$12.50 for 2B, 50% cache hits) on my traffic simulator/distributed physics engine all day, it's found some pretty significant gains without introducing any new problems.

I'm happy

  • blahyawnblah 22 hours ago

    Can you tell me about your engine?

  • p1necone 21 hours ago

    50% cache hit is really low - in a standard agentic loop you should expect like 99%+ cache hit percentage (which should also lower that $12.50 to like a couple of $ for the same amount of tokens).

    If you're using a customised harness you should make sure you don't have something that's e.g. changing your system prompt on some requests or rewriting history - it can be tempting to do stuff like strip old thinking tokens or compact tool call results to reduce context size but it's a trap - you want to never change history because of how cheap cache is, even more so with deepseek because their cache hit pricing is so low compared to most other models.

    • Lalabadie 21 hours ago

      In my experience, that's the OpenRouter tax. Even a session that does everything right to remain sticky ends up getting moved between providers on a few requests, which bills you the full context as input every time the switch happens.

      I assume it's done as load balancing/latency mitigation, but it's put me off of OpenRouter for my use cases (limited use, limited need for changing models).

      • grono 20 hours ago

        Why not pin to specific openrouter provider and disable fallback?

      • julianz 20 hours ago

        There is only one provider for this model, so shouldn't be running into that.

      • p1necone 19 hours ago

        This has not been my experience. Generally I do pin to 1 provider, or 1 provider with a couple fallbacks (especially with deepseek - most providers are 10x the cached token price compared to deepseek themselves), but even when I don't I still usually see 99%+ cache hit percentage. Specifically using pi with various ad-hoc customisations (that I was careful not to break prompt caching with).

        • irthomasthomas 16 hours ago

          then what is the point of using operouter for this model? Just use the deepseek API and save the 5% fee on top of the better caching rate.

          • FridgeSeal 13 hours ago

            Because they don’t want to sign up for 10 different providers and subscriptions/etc, especially if some models are just going to receive light, or rare usage?

      • wut42 18 hours ago

        It is a pain from openRouter if you don't define your providers correctly, but for DeepSeek, surely not- the weights aren't released yet and there's only one provider, DeepSeek.

        • Lalabadie 10 hours ago

          With Deepseek as the provider, there's no issue of course, but that means you don't filter providers for data retention, and you could also choose direct API use with them at that point.

          • wut42 9 hours ago

            Fallback are still very useful and won't poison much your cache hits too much if the provider is down anyway.

      • seunosewa 10 hours ago

        You can set it up to always use the official provider.

      • isqueiros 10 hours ago

        Seems like pro 0813 is exclusively served by Deepseek themselves at the moment so I wouldn't say that's the case?

  • fooblaster 18 hours ago

    Can you explain how you used 12 billion tokens to do useful work?

    • eru 18 hours ago

      (Not the original commenter.)

      You can rack up quite a lot of tokens if you ask it to try out a lot of things, eg for performance investigations and trying out optimisation ideas.

      • fooblaster 18 hours ago

        is there a standard pattern for this? Like spawn an agent for each technique to try?

        • eru 17 hours ago

          Not sure. I usually tell the agent to spawn subagents at will. (And they are doing that on their own anyway.)

        • laichzeit0 16 hours ago

          I regularly do a “go to DynaTrace, look at how this service gets used in production then use a profiler and see if there’s any low hanging optimisations we could make” on stuff. LLMs are really good at doing everything that was fun about software development.

          • eru 15 hours ago

            Ha, I have a whole hobby work-stream going on about finding low-hanging fruit in various open source projects to turn into valuable contributions.

            Two premier sources: (1) look at good contributions someone already tried to make, but that got stuck in review or were otherwise abandoned. (2) look at user reported bugs and see if we can find a user reported bugs, and see if we can reproduce and fix cheaply.

            Both are explicitly scoped as best-effort affairs: move on, if you can't quickly make progress.

            Most of the work I have to do as a human is review and navigating the submission process: tokens are cheap these days, so you really need to make sure the contribution is actually worth someone's time to review.

alecsm 1 day ago

I've been using the last Deepseek Flash update for a week and I'm amazed. It was a capable model for easy tasks but now it looks like it can do some heavy development for peanuts.

I can't wait to try this new one.

  • coredog64 1 day ago

    IME I can't trust it to write it's own plans from a spec, but if I give it a detailed execution plan written by Opus, it's fast and cheap (if chatty) in executing it.

    • stavros 22 hours ago

      This is what I do, and it works fantastically well. Just make sure you have Opus/GPT review after.

    • polski-g 21 hours ago

      Interesting. I use Flash for making the plans and GPT for execution.

      • Kadin 20 hours ago

        Depending on the language you're writing in and the problem domain, the smaller models can do dramatically better or worse.

        I suspect in the future we'll see language-specific small models. "Coding" is still pretty broad as an activity. It'd be nice to be able to load up a model specific to, say, class-based Python and run it on-device.

        • eru 18 hours ago

          Harmonic's Aristotle is a sort-of language specific model for Lean, if you want to see the future you described today.

      • jatora 18 hours ago

        flash for plans?! i don't understand why you wouldnt use something far stronger for the most load bearing point of the project

        • tokai 11 hours ago

          Price obviously

        • ricardobeat 11 hours ago

          There aren’t many “far stronger” models than Flash 0731 now, it’s only beaten by Claude and OpenAI models at high/max effort, and everything that matches it costs 5x-10x more.

  • xnyan 1 day ago

    I find DeepSeek flash incredible for the price and good in general if it has good plans. I will typically plan using Opus or GLM, then implement with DSF

book_mike 1 day ago

What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.

  • okamiueru 1 day ago

    How do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.

    • bikemike026 1 day ago

      If you read Opus 5's output, it is beyond the comprehension of virtually all engineers and developers. That is what I mean by intelligence. Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.

      • logicchains 1 day ago

        You mean Fable 5 right? Opus 5 makes lots of stupid mistakes about anything that requires any domain knowledge.

      • hgoel 1 day ago

        I don't think that's because of its "intelligence". It speaks obtuse techbro-ese: stringing together words that sound smart to obscure the simplicity of the thing it's describing. In many ways it's the opposite of intelligence.

        Opus 5 and Fable 5 in particular suffer from this issue at worse level than most models in the same class.

        • greenchair 1 day ago

          yep, it is so bad i had to create rules to cut down on the techbro language and domain slang.

          • hgoel 1 day ago

            I just canceled my Claude subscription outright. The models are all gairly fungible, it's easy enough to just switch to another provider.

        • sebastiennight 16 hours ago

          This is a "load-bearing" issue recently.

          I think the idea is packing more information into fewer words, but the result is a word salad that is somehow simultaneously very dense in adjectives and adverbs, and still way too verbose.

        • qlte 6 hours ago
            > stringing together words that sound smart to obscure the simplicity of the thing it's describing
          

          Laser targeted at LessWrong posters

      • okamiueru 1 day ago

        I'd have to ask for you to be more specific, otherwise, to take your answer at face value, it comes across as a contradiction.

        > [Opus 5's output] is beyond the comprehension of virtually all engineers and developers

        That would make it pretty bad? The key defining quality of good software, is clarity, and the ability to simplify a complex problem to the point of it seeming trivial.

        > Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.

        The bar here should absolutely be to judge this against the expert level within each domain. I have time and time come across LLM output being woefully underwhelming in every single request where I am an expert. For all areas that I am not, it sure seems plausible. It is far more likely than not, that it is equally inadequate in the areas I lack the necessary knowledge to tell.

        If the AI is being subpar in every field and category compared to an expert in said respective field, then, what a strange gauge of a tool's usefulness. Are we attributing higher value because a single model is "attempting to solve all knowledge and fields at the same time", why is that of any importance, or excuse?

        We should not define "intelligence" as how effectively it can convince a non-expert of something being plausible. That sounds like the absolute worst tradeoff. You'd have to waste the experts time in filtering and refuting incorrect postulations that are cheep to generate. The perfect storm for bullshit asymmetry.

        • bikemike026 1 day ago

          I disagree with points 1, 2, and 3. Point 4, AI is better than average, and sometimes it's better than excellent. Point 5 is irrelevant.

          • okamiueru 11 hours ago

            That seems awfully self deprecating. Surely, you expect better of yourself in at least some area, than the average competency of humans across all areas?

      • nwienert 20 hours ago

        Careful, you may have a bit of psychosis. They are very, very far from incomprehensible, and also very far from the top at least of my field. The best in my field are produce far higher quality results, and I think that's true for all fields. It's just an incredibly good 85% quality machine that experts all use because they can guide it to be up to their quality faster than doing it themselves.

        • okamiueru 11 hours ago

          You could take that even further, to the actual danger of reliance of these tools when you lack the expert knowledge. That is, when you assume it took you 100%, but missed the 15% it got very wrong, or perhaps even worse: subtly wrong. This compounds with the next similar task, and either you've made the actual experts quit their job as it has become to babysit LLM output, or you end up with an unusable mess, deleted production databases, etc.

      • scrollop 16 hours ago

        If you read the many, many complaints about opus 5 on anthropic forums, the sentiment is that opus 5 output is poor and people are back to 4.8 and 4.6.

        You may want to re-evaluate and compare to the older models.

        • pmontra 16 hours ago

          Opus 5 is too verbose.

          I'm using Sonnet 5 on a large porting project and it's good. I switched from Opus 5 to Sonnet 5 on a project of another customer and I didn't notice a decrease in quality. I concede that it's very difficult to assess a difference in quality unless one uses both models on the same task and carefully compare the code, not the output in the terminal. I really don't have the time and the tokens for that. Anyway, Sonnet is still doing a good job.

    • f6v 1 day ago

      My definition is that I can be much less precise with AI the more intelligent it is. It can extract the intent from my fuzzy description of the problem. Which means I can offload some of the thinking effort.

      It wasn't possible a couple years ago. I used to make fun of people who were trying to get ChatGPT to think about the problem when all it could do was write code from the pseudocode you provide.

      But now I can say: "Look at the latest log and make a plan to fix". And it takes it from there.

      • eru 18 hours ago

        Sounds like a good working definition in the context you are using it in.

        > But now I can say: "Look at the latest log and make a plan to fix". And it takes it from there.

        I usually tell the agents to first work on reliably reproducing the problem in the log, and only then even start thinking about a fix.

      • versteegen 18 hours ago

        I think this is the best and most useful way to measure model intelligence. In my experience it's what really sets apart the capable models from the best. A small model can be RL trained to be extremely good at programming or narrow problem solving for its size (eg 5.6 Luna, DS4 Flash, Qwen 3.6 27B), but even Luna is IME comparatively awful at understanding intent and making good decisions with limited guidance.

      • okamiueru 11 hours ago

        I'm not sure if you are aware as to the extent certain processes and functions are being anthropomorphized.

        These systems are not "intelligent" if you follow the dictionary definition. Hence the question posed to get a better understanding of how it is being used in this context.

        They also do not "extract intent". There is for sure some intent behind your input to the service. What follows is a predictive text that uses your input, together with a LLM trained on a corpus with similar relations, that ultimately gives you a series of words.

        That isn't to say a service like this cannot be useful. But I'm often wondering if the people who rely on these, and are particularly enthused by them, are actually aware that the terms they used are in fact anthropomorphized. I start by giving the benefit of the doubt, but it rarely lasts. 'Reasoning', 'agent', 'skill' 'hallucinate', 'know', 'think', 'train', 'learn', 'understand', 'harness', 'attention', 'context', 'prompt'.

    • odig 1 day ago

      so?????

  • tomr75 21 hours ago

    how are you paying for tokens with this setup + what harnesss + how many tokens/day are you consuming?

  • anon7000 14 hours ago

    Opus 5 fucking sucks to talk to and read compared to 5.6 Sol though. I’m fully done with Claude models until they figure this out

    • dustypotato 14 hours ago

      My god the Jargon is so hard to parse. Everytime i resort to cursing it , it understand. Even putting `ASD-STE100` or simplified english in the claude.md doesn't work. It gets the job done but is an anti social asshole

Hardd 14 hours ago

Based on my experience so far, compared to previous models, DeepSeek V4 Pro achieves results equal to or even better than before, but at a lower cost.

  • thunfischtoast 11 hours ago

    Sounds like something a DeepSeek V4 Pro bot would say

    • KoolKat23 11 hours ago

      Passed your personal turing test.

    • twelvechairs 10 hours ago

      You are absolutely right!

      Would you like me to respond in a more naturalistic way for hackernews denizens?

jklmnopqrstuvw 1 day ago

Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project.

Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.

Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

  • ferongr 1 day ago

    [flagged]

    • numpad0 1 day ago

      no he and his stuffs are now considered transparent, no pun intended. I think he deserves it since his minions were persistent with usage of "this ___ has hateful bias against ___" canned response.

  • computerex 1 day ago

    Repeat the test like 5 times for each model and see the results.

    • epolanski 1 day ago

      +1, a single test means little.

      • jklmnopqrstuvw 1 day ago

        I don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed.

        • seunosewa 1 day ago

          Do it a second time at least.

        • computerex 1 day ago

          They are causal autoregressive models, the output is sensitive even to the implementation nuances in inference. Even 1 token that's badly selected could throw off the entire answer.

          • segmondy 1 day ago

            you're thinking of one shot. if they are running an agentic loop then they don't need multiple passes. an agentic loop is multiple passes with tool calls and tools could fail and agent would correct from seeing the failure. a bad model will compound on error and fail, a good model will correct. 1 test is fine to gauge the quality of the model.

            • computerex 1 day ago

              An agent doing a task even with multiple back to back calls like normal without an example is zero shot. An agent doing a task with 1 example is one shot. An agent doing a task with a few examples is few shot. I don't think you are correctly using these terms.

              The multiple back to back LLM calls are done on accumulating context, so if there is a sampling error it could throw the entire session out of whack, because LLM's build on the previous context.

              It's actually meaningless to argue, one could simply sample more than 1 times and let the numbers speak for themselves.

              • gpt5 23 hours ago

                That's not true. An agent in a loop can test itself, review, verify and iterate as much as needed. That's one of the primary reasons more capable models tend to have a higher success rate.

                I don't disagree that multiple tests increase confidence, but it's not correct to argue that an agent in a loop harness is equivalent to oneshotting

                • techpression 22 hours ago

                  It takes more than that. I had Claude do a feature, it took five minutes, then I had Claude do a code review of its changes, that took 65(!) separate agents and 40minutes.

                  I don’t think single agent loops are good enough.

              • nl 21 hours ago

                > An agent doing a task with 1 example is one shot. An agent doing a task with a few examples is few shot. I don't think you are correctly using these terms

                This is a different thing. Yes, giving multiple example is called "few-shot prompting".

                But one-shot vs few-shot benchmarking is different. In this context "one-shot" means "pass at 1 effort" as opposed to "multi-shot". In the literature this is called "pass@k".

                Anthropic has a good explanation here: https://www.anthropic.com/engineering/demystifying-evals-for... (search for "pass@k").

                In this discussion we are discussing pass@1 (single shot) vs pass@(k>1) (multi shot).

                > The multiple back to back LLM calls are done on accumulating context, so if there is a sampling error it could throw the entire session out of whack, because LLM's build on the previous context.

                This isn't really true. In an agentic loop the LLM can correct itself via in-context learning.

                • computerex 3 hours ago

                  Yes, and the reason why pass@k exists is because of self-consistency. There is no guarantee for right answer to be selected or for the LLM to correct itself. While I agree pass@1 is a useful metric, I'd be more interested to know pass@5 so I can better compare the results.

  • NooneAtAll3 1 day ago

    I thought it was impossible to downvote posts?

    • numpad0 1 day ago

      Maybe a tug of war between flags and vouches might work like downvotes?

    • benjiro29 1 day ago

      I thought it was impossible to downvote posts?

      User Posts can be downvoted but you need over 500 karma to have access to the downvote button. A Submission can not be downvoted.

      • Barbing 20 hours ago

        Yes and-

        Submissions can be flagged by anyone and mods/admins can downweight them. (If I’m not mistaken this is common for, say, Flock posts at the moment.)

        Curiosity & repetition are two key factors.

      • NooneAtAll3 17 hours ago

        comments can be downvoted, posts can't

  • Zetaphor 1 day ago

    It's the third link on the front page right now?

  • bigmadshoe 1 day ago

    Why are people giving these n=1 comparisons like they mean anything? The worst offender is that pelican guy. These are non-deterministic systems and a single trial should not update your priors much at all.

    Of course it's significant that your response had a bug and took four times longer, but if you're only going to try once, this isn't real science, it's just vibes.

    • jklmnopqrstuvw 1 day ago

      Months ago I start making this kind of test for my own reference. At beginning I I test each model multiple times, and results always same(pass or fail). Later I test only once for new models, I trust the results.

      • kees99 22 hours ago

        > multiple times, and results always same

        Not my experience at all.

        With smaller models, whenever I see a response that is going into wrong direction, I would just redo that step, and more often that not that brings improvement.

        This effect is less pronounced with SOTA, but still there.

        • shunia_huang 19 hours ago

          Yes not my experience either.

          I've tried or sometimes be stupid to work on bugs/features and ask with almost identical prompts with same modal and harness set, and yes, they generate totally different results.

          Sometimes the output is unusable and even with extended guidance it will still drift away from what I was expecting.

          Sometimes the output is just one shot and follows almost whatever I want.

          I then be used to work like this, if the model and harness set does not work for one time, I just start a new session and do it again. And currently there is one of my task working like this.

    • gnunez 22 hours ago

      I don’t understand why people are calling these transformers models non-deterministic? Are you referring to the temperature parameter? I haven’t played with transformer internals in a while but my understanding is that if the temperature is fixed at a value where the top logit is always picked, then because they weights are fixed, the exact same input should produce the exact same output. Am I missing something?

      • gpm 22 hours ago

        Well, yes and no.

        By non-deterministic I think people really mean "chaotic" in the chaos theory sense. Small perturbations in the input lead to wild and unpredictable changes in the output. Even with temperature parameters a fixed PRNG seed could mean an LLM was just chaotic and not technically non-deterministic.

        But more literally while LLMs are in theory deterministic (though perhaps not inference providers implementations if there's anything like a race condition affecting how things are rounded when added together) - we use the LLMs in harnesses that aren't. There are very likely races in the terminal outputs, dates both intentionally put in the context and accidentally leaked to the context, things like that.

        • gnunez 21 hours ago

          Ok. I see. I guess people are not referring to the raw models themselves when they say non-deterministic, but are also including the harness used in conjunction with the model. Then, in that case, for the exact same input you could get a non-deterministic output. But the model itself and all the mathematical machinery around the model is still very much deterministic.

          I guess if we really needed to, we could construct a deterministic agent harness. But in most use cases we probably want some chaotic behavior to increase our chances of stumbling on the desired results.

          Thank you for the clarification

      • nl 21 hours ago

        > if the temperature is fixed at a value where the top logit is always picked, then because they weights are fixed, the exact same input should produce the exact same output. Am I missing something?

        Yes.

        Your input is part of a batch, and you don't know where in the batch it is. By default batches are not invariant and VLLM only supports invariance at all on some Huwaei Ascend hardware.

        See https://docs.vllm.ai/projects/ascend/en/latest/user_guide/fe...

        • gnunez 21 hours ago

          I totally missed the memo on batching. That changes everything. Thank you for the info.

      • sejje 21 hours ago

        I think people are wrapping that across the English language. In English, these two tasks are exactly the same:

        "Would you hand me that item?"

        "Please hand me that item"

        But when posed to the LLM, they generate different outputs. One character difference in the prompt might be a whole different output. People who aren't programmers mostly don't know that there's any difference. They asked for the same thing, it knows what they want in both cases...but different results.

        • esikich 18 hours ago

          I'm not sure that's true. Sure, in the end I might hand them the item, but my thoughts about what they said will be different. I think you have to consider my thoughts "output" for this comparison to be valid.

        • sebastiennight 16 hours ago

          Not to be too pedantic, but these requests would not be exactly the same.

          There is a bit of indexicality in "Would you hand me that item ?"

          that might cause it to be interpreted as an actual question rather than a request, and might elicit different responses:

          - maybe _I_ would not hand this to you (I'm busy right now), but the person next to me whose hands are free would, so I'd nod to them. However, if you had said "Please hand me that item" I'd put down what I was doing to comply.

          - maybe I would not hand _this_ to you (it's not the right tool IMO), but I'd suggest another option. However, if you had said "Please hand me that item" I'd put my doubts aside to comply.

          - maybe I would not hand this to _you_ (you're not the one who should be handling it), but I'd do the thing myself or hand it to a more qualified member of the group. However, if you had said "Please hand me that item" I'd trust you enough to comply.

          I think this distinction is relevant in that I've found people to sometimes have difficulties understanding how similar LLM prompting is to giving instructions to human colleagues.

          I've had a collaborator who though very highly of his own prompting skills (while his prompts were very ambiguous and of the "make no mistakes, erase everything & correct yourself if you find one" variety) and blamed the models for not being "smart enough", and it was very noticeable that his management style for the juniors on his team was similarly unproductive.

        • MarceColl 15 hours ago

          Yeah, that's bullshit.

          Depending on my mental state, status with the person and many other factors each of them may trigger both many different internal thoughts, looks, body expressions and even outcomes.

      • neosat 21 hours ago

        As the above two comments mentioned this is not true in practice due to batch effects (you can read about some interesting work published by Thinking Machines on this), as well as calculation drift that happens across computations esp. now with inference optimization becoming common.

      • LeBit 9 hours ago

        You could answer your own question really, really quickly.

Gecko4072 1 day ago

Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.

  • Jsttan 1 day ago

    What is the new price through?

    • Gecko4072 1 day ago

      https://api-docs.deepseek.com/quick_start/pricing/

      edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount

      • minraws 1 day ago

        isn't it the same old pricing? did they increase V4 Pro pricing already?

      • nchmy 1 day ago

        i dont see any price increase there... what am i missing?

        • alecsm 1 day ago

          Right below the pricing it is stated that they plan to increase the prices in the near future.

          • nchmy 1 day ago

            "near future" is not "today"

            • alecsm 1 day ago

              It can be because the message has been there for some weeks now.

        • vdfs 1 day ago

          It's a big confusion, some[0] say an email was sent about significant price increase, personal I haven't seen anything official

          [0] https://finance.yahoo.com/technology/ai/articles/deepseek-pl...

          • surgical_fire 1 day ago

            The email is real, I received it from DeepSeek itself. I probably received it because I buy tokens directly from them.

            No actual price increase however.

          • lionkor 14 hours ago

            Also got the email. It warned of a future large price increase, and to carefully watch usage.

            I read it as a "hey we will make stuff more expensive, don't miss it"

        • GrinningFool 1 day ago

          The banner on account settings; and a blurb on the pricing page: "We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice."

    • anigbrowl 4 hours ago

      It has been updated now, and will take effect next week:

      https://api-docs.deepseek.com/quick_start/pricing/ Briefly Pro is $2/1m output in off-peak periods, $4 in peak. Flash is $0.66/$1.32. Input tokens are still much cheaper.

      I don't mind these prices but I find the need to check against two different time brackets of unequal length an annoying distraction. I guess I need to make some little background app or plugin.

  • Eueudhsbsj32 1 day ago

    What's the new pricing?

    The prices on OpenRouter still look the same.

    • notatoad 1 day ago

      nobody is saying. just "more".

      but openrouter says they don't expect the price to change other than through the deepseek api, other people hosting the same model will keep charging the same price.

      • Eueudhsbsj32 1 day ago

        Unfortunately cache reads with third party providers are all 10-50x more expensive than with DeepSeek, so they're not even close to as cost efficient for multi-round agent use.

      • monster_truck 22 hours ago

        It's just a flat 1.5x during peak hours, they emailed this to everyone 2 months ago.

        So still effectively limitless.

        • anon373839 17 hours ago

          Yeah, Dax from OpenCode said that it appears to just be traffic shaping, nothing to do with the inference economics. He also said that OC have already replicated the inference cost in internal experiments.

  • igravious 1 day ago

    yup :)

    i'm doing opencode <-> openrouter <-> official deepseek api (i don't get the opencode hate, i like it)

    how are you doing it?

    am also using Kimi K3 via kimi-code

    and also GLM 5.2 via ZCode

    happy with all three, they're trailing frontier but i figure if i'm running GNU/Linux then i ought to favour open weights models with my €s -- reduced my usage of claude/gpt to the ~$20 tier just to keep abreast of claude_code/codex developments

    • literallyroy 1 day ago

      > i don't get the opencode hate, i like it

      When the company I work for was evaluating it, there were multiple rough points. Their terms and conditions allowed training on prompts, the default behavior was to route prompts to their servers for conversation summary/labeling. One of their lead maintainers is also super toxic on many issues.

      Sorry this is all baseless with no links, I’m on my phone and locating those issues again isn’t something I have time for.

      It’s a good tool I just don’t like the privacy policies nor maintainers attitudes.

      • HDBaseT 23 hours ago

        1. The privacy policy was a bit misleading, but it has since been updated to reflect the exact state of things. [1]. For example, DeepSeek models have ZDR, although their ZDR contract is renewed monthly. It COULD change. You need to toggle a Setting in your account to use DS.

        2. At one point (apparently) summary and title generations were handled by Grok. This has changed, by default it uses your 'small_model' configured in your config. By default, it will use a cheap model provided by your provider. E.g. if you have ChatGPT API connected, it will use the cheapest ChatGPT model. OpenRouter users MAY see it routed to a free model however. [2] [3]

        [1] - https://opencode.ai/docs/go/#privacy [2] - https://github.com/anomalyco/opencode/blob/9b805e1cc4ba4a984... [3] - https://opencode.ai/docs/config/

  • eli 1 day ago

    The Deepseek official API is good with excellent caching.

    But their privacy policy is unusually bad - they can train off your prompts and completions.

    • trollbridge 1 day ago

      Use another provider from OpenRouter.

      I really don’t care if they train off my prompts.

      • stanac 1 day ago

        V4 Pro 0813 isn't offered by other providers. I can't find this model on hugging face. It's probably not open, or not open yet.

  • sschueller 1 day ago

    Deepseek seems to have gotten too cheap. I have been using it for a long time and it's at a point now where my credits balance barely moves even at max setting.

  • killingtime74 20 hours ago

    Just use opencode go, you get more bang for your buck. Same api

scrlk 1 day ago

Benchmarks:

    | Benchmark                | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2   | Kimi-K3   | Opus-4.8  | Fable 5       |
    |                          | 0813      | 0731        | Preview   | Preview     |           |           |           | (w/ fallback) |
    |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------|
    | HLE (wo/w tools)         | 42.7/60.0 | 37.8/51.5   | 37.7/48.2 | 34.8/45.1   | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0     |
    | Terminal Bench 2.1       | 87.9      | 82.7        | 72.1      | 61.8        | 81.0      | 88.3      | 85.0      | 88.0          |
    | NL2Repo                  | 61.5      | 54.2        | 38.5      | 39.4        | 48.9      | -         | 69.7      | -             |
    | Cybergym                 | 83.3      | 76.7        | 52.7      | 38.7        | -         | 80.0      | 78.3      | 83.1          |
    | DeepSWE                  | 62.7      | 54.4        | 12.8      | 7.3         | 46.2      | 67.5      | 58.0      | 70.0          |
    | Toolathlon-Verified      | 74.1      | 70.3        | 55.9      | 49.7        | 59.9      | 76.5      | 76.2      | 77.9          |
    | Agents' Last Exam        | 25.7      | 25.2        | 16.5      | 15.8        | 23.8      | 27.6      | 25.7      | -             |
    | AutomationBench (Public) | 31.8      | 25.1        | 12.8      | 10.8        | 12.9      | 30.8      | 27.2      | 29.1          |
    | DSBench-FullStack        | 71.1      | 68.7        | 41.8      | 37.0        | 61.8      | 73.7      | 71.6      | 77.2          |
    | DSBench-Hard             | 67.2      | 59.6        | 31.1      | 25.8        | 54.5      | 63.0      | 71.7      | 68.3          |

Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4...

  • parsimo2010 1 day ago

    The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence...

    For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max.

    - 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse.

    - 86.6 on Terminal Bench 2.1. Pro 0813 is better.

    - 55.9 on NL2Repo. Pro 0813 is better.

    - 27 on Agent's Last Exam. Pro 0813 is a little worse.

    - 72.5 on Toolathon-Verified. Pro 0813 is better.

    - 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better.

    - 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better.

    I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that.

    • trollbridge 1 day ago

      By that standard, the release of Grok 4.6 was also timed on the same day.

      Given how I think DeepSeek operates... I think they just release it when they feel it's ready, and don't even seem that concerned with what other people are doing.

      • somenameforme 1 day ago

        Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just lets them shrug and cancel the fund raising round after the leaks came from said funding round.

        • surgical_fire 1 day ago

          Their stance on LLM development is why they earned my respect in a time when OpenAI and Anthropic only earn my mistrust.

          That, and the fact that DS is an insanely capable model.

        • scrlk 1 day ago

          Benefits of having a well performing hedge fund funding DeepSeek.

          IIRC, Demis attempted to start a fund inside DeepMind but it was killed off. In an alternative world where he manages to pull that off, perhaps DeepMind would still be independent with Demis at the helm.

        • trollbridge 1 day ago

          The founder of DS's stated goal is to get to AGI. He thinks this is the path to get there.

          Kind of interesting, when compared to the hubris from American frontier labs.

          • johnvanommen 1 day ago

            > Kind of interesting, when compared to the hubris from American frontier labs.

            One Man’s “hubris” is another man’s “marketing campaign.”

            Drama sells.

          • ngl999 20 hours ago

            When one needs money, an infinitely remote goal is the best cause.

      • parsimo2010 1 day ago

        Actually, yes. I just didn't know about Grok's release because they aren't on the front page of HN.

    • eli 1 day ago

      Official pricing only kinda matters for an open weight model, no?

      • parsimo2010 1 day ago

        It still matters as a point of comparison until other providers come online. If the consensus price from other providers is much different that can be compared then. But for now we have $0.435 / $0.87 for v4 Pro 0813 (with increase announced but we don't know the new pricing), and $2 / $6 for Qwen3.8-max. So until we get other data points that is what we have to look at.

        • eli 1 day ago

          I wondered if the promised change in pricing is actually going to be deepseek bringing up their cached costs. They're extremely inexpensive.

    • maherbeg 1 day ago

      I mean at the rate of model releases happening, I think a lot of these will collide more often than expected!

  • bel8 1 day ago

    So it's a Fable class LLM?

                                 DSV4Pro vs Fable5
        HLE w tools              60.0 vs 63.0
        Terminal Bench 2.1       87.9 vs 88.0
        Cybergym                 83.3 vs 83.1
        DeepSWE                  62.7 vs 70.0
        Toolathlon-Verified      74.1 vs 77.9
        AutomationBench (Public) 31.8 vs 29.1
        DSBench-FullStack        71.1 vs 77.2
        DSBench-Hard             67.2 vs 68.3
    • eli 1 day ago

      Fable's guardrails would never let it do something like Cybergym so at least for that one it's measuring Opus 5

      • wren6991 1 day ago

        We have a first-party figure from the system card [1]:

        > Mythos 5 reproduced 83.8% of targeted vulnerabilities on a single try, and produced at least one crash in 99.4% of tasks. This is comparable to Claude Mythos Preview, which reproduced 83.1% of targeted vulnerabilities and produced a crash in 97.1% of tasks. By contrast, Claude Opus 4.8 achieved a score of 78.1% (95.7% any crash).

        So their quoted figure exactly matches the figure for Mythos Preview, although they don't state the provenance. It could also quite possibly be an independent measurement of Opus 5.

        [1]: https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e90826608...

    • nikcub 1 day ago

      that DeepSWE result is likely most indicative of how you'll find real world usage

  • goldenarm 1 day ago

    Geometric mean of all these benchmarks :

    * GPT-5.6 Sol: 65.5

    * Fable 5 (w/ fallback): 64.5

    * Opus 5: 64.0

    * DS-V4-Pro 0813: 62.5

    * Kimi-K3: 62.3

    * DS-V4-Flash 0731: 55.8

    * GLM-5.2: 47.3

    • svachalek 1 day ago

      Maybe it's me but I don't see how DS Flash is better than GLM at all, much less by a huge gap. I'd probably protest less against Fable and Opus being put at the same level than many would, but there's no denying the two models are a very different experience from each other. I guess where I'm going is no one should pick a model by the benchmarks.

      • platinumrad 1 day ago

        I think instruction following carries outsized weight in these evaluations.

      • spijdar 1 day ago

        I'm not the most LLM-savvy person around, and I'm not gonna say I've put a ton of effort into practically compared these open models. But, a month or two ago I did do some "practical evaluates" testing GLM 5.2 versus DSv4 (flash/pro) with OpenCode's subscription with some late 80s Unix clone-type work, and this jives with my experience.

        GLM ended up being far slower, and far more expensive, for approximately the same results. There was never a problem that GLM could solve that DS couldn't solve, faster, and significantly cheaper.

        I strongly agree that you shouldn't pick a model based on benchmarks. But for me, I found GLM really underwhelming given its cost and speed.

        DSv4 isn't as good as GPT or Claude or what have you, but it's fast, and pretty darned effective. I can run a 3-bit quant of DSv4 locally on my system with ~15 tokens per second, and for a local model it might be the most overall effective at coding. For what it is, it's extremely impressive.

        • ApolloFortyNine 1 day ago

          My experience is the same.

          Imo it has a lot to do with you/the harness tries to get it to test itself. Deepseek v4 flash seems more than capable of understanding when something has failed, and making changes until it works. I've definitely seen it make mistakes I would expect something like Opus to find, but it works through them on it's own (and for literal pennies).

          At the end of the day, I think that's one of the most important features of a model.

        • wut42 18 hours ago

          Exactly the same experience. I really loved GLM5.2 for a while, but after trying it again after riding DSv4 flash (new) for a while, they're mostly at the same capabilities, with GLM being slower and much, much more expensive. A task cost me 2$ where it did very wrong, whereas Flash nailed it almost instantly for like a rounding error on my billing page.

        • kaeluka 9 hours ago

          Interesting!

          GLM 5.2 is slower for sure (although they offer a fast version), and it's more expensive. But in my experience, it's universally better than Deepseek V4-flash-0731. Don't get me wrong, the new Flash version is amazing.

          But the use cases I have looked at are about source code understanding, bug finding, etc. - GLM 5.2 is clearly better.

          I think by using some prompt engineering, you will probably be able to close this gap, but some extra work is needed.

          And I'll say it again: the new Flash version is amazing. I love it. That level of intelligence for the price is unprecedented, and the fact that it's open weights and runs locally makes me genuinely happy.

      • spiffytech 1 day ago

        In my little social circle DS4F generally substitutes for GLM 5.2 except it's the next best thing to free.

      • segmondy 1 day ago

        It isn't. I run both at home. GLM5.2 Q4 crushes DSv4Flash0731 Q8. I reach for DS for speed and for medium effort level work. If I care about quality I'll reach for GLM5.2 Looking at this release, I'm comparing it to GLM5.2 and it seems to beat GLM5.2, only time/experience will show. If true, then I'm happy. It's much easier to run than Qwen3.8/KimiK3

  • NietTim 1 day ago

    In classic reddit fashion the post you linked to is now deleted

    • SV_BubbleTime 1 day ago

      To be fair… I don’t know who still needs to figure out that AI benchmarks are almost all entirely fucking trash, but the great number would surely surprise me.

  • myworkaccount2 1 day ago

    IMO the HLE scores without tools seem to align better with real world performance of the models.

    To me it feels like the difference between "RL performance" and the pretraining / base "knowledge".

    Yes you can RL terminal bench to the moon but does the model hold up on out of distribution tasks?

    Kind of like trying to navigate a dark room with a laser light, vs a flashlight. Laser is going to go a lot farther much more efficiently but only if you are already pointing it at the right place.

  • andai 1 day ago

    The most interesting part of this is how Flash scores almost as well on all of them.

    Haven't tried the new DeepSeek models but I'm assuming the difference is more than these numbers show!

swingboy 12 hours ago

I’ve found Flash 0731 to be pretty great recently. I do feel like a lot of the models are pretty close in terms of ability. I often run code through multiple different models _and_ harnesses for code reviews and they all typically find the same things as each other.

eshack94 1 day ago

It appears that the only available endpoint (as of this writing) requires enabling "Allow paid endpoints that train on request data" in the OpenRouter privacy settings. I hope additional paid providers will become available that don't require training on data.

  • cdolan 1 day ago

    That is likely because Deepseek themselves is the only host.

    In 24-48 hours there will be other options I presume

  • jubilanti 1 day ago

    Their privacy policy doesn't forbid them from just straight up publishing your raw prompts as training data.

    My threat model is that anything I POST to DeepSeek I treat as public to the web, as much as a public GitHub repo is.

XCSme 1 day ago

Again, I will wait until there's a provider that doesn't train on prompts before I will benchmark.

  • dakolli 1 day ago

    psst.. they all do. Also, what kind of IP are you protecting, are you protecting some crazy discovery, nothing you're throwing at them is special, they aren't going to steal your CRUD pomodora app.

    If anything Deepseek is the only company I'd want to consent to training on my data, they're by far the most altruistic. Atleast they give back all their IP in the form of research and open source weights. It's not like they're hoarding your data for them to make money, they're basically giving everything out for free. The only reason you even have the option of waiting for another provider is because they release weights.

    They're releasing all their IP, which is a trillion times more valuable than anything you're providing, you people are just greedy and oddly self centered.

    • XCSme 1 day ago

      I "trust" what they say on OpenRouter for the provider, for some it says they retain prompts, for other that they retain but can also use them for training.

      It's not any crazy IP, just my own benchmarks/tests, once they are in the training set it defeats the purpose of the tests, and I have to make new ones.

    • diydsp 1 day ago

      >are you protecting some crazy discovery

      Yes. If someone figured out my current project they would have a huge scoop.

  • LeBit 1 day ago

    The good thing is that there seems to be quite a lot.

    Let’s just wait a bit for this one.

  • ljlolel 23 hours ago

    I built TrustedRouter so this can fail closed. min_privacy=zdr rejects the request when the model has no ZDR provider; confidential requires provider-side confidential compute. https://trustedrouter.com/blog/how-confidential-computing-pr...

    • XCSme 22 hours ago

      I don't understand how this works?

      Is it another proxy on top? What stops the provider from reading/storing the prompts at the LLM execution level?

      • ljlolel 19 hours ago

        it's confidential compute, it's open source and you can verify yourself that it's not reading the prompts

        • XCSme 18 hours ago

          Open source doesn't matter if someone else is running it, right? They can change it?

          As long as the prompt is not encrypted at some point, and I don't think LLMs can run on encrypted prompts, then it can be read.

          • ljlolel 11 hours ago

            With confidential compute / TEEs you can guarantee that the code is running, it's verifiable with remote attestation

            • XCSme 11 hours ago

              So where does the guarantee stop? At the GPU driver level? Firmware level?

              What if the GPU has a custom bios flash that somehow logs the unencrypted prompts?

            • inigyou 11 hours ago

              SGX has been cracked

schmorptron 15 hours ago

The past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder where that's from? Maybe they're overindexing on coding even more than others? I can't say I've noticed it in my (coding) usage so far, has anyone seen it make up potential root causes or other speculative stuff more than other models?

  • gunalx 15 hours ago

    Yeah. I stopped ising deepseek v4 flashbecause it is awful (even worse than my local qwen3.6 35B model) at multilingual prose.

    • isqueiros 10 hours ago

      Working on a language related app makes me realize that all these supposed language models don't have many good language benchmarks

cjg007 1 day ago

Before DeepSeek-V4-Pro-0813's price goes up, I expect a surge of frantic traffic — hope the servers can hold up.

  • cjg007 9 hours ago

    DeepSeek raised their prices. Oh my god.

aabdi 1 day ago

https://api-docs.deepseek.com/quick_start/pricing/

Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.

  • swiftcoder 1 day ago

    How does it stack against the updated Deepseek Flash version?

    • k__ 1 day ago

      Around 5 percentage points better. (E.g., 87% instead of 82%)

      • Gecko4072 1 day ago

        So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.

        • k__ 1 day ago

          I tried the previous Pro model and in the end it was 50% more expensive than the previous Flash.

          Wasn't worth it.

        • saaga 1 day ago

          Yea that's what I was thinking. Flash is nuts. I find I have to be a more precise and specific with it but damn. It's crossed a threshold of production grade coding for sure.

          I was running a session over a couple days and it didnt cross a dollar lol.

        • networked 1 day ago

          I haven't tried DeepSeek V4 Pro 0813 yet. Recent experience tells me that larger models are worth it in non-obvious ways. MiMo-V2.5-Pro solved problems that DeepSeek V4 Flash 0731 couldn't solve for me: for example, adding a live counter for elided reasoning lines to a terminal-based coding harness. You wouldn't be able to tell from the scores on their respective Artifical Analysis page (https://artificialanalysis.ai/models/mimo-v2-5-pro, https://artificialanalysis.ai/models/deepseek-v4-flash). I like the DeepSeek V4 models, though. They critiqued my engineering decisions better than MiMo, and they seem to have a distinct aesthetic in the SVGs they write.

          • trollbridge 1 day ago

            Interesting - I've been dropping into MiMo-V2.5-Pro-UltraSpeed whenever Flash seems to be "stuck" and it usually figures it out. I use UltraSpeed just because I'm so frustrated by then that I'm impatient.

            I still find 5.6-Sol can solve some things neither of those can, but it's so slow (and it's so hard to trace / debug the reasoning) that I just let it run overnight.

            • networked 1 day ago

              What about 5.6 Terra and especially Luna? Luna scores pretty high on benchmarks and seems to have different habits (like a denser pattern of tool use) and blind spots.

              I'm trying out a development workflow where I generate mundane code with MiMo and Luna (and soon V4 Pro 0813?) and have Opus 5, which is running on only a Pro subscription, review and refactor it. I'm not sure it will justify the context switching, but it's an interesting exercise.

              • trollbridge 1 day ago

                Terra and Luna are fine, but they’re quite slow (OAI seems to be really slow lately) and don’t have the reasoning traces. My workflow really depends on them or I can’t switch models effectively.

        • npn 1 day ago

          I still believe this is not the full potential of pro models. I expect they will release another checkpoint later this year.

        • eli 1 day ago

          Opus 5 medium to Opus 5 max is only 3 points, if that puts it in context

      • sparkling 1 day ago

        deepseek-v4-flash feels so fast and snappy, i'm loving it. Happy to trade speed for the the 5% degraded benchmarking performance.

        • k__ 1 day ago

          I wouldn't exactly call it snappy, but faster than Pro, yes.

          • ericd 1 day ago

            Single request depth on vllm with dspark, I'm getting ~200 tps, I'd say it's pretty snappy.

            • JacobAsmuth 1 day ago

              Well sure but you're running on tens of thousands of dollars of hardware.

              • ericd 1 day ago

                It's much faster than other models on that same hardware in the same size class. I've tested a few, it's by far the fastest I've tested.

                And it wasn't tens* until recently. Didn't expect this to be one of my best performing assets this year.

            • k__ 1 day ago

              I get like 80.

              • ericd 19 hours ago

                What's your setup? Happy to try to point you in the direction that worked for me.

        • saaga 1 day ago

          I feel the same too. I like the speed. I'm also a big fan of glm 5.2 fast. I can't wait for like 2000 t/s on these haha.

    • pixelesque 1 day ago

      I've found Pro to be a lot better per "task" than the recently released Flash for code reviews and things (via OpenRouter running in pi.dev).

      Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.

      • swiftcoder 1 day ago

        yeah, I've definitely noticed one has to be quite precise to keep Flash on the straight-and-narrow

        • RALaBarge 1 day ago

          Every plan and every code checkpoint finds me saying "Check with Grok and Fable latest to critique our strategy/code review" with pretty much every model. I havent ran into any deal breakers with the new Flash version yet (like it not running a tool properly or coming back with something completely daft)

      • surgical_fire 1 day ago

        I use a plan -> implement wotkflow for this reason.

        pro plans, flash implements. I am super happy with how flash behaves like that.

  • JacobAsmuth 1 day ago

    Per token. You need to look at pricing per task.

    • trollbridge 1 day ago

      ... which still comes out cheaper, since DeepSeek caches so much more.

      I keep track of my token consumption even on subscription plans and my equiv. cost for my 5.6-Sol usage is around $4000-$8000 a month.

      • dgellow 1 day ago

        How much do you pay for the subscription?

        • RALaBarge 1 day ago

          Not them, but I payed 10 dollars to DeepSeek directly to use their Reasonix tool. I worked all weekend and the past few days, billions of tokens, I still have 3 bucks left!

  • xynelius 1 day ago

    If that wasn't impressive enough, it's actually ~60x cheaper if you take into account the typical cache-read/input/output split in agentic coding, and the deep discount for cache reads offered by DeepSeek. Opencode has some public data on the typical split [1]:

    For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached.

    Cost per request for V4 Pro: $0.000875 per request.

    Equivalent Opus cost (w/o taking into account cache write costs): $0.052 per request.

    [1] https://opencode.ai/docs/go/#usage-limits

    • taosx 1 day ago

      I created a simulation for coding harnesses based on my own pi sessions. When taking into account all factors, DS-v4-Pro is cheaper than gpt-5.6-luna due to caching. Look at the bill segments difference for cache read cost and uncached cost between deepseek and the other models. At this point is cheaper to use ds-v4-pro than the luna models from openai.

      ignore the numbers except the classic and keep in mind that classic is based on pi with the only change limiting tool output to 10kb

      https://harness.eveid.com/lazy-harness-cost-simulation

      * I built this for getting an initial estimate between different checkpoint/ compaction methods for the harness.

      • RALaBarge 1 day ago

        Hey this looks good! Maybe consider adding a hover-over popup for the rectangles explaining what each thing means to a lay person. I see it at the bottom, but that is below the fold.

        • taosx 1 day ago

          Done, I'll take any other suggestions and apply them later, I will also split it a bit for different usecases as this was initially a throwaway prototype but found it useful. Basically it needs a bit more human touch.

    • HDBaseT 23 hours ago

      Can we have a conversation about subscription plans for a minute?

      I don't mean to hype up the US AI firms, but if a ChatGPT $200/m subscription can get you $16,000 in effective API costs, doesn't effectively every model get destroyed by the subsidized Claude/ChatGPT models? Both in price and intelligence.

      • polski-g 21 hours ago

        Yeah pretty much. I spent half a billion in tokens one night on a huge refactor with DSFlash, cost $11.

        If I spent that every night it would be 3x my GPT subscription.

        • nchmy 10 hours ago

          That seems too expensive to be honest. Did you do it with official deepseek api or a 3rd party provider? Because official has 10x cheaper cache reads than the rest. I've done similar sized chats for like $1

      • xbmcuser 17 hours ago

        That is the problem currently the subscription plans are being subsidized by VC money and token buyers. When Open weight get good enough token buyers build their own servers instead of buying tokens then no one to subsidize the subscriptions

        • adventured 13 hours ago

          99.9%+ of the tech worker population will never be able to build their own servers to run future frontier models. Kimi 3 is an indication of what's coming. These models will keep getting drastically larger. The hardware isn't getting cheaper anytime soon (no matter what China does; that goes for memory and GPUs).

          Cycle forward to Fable 7, Kimi 5, GPT 7 a couple years out. Forget about it unless you own a datacenter.

          • zozbot234 13 hours ago

            > 99.9%+ of the tech worker population will never be able to build their own servers to run future frontier models.

            A single local user can run frontier models slowly on a 24/7 basis, which drops hardware requirements by orders of magnitude compared to a datacenter setup for just-in-time inference. This is not a real alternative to subsidized subscriptions at present, but it's a great insurance policy against future VC-driven rug pulls.

          • xbmcuser 6 hours ago

            90% of the tech worker population does not work for themselves they work for someone that pays them 1000s in salary for those paying those tech worker spending $40-50k on a server that helps them not pay for 2-3 tech workers is not that big a deal.

  • segmondy 1 day ago

    ... and mere mortals can run this at home or rent a GPU, you can't do so with Sol or Fable.

Myzura 1 day ago

This model is not very good at coding, but it is quite good at research, evaluation and action, I don't write code, but it really goes head-to-head with the most expensive models in searches such as stock market and forex

hemkeshr 10 hours ago

I had high expectations for V4 Pro, especially since DeepSeek V4 Flash 0731 performed so well compared with other Flash models. What a letdown.

  • erichocean 10 hours ago

    Bizarre, it has a nearly identical improvement as the flash model.

  • aqme28 10 hours ago

    Wait, what are you disappointed by? Seems like a significant jump in performance, and it competes handily with other models.

arj 11 hours ago

Been testing this on hobby project https://github.com/arj03/seedkernel/. Latest flash was a big step up. Pro feels really slow compared. Claude opus is still better day this level.

  • arj 9 hours ago

    That said. It is really good at security review at max settings. Just ran one, it came to 0.15$

minraws 14 hours ago

Deepseek V4 Pro 0813 is the most unreliable model I have tried, it works on pass@3 shockingly well you can get it to match Sol or Fable perhaps in task done, but it's horrendous at pass@1 very prone to going wrong and doing horribly at most benches.

I am not sure what it is buy I suspect it might be GRPO.

  • seunosewa 10 hours ago

    Could you try setting the temperature very low e.g. 0.0?

nthypes 1 day ago

Still behind Kimi-K3 in almost half of the benchmarks

  • segmondy 1 day ago

    Much easier & cheaper to run than Kimi

code51 9 hours ago

OpenRouter is a place with zero support.

Suddenly get a big debt on your account with nobody to respond.

As an early adopter of OpenRouter, I'm afraid they are in shambles.

  • htrp 4 hours ago

    Isn't that the model for all of the labs though?

jannishan 8 hours ago

Why do I feel that the Pro version's effects are inferior to Flash's? Is it just my imagination?

big-chungus4 15 hours ago

So flash is 52 points on artificial analysis, and pro is 53

  • WASDx 15 hours ago

    This was a disappointed to me. Why would I use pro over flash now? Is there some area where the difference is significant?

    • lionkor 14 hours ago

      The idea is, I believe, that the Pro model is a larger model (more parameters, or less quantization) in general. What implication that has, I couldn't tell you.

      For tasks like pondering on something, reviewing code, etc. I use Pro, just because it feels like the right model for that.

  • WASDx 4 hours ago

    On DeepSWE it's now 53% vs 63% which is one of the coding benchmarks I trust the most. DS own measurements also show a more significant increase so I suspect AA might update when they release an article.

    Surprisingly DeepSWE currently shows a lower total cost for pro so that might also update I guess. As usual, don't trust the benchmarks and try for yourself.

ernsheong 1 day ago

These people can't version control properly, V4.1 or V5 would be more appropriate.

nullbyte 1 day ago

Even though cost-per-token is low, Deepseek v4 tends to burn an immense number of tokens to accomplish tasks.

  • SwellJoe 1 day ago

    It still ends up being one or two orders of magnitude cheaper per task on benchmarks.

  • dools 14 hours ago

    I can go for days on end without topping up my deepseek account. When it can’t solve a problem I switch to GPT and have to top up in real time.

Perenti 22 hours ago

Graphs without labels and/or scales on the axes are useless. I know less after viewing that page than before, but I got to see some pretty lines that I guess must mean something.

  • CGamesPlay 21 hours ago

    The only graphs that don't have axes require a mouse to use. They show the values on hover—or if you tap the fullscreen button, that version also has axes. One graph is a 3-month time series showing a single day, so it looks like it doesn't have an X axis, but it does.

Readerium 1 day ago

V4 Pro has vision correct?

  • trollbridge 1 day ago

    No.

    • coredog64 1 day ago

      Saw somewhere that they don't believe vision advances the AGI work they're doing, so it's not on the roadmap.

nimsarajay 1 day ago

I'm Satisfied with this model (in opencode)

magekinnarus 15 hours ago

DS Pro is what I hoped it would be. I have used it as an auditor for a couple of implementations, and the work is solid. This will allow me to split the work between 5.6 Sol and DS Pro.

m00dy 8 hours ago

DeepSeek is moving to peak/off-peak API pricing. Off-peak rates are 50% of peak rates.

Peak: 01:00–04:00 UTC and 06:00–10:00 UTC Off-peak: all other hours

New pricing takes effect August 16, 2026 at 16:00 UTC.

Model Period Cache hit Cache miss Output (input / 1M) (input / 1M) (/ 1M)

deepseek-v4-flash Off-peak $0.007 $0.22 $0.66

deepseek-v4-flash Peak $0.014 $0.44 $1.32

deepseek-v4-pro Off-peak $0.022 $0.66 $1.98

deepseek-v4-pro Peak $0.044 $1.32 $3.96

For batchable workloads, scheduling outside those two UTC windows cuts token costs in half.

halyconWays 19 hours ago

You mean to tell me it's been 10 hours and there's no unsloth quant? I've been running 0731 and love it.

xbmcuser 17 hours ago

Did you guys read the fine print they plan to increase prices significantly in the future

  • mrbnprck 10 hours ago

    Take a look at openrouter there are a few alternative deepseek model providers that follow an even cheaper pricing strategy.

    Funnily even if deepseek themselves increase price 2-3x they are still more affordable.

gigatexal 22 hours ago

Welp gonna give Deepseek more money. This is very cheap indeed. And I’ve been using them and kimi for a bit now not via open router but on my own and have found them on part with sonnet 5 though sonnet 5 these days I think has gotten worse.

At work I had to move to Fable to get decent work results.

moritzwarhier 1 day ago

Is having padded version numbers with a leading zero a common thing?

Wondering, sorry if it's a dumb triviality to ask.

Is this even a (sub-)version number? I mean the major version is clearly 4.

  • gorfxx 15 hours ago

    its for the month and day the model released in 2026. 0813 -> Aug 13th. I assume that if they use it internally the padded 0 makes finding the newest model easier cause all the numbers for the date line up instead of the zig zag you get without it once you get to 2 digit months.

nicman23 14 hours ago

mannnn i just downloaded 0731 ffs

Palmik 1 day ago

Why does this link to OpenRouter, which has no useful information on its own? Linking to the official API or the benchmarks would make more sense:

- https://api-docs.deepseek.com/

- https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)

  • echelon 22 hours ago

    Moreover, OpenRouter is NOT Open Source, fair source, source available, etc. It's a proprietary cloud service that got first place in the API aggregation distribution game.

    Link to DeepSeek!

    • zamadatix 20 hours ago

      Open doesn't always refer to the code. Just like their previous project, it refers to an open marketplace where anybody can sign up to sell access to models.

      But it'd still be nice to post to wait an extra minute to find some other page/new url from deepseek for it instead of posting that it exists somewhere.

      • ljlolel 19 hours ago

        it's not open, I know people rejected by them (then they went to my site to be listed Trustedrouter.com)

        • zamadatix 15 hours ago

          You accept any and every registration? I think that goes a bit beyond an open market and into a completely unregulated one I'd never want to route my sessions through. Thanks for providing the conflicting interest disclosure though, too often people don't bother.

    • ljlolel 19 hours ago

      TrustedRouter is hosted and full opensource!

      • scrollop 16 hours ago

        Do you know if it's more expensive than open router?

        • ljlolel 11 hours ago

          5% instead of 5.5%

    • shostack 16 hours ago

      It is great for getting aggregate insights about the privacy and security picture of model providers.

      For example if you want zero data retention and US -based hosting you can find that easily. You will not find that through Deepseek.

  • simonw 21 hours ago

    DeepSeek really need to provide a PAGE for this model release. There's no blog post, there's not even a tweet. It's very unclear what we can link to!

  • zxilly 21 hours ago

    There's no new page for this model. Hackernews didn't allow the same link be posted twice.

    • ronbenton 20 hours ago

      Bogus query params could work maybe

      • manbun 18 hours ago

        i disagree

      • tomhow 15 hours ago

        Please don't do that. Email us instead (hn@ycombinator.com) and we'll review.

    • beltsazar 18 hours ago

      Is this a new rule? I've seen some popular blog posts reposted here several times and still highly upvoted.

      • squirrellous 16 hours ago

        I _think_ that has always been based on discretion of the mods. It’s only allowed after some period of time and if popular enough.

        • tomhow 15 hours ago

          The software detects when an identical URL is submitted and will mark it as a duplicate (with certain conditions like how recently it was last submitted and whether it had any front page time and discussion).

          When similar URLs or different versions of the same story are submitted and get votes/comments, we have to use discretion to work out which URL is best, who submitted first, which discussion is most active/healthy, and we'll try to consolidate the discussion into one thread with the most informative/canonical URL as the main link.

  • alexwwang 18 hours ago

    Don’t you find the official document website was out of service for a long time since the new model was published soon?

  • sinuhe69 17 hours ago

    I don’t know about you but I find the information about prices, effective price (weighted average), providers and performance, benchmarks (down bottom) very useful. With openrouter I can even test it right away and compare with other models (use the chat functions).

    • Palmik 14 hours ago

      As of now, there is only a single provider for this model and that's the official DeepSeek API.

      When GPT 6 comes out, would you expect the top thread to link to OpenRouter?

LeonKnst 1 day ago

I find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard

  • HawtAds 1 day ago

    Hacker News is very Bay Area/US tech centric where spending a few hundred a month on AI is just pocket change. The weaker AI models with more questionable data retention policies are popular in developing countries. I think the new Facebook muse model will be similarly popular.

  • BlackRabbit1 1 day ago

    A lot of it/infrastructure departments aren't aware that you can use Asian models hosted within the US or even EU.

  • spacebanana7 1 day ago

    In an enterprise setting Chinese models are often discouraged due to political risk. They don't want to need to remove a model that's deeply embedded in their stack. And it's entirely feasible that the US gov bans federal contractors from using them in the next 6 months for example, or that EU AI safety rules effectively ban them too.

    • BlackRabbit1 1 day ago

      There are EU/US providers offering Deepseek/Qwen/Kimi/etc.-as-a-Service. With zero ties of their infrastructure to China.

      Fully compatible with the well known Antrophic API.

      You only have to replace the URL and your key.

      • odo1242 1 day ago

        Based on what the political climate looks like nowadays it's entirely possible the US bans federal contractors from associating with any company that uses the models themselves, regardless of data provenance or where they are hosted. Or they create AI safety rules that make it impossible to release open source models (for example, making it so that closed-source models can be evaluated with a harness but open-source models need to pass the benchmark with the weights alone, which isn't really possible). Or they just declare Chinese models a security risk like TikTok (claiming that the model would be trained to respect Chinese interests).

        It may not be likely but it's definitely possible enough to be something people worry about.

    • trollbridge 1 day ago

      Then run the DeepSeek or Qwen model on AWS GovCloud, etc., and you won't have any risk of exposure to "China".

      I'm not even sure what "EU AI safety rules" are. Can't people in the EU just use whatever they want?

      • hgoel 1 day ago

        Running on AWS GovCloud isn't necessarily an option, some places prohibit running Chinese origin models even locally.

  • sinuhe69 1 day ago

    Well, one reason is that we always have to work with the quirks of each model. So, a know model is often preferred over a new/unknown one because we have to be vigilant again. (Negative) surprises are mentally exhausting in the long run. IMO, you can work much better when you know the model.

  • ianm218 1 day ago

    I suspect if you follow dev groups in developing countries people are much more focused on token/ price efficiency.

    For funded startups it mostly just doesn’t matter a ton unless you are passing on inference in your product at scale

  • krlx 1 day ago

    Well things may change soon. I've been testing Coding fulltime with Deepseek Flash this week to evaluate an eventual shift for the whole company away from anthropic. It has been quite positive and I can't wait to try pro tomorrow. If our data has to be used by either US or China, we might as well go the cheaper and unwalled garden. If only it supported image input ...

  • spacephysics 1 day ago

    Most of my model usage comes from my work’s model selection (which is now down to just Claude models)

    I’ll try out the latest models, but mainly stick with Claude only because I’m most used to its quirks and how to work around them. I imagine this is part of these hyperscalers playbook.

    I will say though, I miss Sol model at work. It with Codex was amazing at first-shot understanding. Claude i need to scope out where to look otherwise a large portion of my token budget is eaten up

  • cortesoft 1 day ago

    I keep using Claude and Codex simply because the subscription rates are SO MUCH cheaper than per-token rates, even with the cheaper models

  • numpad0 1 day ago

    There's just no place for models that are neither SoTA nor truly crazy cheap in today's public mental health climate.

    If it's 500x cheaper than US models for similar ballpark performance just because it's hosted in China, sure whatever. If it's name brand like Anthropic/OpenAI/Google, that's kinda fine too.

    If it's neither, like merely 50% cheaper than latest OpenAI whatever, however massive loss that pricing may be incurring to its provider, it wpuld be considered not worth any attention.

yipinwong 1 day ago

Worse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)

  • Eueudhsbsj32 1 day ago

    Unless you're Chinese, why would you care if they see your data?

    As an American, I'd much rather have my data kept outside the country than here where companies and the government have a lot more leverage over me.

    • segmondy 1 day ago

      This! It's always amusing when folks say "But China", my data in the hands of my government and their billionaire friends is more than dangerous than in China. I mean, if it's an IP sort of thing then go local.

    • akman 1 day ago

      I do think this question comes up a lot-- I can understand why.

      For some well-explained reasons, check out https://darioamodei.com/essay/the-adolescence-of-technology and search for "CCP".

      • Eueudhsbsj32 1 day ago

        So is your concern more about reducing the risk of an authoritarian China "winning" the AI race? And less about reducing the risk of your data being used against you personally?

        To me, the risks of an individual helping China to continue to develop their AI by being a customer is pretty marginal compared with the personal risks of my data being used against me.

        • akman 23 hours ago

          I see. Though if you follow the argument set forth by Dario, it seems you'll not only have your concern to worry about (i.e., personal risks of your data used against you), but many more as well on top of that.

      • vrganj 1 day ago

        As somebody from neither the US nor China, this argument would be much stronger if the US hadn't started acting like a rogue state - starting wars of aggression and messing up the world's energy supply, actively speeding up climate change, kidnapping leaders of sovereign nations, threatening its allies (!) with invasion, etc etc.

        The CCP's not great either, sure. But the Americans don't really have a leg to stand on anymore.

        • akman 23 hours ago

          As far as AI is concerned, it looks like you will need to pick 1?

          • Cookingboy 22 hours ago

            Then I pick the country that hasn't been bombing people and starting wars nonstop over the past 40 years.

            Hint: It's not the U.S.

        • theyliesoeasily 23 hours ago

          To be fair the US has been doing this its entire history, it just stopped pretending. C.f. Hawaii, Guatemala, Cuba, Chile, "The Jakarta Method", etc. etc. etc.

      • HDBaseT 22 hours ago

        A blog by Dario of all people. Totally not bias towards non-US models.

    • yipinwong 7 hours ago

      Premise is wrong, "unless you're chinese".

      Don't matter whether you are chinese or not, anything going to China will be used against you.

      And also you are cherry-picking my comment on "China" as my emphasis was on Luna >>> DeepSeek.

      You just are not reading my intention.

  • iammrpayments 1 day ago

    It’s either chinese in the US or chinese in China anyway

  • comandillos 1 day ago

    At least you can run it for relatively cheap hardware. I guess OpenAI doesn't let you do that.