wingman-jr 1 day ago

While it's not perhaps clear to me that this is a true Show HN, I enjoyed the writeup and it's good to see our German colleagues across the pond taking a good shot at this. I also appreciated the brief description on the Merlin-Arthur protocol - seems like a clever way to try to tackle the "I don't know" problem.

Ey7NFZ3P0nzAe 13 hours ago

I'm always wondering why new models don't always adopt deepseek's KV tweaks. It's insanely valuable to have such powerful prefix caching and so cheap at inference time.

Are there drawbacks to this?

pu_pe 1 day ago

A blog post is not a Show HN topic.

  • embedding-shape 1 day ago

    > the weights are on Hugging Face

    > Show HN is for something you've made that other people can play with

    • pu_pe 1 day ago

      The author did not make this model though, they used Claude to write a blog post about it.

rolymath 1 day ago

This is not a Show HN

peterBlue75 1 day ago

The thing to note here, besides the transparency and the fact that it’s actually a good model that also works well on coding and agentic tasks, is that it’s the first release by a team formed less than a year ago, with a strong focus on iteration velocity. There’s more to come.

disclaimer: I‘m part of the training team, happy to answer any questions

  • dang 1 day ago

    Please don't copy-paste comments - it makes merging threads a pain! (and indeed the temptation to copy-paste is indication that a discussion needs merging)

    p.s. It's a great comment! no problem on that level