Agree, and both are on the front page. Also, Show HN is usually for your own project, but this appears to be a blog post about someone else's project (albeit "a good friend").
While it's not perhaps clear to me that this is a true Show HN, I enjoyed the writeup and it's good to see our German colleagues across the pond taking a good shot at this. I also appreciated the brief description on the Merlin-Arthur protocol - seems like a clever way to try to tackle the "I don't know" problem.
I'm always wondering why new models don't always adopt deepseek's KV tweaks. It's insanely valuable to have such powerful prefix caching and so cheap at inference time.
The thing to note here, besides the transparency and the fact that it’s actually a good model that also works well on coding and agentic tasks, is that it’s the first release by a team formed less than a year ago, with a strong focus on iteration velocity. There’s more to come.
disclaimer: I‘m part of the training team, happy to answer any questions
Please don't copy-paste comments - it makes merging threads a pain! (and indeed the temptation to copy-paste is indication that a discussion needs merging)
p.s. It's a great comment! no problem on that level
We've merged (most of) the comments into this thread, which is currently on the frontpage:
Kolibri: A Sovereign Open-Weight Model - https://news.ycombinator.com/item?id=49942706 - Oct 2026 (46 comments)
IMO the announcement is better https://aleph-alpha.com/en/blog/kolibri-has-landed-a-soverei...
Agree, and both are on the front page. Also, Show HN is usually for your own project, but this appears to be a blog post about someone else's project (albeit "a good friend").
It's here: https://news.ycombinator.com/item?id=49942706. I'm currently merging the threads.
While it's not perhaps clear to me that this is a true Show HN, I enjoyed the writeup and it's good to see our German colleagues across the pond taking a good shot at this. I also appreciated the brief description on the Merlin-Arthur protocol - seems like a clever way to try to tackle the "I don't know" problem.
I'm always wondering why new models don't always adopt deepseek's KV tweaks. It's insanely valuable to have such powerful prefix caching and so cheap at inference time.
Are there drawbacks to this?
A blog post is not a Show HN topic.
> the weights are on Hugging Face
> Show HN is for something you've made that other people can play with
The author did not make this model though, they used Claude to write a blog post about it.
This is not a Show HN
The thing to note here, besides the transparency and the fact that it’s actually a good model that also works well on coding and agentic tasks, is that it’s the first release by a team formed less than a year ago, with a strong focus on iteration velocity. There’s more to come.
disclaimer: I‘m part of the training team, happy to answer any questions
Please don't copy-paste comments - it makes merging threads a pain! (and indeed the temptation to copy-paste is indication that a discussion needs merging)
p.s. It's a great comment! no problem on that level