Step models were IMO the first local model you can run on 128GB shared memory that worked well. Really excited to see how it compares to Qwen Flash Next.
Edit: bummer, didn’t know it’s 600B-A27B. No way to run that on 228GB.
Well, according to Artificial Analysis (which I'll admit I've been using as a bit of a mental crutch to avoid comparing models myself, so YMMV), it's smarter and slightly cheaper than Gemini 3.8 Flash, which has been my benchline for "cheap and smart enough", I'll give it a try on OpenCode for the week but I'm not sure I'll be compelled enough to switch from Muse Spark 1.3.
Earlier this year, Step 3.5 Flash and then Step 3.7 Flash were easily my favourite local models. Step 5 is far too large to run on my DGX Spark-like, but I’m excited to try it out. It’s benchmarks look promising, and I’m hoping it’s a lot faster than GLM is
Seems like it was an under-noticed model back when it came out because there were so many new Qwen models coming out around the same time. I tried it some on my Strix Halo box, but it was right on the edge of fitting.
Indeed - Even with llama.cpp you can store n-grams to disk and use it. I m using unsloth's llama fork atm but a new llama release is out that may be better - hadn't had time to test it yet.
Back in their 3.5/3.7 era [1], they had a small+fast model, specifically eschewing knowledge, and instead focused on "general intelligence" like managing tool calls, orchestrating sub-agents, etc. Could be a very good substrate for running a personal assistant (time will tell what's the right approach); so I'm curious to try that out and see how they've come along since then.
I like trying new models but I wish we’d get something actually new. Like a new architecture or something. LLMs are just so sloppish. We can do better.
It's fast. The average speed is 115 tokens/sec according to OpenRouter. I haven't tested the model to see how it is in practice, but I'd certainly pay a little extra for faster inference.
Edit: Though the average latency of 1.5s isn't very low, so it might not be that fast in practice for agentic work. Also, I don't know how much thinking it does, as that's generally been the drawback to Chinese models.
Model | Reasoning | Intelligence Index | Artificial Analysis million output tokens for the intelligence index
-|-|-|-
Step 5 | ? | 44 | 160
GLM 5.3 | Max | 45 | 210
MiMo V2.6 Pro | ? | 46 | 140
Kimi K3 | Max | 44 | 160
Qwen Max 0902 | ? | 45 | 190
DeepSeek 4.1 Flash | Max | 39 | 250
GLM 5.3-flash | Max | 42 | 180
GPT-6 Astra | Low | 46 | 10
It looks reasonable by open model standards. This doesn't capture the fact that DeepSeek and Step 5 have much higher token/s than the rest, other than MiMo Ultraspeed. MiMo V2.6 is either slow but cheap, or fast but expensive.
Just based on these numbers, it looks good.
Astra-low is one of the fastest and cheapest because it doesn't use many tokens, but I've never tried it.
I liked DeepSeek and GLM when I used them.
Artificial Analysis has some very specific biases or perspectives on what they are measuring, so I wouldn't take their benchmarks as the final word on model quality.
Composite indices are only useful for model companies which want to build one horizontal capability for N use cases, or naive users who don't want to bother with the effort of carefully pairing models with use cases. If you're a power user looking to understand and make deliberate choices, then you want pointed evaluations -- not general composite indices. You wouldn't hire the same person to do your taxes and mow your lawn, so why is it any different with LLMs? Only if you come at the problem with the folklore around "AGI" do you start making composite benchmarks.
As I mentioned in a sibling comment... Back in their 3.5/3.7 era [1], Stepfun had a small+fast model, specifically eschewing knowledge, and instead focused on "general intelligence" like managing tool calls, orchestrating sub-agents, etc. Could be a very good substrate for running a personal assistant (time will tell what's the right approach); so I'm curious to try that out and see how they've come along since then.
I've always wondered if there were niches that some models are better at. I use them for software, electrical and mechanical engineering plus other things. What are some notable subject-specific differences you found?
Step models were IMO the first local model you can run on 128GB shared memory that worked well. Really excited to see how it compares to Qwen Flash Next.
Edit: bummer, didn’t know it’s 600B-A27B. No way to run that on 228GB.
Well, according to Artificial Analysis (which I'll admit I've been using as a bit of a mental crutch to avoid comparing models myself, so YMMV), it's smarter and slightly cheaper than Gemini 3.8 Flash, which has been my benchline for "cheap and smart enough", I'll give it a try on OpenCode for the week but I'm not sure I'll be compelled enough to switch from Muse Spark 1.3.
Earlier this year, Step 3.5 Flash and then Step 3.7 Flash were easily my favourite local models. Step 5 is far too large to run on my DGX Spark-like, but I’m excited to try it out. It’s benchmarks look promising, and I’m hoping it’s a lot faster than GLM is
Totally. I think Step 3.5 was the first model that made me sit up and realise that LLMs weren’t a fad
It should be mandatory to put benchmark results on the beginning of the page describing new model release.
How am I supposed to know if it's even worth my time?
Why is that interesting?
Stepfun made a close to SOTA model with 3.5. They got overtaken quickly but have shown enough to be given attention when a new model releases.
It’s also an open model. Even if you only use Opus and Sol, these open models help push them and the frontier.
Seems like it was an under-noticed model back when it came out because there were so many new Qwen models coming out around the same time. I tried it some on my Strix Halo box, but it was right on the edge of fitting.
3.5 Flash was decent, but I was using it on OpenRouter. What model are you running on your Strix Halo nowadays?
Qwen-3.8-flash-next @Q4
Not sure why this is downvoted.
It is fully possible to run the 180B Owen model on strix halo using strata and even get reasonable speed out of it.
https://github.com/Niko1221/Strata
Indeed - Even with llama.cpp you can store n-grams to disk and use it. I m using unsloth's llama fork atm but a new llama release is out that may be better - hadn't had time to test it yet.
Yes, this on Halogen. I can get ~45tok/sec which is quite usable.
well, it's free for now , beyond that eh.
Back in their 3.5/3.7 era [1], they had a small+fast model, specifically eschewing knowledge, and instead focused on "general intelligence" like managing tool calls, orchestrating sub-agents, etc. Could be a very good substrate for running a personal assistant (time will tell what's the right approach); so I'm curious to try that out and see how they've come along since then.
[1] https://news.ycombinator.com/item?id=47069179
Wonder if this was the space bunny alpha model
After some quick tests, it likely isn't. Additionally, Space Bunny Alpha had hints it was a Minimax-variant model.
Its been on ZenMux for at least 2 weeks so I doubt they are hiding
I like trying new models but I wish we’d get something actually new. Like a new architecture or something. LLMs are just so sloppish. We can do better.
welcome to the sigmoid...
Talk is cheap. You have to actually beat transformers first. All alternate architectures that have sprung up are sidegrades at best.
We have something better than transformers, but not better enough to shift investement to it.
It doesn't look competitive along any dimension: https://artificialanalysis.ai/models/step-5#intelligence-com...
Better luck next time.
It's fast. The average speed is 115 tokens/sec according to OpenRouter. I haven't tested the model to see how it is in practice, but I'd certainly pay a little extra for faster inference.
Edit: Though the average latency of 1.5s isn't very low, so it might not be that fast in practice for agentic work. Also, I don't know how much thinking it does, as that's generally been the drawback to Chinese models.
Using Artificial Analysis
Model | Reasoning | Intelligence Index | Artificial Analysis million output tokens for the intelligence index -|-|-|- Step 5 | ? | 44 | 160 GLM 5.3 | Max | 45 | 210 MiMo V2.6 Pro | ? | 46 | 140 Kimi K3 | Max | 44 | 160 Qwen Max 0902 | ? | 45 | 190 DeepSeek 4.1 Flash | Max | 39 | 250 GLM 5.3-flash | Max | 42 | 180 GPT-6 Astra | Low | 46 | 10
It looks reasonable by open model standards. This doesn't capture the fact that DeepSeek and Step 5 have much higher token/s than the rest, other than MiMo Ultraspeed. MiMo V2.6 is either slow but cheap, or fast but expensive. Just based on these numbers, it looks good. Astra-low is one of the fastest and cheapest because it doesn't use many tokens, but I've never tried it. I liked DeepSeek and GLM when I used them.
Artificial Analysis has some very specific biases or perspectives on what they are measuring, so I wouldn't take their benchmarks as the final word on model quality.
Composite indices are only useful for model companies which want to build one horizontal capability for N use cases, or naive users who don't want to bother with the effort of carefully pairing models with use cases. If you're a power user looking to understand and make deliberate choices, then you want pointed evaluations -- not general composite indices. You wouldn't hire the same person to do your taxes and mow your lawn, so why is it any different with LLMs? Only if you come at the problem with the folklore around "AGI" do you start making composite benchmarks.
As I mentioned in a sibling comment... Back in their 3.5/3.7 era [1], Stepfun had a small+fast model, specifically eschewing knowledge, and instead focused on "general intelligence" like managing tool calls, orchestrating sub-agents, etc. Could be a very good substrate for running a personal assistant (time will tell what's the right approach); so I'm curious to try that out and see how they've come along since then.
[1] https://news.ycombinator.com/item?id=47069179
> carefully pairing models with use cases
I've always wondered if there were niches that some models are better at. I use them for software, electrical and mechanical engineering plus other things. What are some notable subject-specific differences you found?
Come on... That chart shows Sonet 5.5 as better than Qwen3.8-Flash-Next. I have observed the opposite.