StepFun's Step 5 Preview: 600B MoE agent model, 44 on the Artificial Analysis Index, open weights on October 15
- StepFun launched Step 5 Preview, a sparse Mixture-of-Experts model with 600B total parameters and 27B active per token, a 1M-token context window, vision input, and open weights promised for October 15.
- It trails Claude Opus 5 and GPT-6 Astra on nearly every published benchmark: DeepSWE v1.1 67.7 against 74.0 and 74.1, StepCodeBench 49.0 against 63.9 and 61.0, Terminal-Bench v4 33.3 against 52.3 and 57.9, and GDPval-AA v2 1571 against 1735 and 1580.
- Among the Chinese labs it beats Kimi K3 and GLM-5.3 on StepCodeBench (49.0 versus 43.9 and 40.2) and ProgramBench (80.5 versus 77.8 and 72.0), but loses Terminal-Bench v4 to GLM-5.3's 41.9.
- StepFun says Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index and that its cost per task at that intelligence level is substantially lower than similarly capable models.
- Roughly 70% of internal and external expert evaluators judged the model able to autonomously solve moderately high complexity coding tasks.
Hacker News opinions
For anyone else hunting for the pricing, it's on the platform.stepfun.ai docs page. Took me a minute to find it.
I regularly burn 200-300M cached reads a day on some models, and 7-800M a couple of times. At 8-12 a day just for cached reads, so cheap headline pricing doesn't tell you much.
Huh, wonder why they skipped 4?
Chinese labs sometimes skip 4 because it's considered unlucky.
Right, 四 sounds like death in Chinese, so it gets avoided.
Good to see FireRed used as a benchmark again. Step 5 Preview ran 3,000+ turns and 6 million tokens with no Pokemon-specific tuning, and by turn 3,082 it had Cut, three gym badges, and Lt. Surge down, about a third of the main story. I bet Astra can beat that in 18 hours, not sure how it compares though.
600B total params, 27B active, 1M context, 44 on the Artificial Analysis Index, open weights on October 15. Skipping 4 is a Chinese-company thing, but it also puts them on the same iteration number as Claude Opus 5. Curious if Kimi and Moonshot do the same.
Moonshot already teased K3.1, so probably not.
Could just be the unlucky number thing, 4 carries that weight in traditional Chinese culture.
Their positioning is the smart part. Instead of saying they're cheaper and a bit less capable, they say they're the best among the cheaper and slightly less capable ones.
Their AA Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is 2.70 per million tokens, open weights October 15.