> Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.
> Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.
> The model will be released with open weights on October 15.
I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5).
I wonder if other Chinese labs like Kimi/Moonshot will follow suit.
Parent commenter hinted at that. Yet DeepSeek has released their V4 which was hugely successful, and even their new architecture is marked V4.1. Qwen internals mark their Flash-Next model, also very compelling, as "qwen4exp". So both of them are bucking the negative stereotype.
600b-a27b doesn’t sound enticing. Also with the higher number of active parameters compared to GLM 5.3 flash and DeepSeek V4/4.1 flash, I don’t see how they want to be more efficient at inference.
nh43215rgb · · focus · HN ↗
Bolwin · · focus · HN ↗
dannyw · · focus · HN ↗
It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).
JohnsonZou · · focus · HN ↗
zozbot234 · · focus · HN ↗
torginus · · focus · HN ↗
Tepix · · focus · HN ↗
NooneAtAll3 · · focus · HN ↗
NetOpWibby · · focus · HN ↗
I agree with you though, ChronVer all the way.