> Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.
> Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.
> The model will be released with open weights on October 15.
I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5).
I wonder if other Chinese labs like Kimi/Moonshot will follow suit.
600b-a27b doesn’t sound enticing. Also with the higher number of active parameters compared to GLM 5.3 flash and DeepSeek V4/4.1 flash, I don’t see how they want to be more efficient at inference.
nh43215rgb · · focus · HN ↗
Tepix · · focus · HN ↗