Step 5 Preview: 600B-Parameter Model Pushes the Pareto Frontier in Coding and Finance

Step 5 Preview: Advancing the Pareto Frontier

Step 5 Preview: 600B-Parameter Model Pushes the Pareto Frontier in Coding and Finance

StepFun's Step 5 Preview is a sparse Mixture-of-Experts model with 600B total parameters and 27B active per token, featuring a 1M-token context window and vision input. It delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. On DeepSWE v1.1, it scores 67.7, just behind GPT-6 Astra (74.1) and Claude Opus 5 (74.0). It also excels in long-horizon tasks, optimizing an MLA GPU kernel to 508 TFLOPS in 22 hours and improving a post-training data loop to 60% on AIME24.

Progress begins when that boundary shifts outward.
  1. BoppreH

    In their first demo video, to make a 3D render of the photo, the thinking trace gives away the game:

    > Interesting! It turns out there's already an existing project here [...] The project is fully built [...]

    I'm always astounded how little effort is put into checking the AI answers displayed in these announcements. Back when I paid more attention, I remember OpenAI's and Google's demos constantly showed their AIs giving wrong answers.

  2. nh43215rgb

    > Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.

    > Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.

    > The model will be released with open weights on October 15.

    I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5).

    I wonder if other Chinese labs like Kimi/Moonshot will follow suit.

  3. bethekind

    > Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story.

    Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.

  4. garo-pro

    IT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.

  5. InsideOutSanta

    GLM-5.3 and Kimi K3 are just below where I can use them to completely replace frontier models. Oddly,* SWE-2 is there for me.

    If this performs similarly in the real world, we're approaching a level of capability where for most devs, it only makes sense to pay for Anthropic or OpenAI subscriptions if they are heavily subsidized and actually cheaper than these alternative options.

    * Oddly, because I perceived Devin as being kind of a joke before trying SWE-2.

More from this day

2026-09-20