StepFun released Step 5 Preview on September 20, 2026, priced at $1 per million input tokens, about a quarter of what several frontier labs charge. It only turns on 27 billion of its 600 billion parameters per token, which is what makes that price work. The API is live now; open weights follow October 15, license unannounced.

  • Sparse mixture-of-experts model: 600B total parameters, 27B active per token (about 4.5%), 1 million token context, text plus image input.
  • Pricing: $1.00 per million input tokens, $2.70 per million output tokens, 95% discount on cached input (about $0.05 per million).
  • Artificial Analysis scores it 44 on its Intelligence Index, tied with Kimi K3 (max) at roughly 65% lower cost per task, ranked 27th of 653 models.
  • No SWE-bench Verified score exists yet, so it enters GenZTech's AI Coding Leaderboard as verifying, with no rank.
Sparse activation vs dense serving cost27B of 600B parameters activate per token, about 4.5%, priced at $1.00 input and $2.70 output per million tokens versus a costlier dense build.600B totalstored weights27B active~4.5%, per token$1.00 / $2.70per million tokensSame 600B, denseall params, every tokenCosts moresize sets the billStepFun’s architectureHypothetical dense buildgenztech.blog
Fig 1 27B of 600B parameters activate per token, about 4.5%. Cost tracks what computes, not what sits in memory, which is the lever behind a $1.00 input price.

What exactly did StepFun ship?

Step 5 Preview is StepFun’s new flagship, built for agentic work rather than chat. The Shanghai lab’s announcement calls it a model "delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance," under the tagline "Advancing the Pareto Frontier": more capability per unit of cost. It is live now at platform.stepfun.ai, with open weights following October 15, license unspecified.

RelatedLaguna S 2.1: 8B Active Params, 70% Terminal-Bench

Why does 4.5% activation matter?

A mixture-of-experts model routes each token through a handful of specialized experts, instead of running every parameter for every word the way a dense model does. Step 5 Preview activates 27B of its 600B stored parameters, about 4.5%. That ratio is the whole economic argument: cost tracks what computes, not the weight count in memory. A dense 600B model would run all of it every token, a heavier, pricier job. Idle most of the network per token, and $1.00 input, $2.70 output pricing stays a business rather than a loss.

How does it stack up against Kimi K3 and StepFun's own last model?

The Kimi K3 row is sparse on purpose: Artificial Analysis ties the two on one number, a shared Index score of 44, with Step 5 Preview reaching it at roughly 65% lower cost per task, and nothing else about Kimi K3 disclosed here. Step 3.5 Flash, StepFun's own February model, is the sharper contrast: smaller, fully open weight, built for local deployment and coding agents.

ModelStep 5 PreviewKimi K3 (max)Step 3.5 Flash
Total / active600B / 27BNot published196B / 11B
Context1M tokensNot published256K tokens
Price (input)$1.00 / MNot publishedNot published
Open weightsOct 15, 2026, license TBDNot publishedYes
AA Index4444, ~65% higher cost/taskNot published

Where does the benchmark story get complicated?

Independently, Artificial Analysis put Step 5 Preview at 44 on its Intelligence Index, 27th of 653 models tracked, at a weighted cost of $0.71 per task, a real third-party number. StepFun's own table is a different category: 33.3 on Terminal-Bench v4 against Opus 5's 52.3 and Astra's 57.9, and 29.5 on Agents' Last Exam against Astra's 33.3, a real gap on agentic work. Finance closes it: 66.4 on FrontierFinance against Opus 5's 69.7 and Astra's 55, and 83.3 on DRACO against Opus 5's 87.6 and Astra's 76.8. StepFun's claim: fewer parameters, competitive on two benchmarks, behind on the rest. A vendor picks its own benchmarks, and cheap tokens stay cheap only if the agent finishes the job instead of retrying. No SWE-bench Verified score exists from anyone yet, which is why Step 5 Preview sits on GenZTech's AI Coding Leaderboard as verifying, without a rank.

StepFun's benchmark table: Step 5 vs Opus 5 vs AstraVendor-reported. Terminal-Bench v4 and FrontierFinance scores for Step 5 Preview, Claude Opus 5 and GPT-6 Astra.33.352.357.966.469.755.0Step 5Opus 5AstraStep 5Opus 5AstraTerminal-Bench v4FrontierFinanceVENDOR-REPORTEDgenztech.blog
Fig 2 · vendor benchmark StepFun’s own table shows Step 5 Preview well behind Opus 5 and Astra on Terminal-Bench v4, but within a few points on FrontierFinance, the split its finance pitch is built around. Not independently verified.

How did StepFun get here?

  1. 2023StepFun founded. Shanghai, by Jiang Daxin, ex-Microsoft (Bing, Cortana, Azure Cognitive Services).
  2. Dec 2024Series B. Several hundred million dollars, led by Fortera Capital.
  3. Jan 2026Series B+. More than $718M, led by Shanghai SDIC Leading Fund, China Life Private Equity Investment and Pudong Venture Capital.
  4. Feb 12, 2026Step 3.5 Flash ships. Open-weight MoE, 196B/11B, 256K context.
  5. Sep 20, 2026Step 5 Preview API goes live. 600B/27B MoE, $1.00 / $2.70 per million tokens.
  6. Oct 15, 2026Open weights due. License not yet specified.

Who is this actually built for?

The finance framing is not incidental. FrontierFinance and DRACO are the two vendor benchmarks where Step 5 Preview closes most of the gap to Claude Opus 5, and StepFun is pitching exactly the buyers who care: trading desks, research shops and finance-adjacent agent builders who need a model to chew through long documents cheaply and often, not necessarily win every coding contest. StepFun is private, founded in 2023 in Shanghai by CEO Jiang Daxin, 16 years at Microsoft across Bing, Cortana and Azure Cognitive Services, with a computer science PhD from the University at Buffalo. The funding behind that pricing is in the timeline below.

What does an unspecified license mean for builders?

StepFun says open weights arrive October 15, then stops short of naming the license. Qwen is the cautionary precedent: some of its weight releases carried research-only terms that blocked commercial use developers assumed "open weights" already covered, discovered only after integrating the model. A license that restricts commercial use or caps usage by company size changes the calculus against just paying StepFun’s API price. Until the terms land, open weights is a promise with an asterisk.

RelatedKimi K3 Is the Largest Open Model Ever, and It Is Not Cheap

What it means for the market

StepFun is private, so there is no ticker to point at. The read-through is pricing pressure, not a stock pick. A $1 input price on a 600B model with a 1 million token context leans on what incumbents like Anthropic and OpenAI can charge for comparable agentic workloads. It extends a pattern started by DeepSeek and continued through Qwen and Kimi: Chinese labs shipping open-weight or discounted frontier-adjacent models on a cadence Western labs have not matched. The signal for investors is margin compression at the inference layer, not a winner between two named companies.

What to watch · 2026-2027
  • SWE-bench Verified. Unranked on GenZTech's leaderboard until a score appears.
  • The October 15 license. Restrictive terms turn open weights into a marketing line.
  • The agentic gap. Terminal-Bench v4 and Agents' Last Exam sit furthest from the frontier.
  • Finance uptake. FrontierFinance and DRACO are StepFun's strongest argument.
  • Rival pricing. $1 per million tokens is the kind of number that shows up on the next pricing page.

Our take

27B active out of 600B is a real reason a model this size can be priced this low, and an independent 44 at a fraction of Kimi K3's cost per task backs the pricing as more than marketing. But StepFun chose its own benchmark table, and a lab always chooses the ones that flatter it. Terminal-Bench v4 and Agents' Last Exam show a real gap to Opus 5 and Astra on exactly the agentic tasks this release is built for. The finance strength reads as genuine: a narrow claim backed by two named benchmarks, not a claim of being great at everything. October 15 is the detail worth holding StepFun to. A workable license makes this a real self-host option; a restrictive one makes today a price cut with extra steps.

Primary sources

Original analysis by GenZTech. Source: StepFun.