🔥 StepFun's Step 5 Scores 44 to GPT-5.6's 47 — at One-Seventh the Output Price
StepFun's Step 5 Preview scores 44 to GPT-5.6 Sol's 47 on the intelligence index at roughly one-seventh the output price, with full open weights due October 15.

Here’s a number to sit with: $2.70. That’s what StepFun will charge you for a million output tokens from Step 5 Preview, the 600-billion-parameter model its Shanghai team put on public API this morning. The closest Western model on the intelligence charts, OpenAI’s GPT-5.6 Sol, runs $20 for the same million tokens.
On the benchmark that matters most to buyers — the Artificial Analysis Intelligence Index — Step 5 Preview scores 44. Sol scores 47 at maximum effort, and drops to 44 when you dial its reasoning down to match Step 5’s speed. So you’re looking at a three-point gap at the top, no gap once you normalize for latency, and a price that’s roughly 7.4× cheaper on output.
And on October 15, StepFun says it will publish the full weights. Not an API you rent — the actual model, yours to download and run.
The thesis: near-frontier intelligence just stopped being a thing you can only rent from three American companies.
đź§ Why This Matters
For two years the story of frontier AI has been a story about who could afford it. The best models lived behind metered APIs, priced like enterprise software, and the gap between “good enough to demo” and “good enough to ship” was measured in dollars per million tokens that added up fast when your agent made a thousand calls a day.
Step 5 Preview attacks that gap directly. StepFun built it for agentic work — the long, tool-heavy chains where a model reads code, runs a terminal, checks a spreadsheet, and loops. Those workloads are exactly where output tokens pile up, and exactly where a 7× price cut compounds. If your product spends $200,000 a month on inference, the same workload at Step 5’s rates lands closer to $30,000.
“Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance.”
— StepFun, launch announcement
The open-weights promise raises the stakes again. A rented model can be repriced, rate-limited, or retired. A model you’ve downloaded is a fixed cost you control — and for banks, hospitals, and anyone who can’t send data to someone else’s server, that’s the whole ballgame.
📊 Deep Dive
Step 5 Preview is a sparse mixture-of-experts model: 600 billion parameters total, but only about 27 billion active per token (roughly 4.5%). That’s how you get big-model quality without big-model compute on every request. It ships with a 1-million-token context window and takes text and images in, text out (DataStudios).
Here’s how it stacks up against GPT-5.6 Sol on independent third-party evals, per OrcaRouter’s reading of the Artificial Analysis suite:
- Intelligence Index: Step 5 Preview 44 · GPT-5.6 Sol 47 (max effort), 44 (matched effort)
- Input price: $1.00/M vs $4.00/M — 4× cheaper
- Output price: $2.70/M vs $20.00/M — ~7.4× cheaper
- Cached input: ~$0.05/M (a 95% discount) vs $0.40/M
- SciCode: Step 5 59% · Sol 57% — a rare outright win
- Terminal-Bench 4.0: Step 5 33.3% · Sol 39.9% — Sol keeps the lead on raw terminal tasks
- Output speed: 99.8 tokens/sec vs 61.5
Read the board and the picture is consistent: Step 5 trades a little peak reasoning for a lot of speed and a big drop in price, and it ranks #24 of roughly 200 commercially available models on the index (Eastern Herald). That’s the frontier’s second tier, sold at the discount rack’s prices.
⚠️ The Catch
Benchmarks are not deployments. An Intelligence Index of 44 puts Step 5 a genuine notch below the very best proprietary models on the hardest reasoning, and on Terminal-Bench — arguably the most agentic test here — Sol still wins by six points. If your use case lives in that top slice, three index points can be the difference between a demo and a shipped feature.
There’s also the gap between vendor numbers and reality. StepFun and every lab quote their own internal benchmarks, which reliably run higher than independent evals; OrcaRouter flags that vendor-reported scores on tests like Terminal-Bench lack third-party verification. Trust the outside numbers, not the launch slide.
And “open weights on October 15” is a promise with a date attached, not a file you can download today. A 600B model is also not free to run — you need serious hardware to self-host it, which means the headline price only applies while you rent, and the self-host savings only apply if you’ve got the GPUs.
🎯 What Happens Next
Watch the pricing pages. When a credible model lands at one-seventh the output cost, the incumbents rarely match it head-on — they cut prices on their older tiers and push customers toward cheaper “mini” variants. Expect Step 5 to show up first in cost-sensitive agent products: coding assistants, document-processing pipelines, and finance tooling, where StepFun says its model is strongest.
The October 15 weights drop is the real event. If they ship on schedule and match the API’s quality, Step 5 joins the small club of open models good enough to build a business on — and every team currently paying frontier prices for routine agent calls will run the math (Pandaily).
đź§© Bigger Picture
StepFun was founded in April 2023 by Jiang Daxin, a former Microsoft executive, and is one of the Chinese startups investors have taken to calling the country’s “AI Tigers.” Backed by Tencent and Qiming Venture Partners at a valuation around $10 billion, it sits in a crowded field of well-funded labs (Wikipedia).
The strategic move here is the open-weights play. Renting a model keeps customers dependent; publishing the weights trades that recurring revenue for reach and mindshare, betting that being the default open model you can actually run beats being one more paid API in a saturated market. It’s the same wager several open-weight labs have made this year, now made with a model that scores within striking distance of the proprietary top.
For buyers, the takeaway is simpler than the geopolitics: the price of “good enough to ship” just fell by a factor of seven, and in six weeks you may not have to ask anyone’s permission to run it.
The frontier used to be a subscription. StepFun just turned part of it into a download.
Sources
- OrcaRouter — Step 5 Preview vs GPT-5.6 Sol: 44 vs 47, benchmark and price breakdown
- DataStudios — StepFun launches Step 5 Preview: 600B params, 1M context, open weights Oct 15
- Pandaily — StepFun Launches Step 5 Preview: 600B Sparse MoE, 1M Context
- Eastern Herald — StepFun Step 5 Preview: model priced at $1 per million tokens
- Wikipedia — StepFun (company background and funding)


