StepFun releases preliminary version of Step 5 model: agent with 600 billion parameters and open weights expected on October 15
Read more
Pandaily
pandaily.com

StepFun releases preliminary version of Step 5 model: agent with 600 billion parameters and open weights expected on October 15

StepFun introduced the preliminary version of Step 5 on September 20, 2026, as its next flagship foundational model designed for long-horizon agents, according to information from Tencent Tech, DataLearner model cards, and company materials available at stepfun.com. The official English name of the model is StepFun / Step 5.

It is important to note that although the model identifier step-5-preview is already available via the product API and the open StepFun platform, the open weights in BF16 format are scheduled only for October 15, 2026, and were not available at launch. Therefore, availability via API and open weights should be considered as separate points.

The model architecture features a Mixture of Experts (MoE) sparse design with an approximate total of 600 billion parameters. Approximately 27 billion parameters are activated per token within a 92-layer 'narrow and deep' Transformer. The choice of depth is driven by agent requirements, as longer information paths through layers contribute to improved multi-step implicit reasoning and handling of long tool outputs.

The model supports a one-million-token context window, accepting both text and image input (though video input is mentioned on third-party cards). It is oriented towards applications in AI coding, software development, financial analysis, and professional agent usage. To maintain practicality with the million-token attention, sparse GQA was applied in combination with token block merging, reducing the cost of indexer and top-k selection by approximately one eighth.

During training, special attention was paid to bit-level alignment between training and inference to ensure stable MoE routing, and load-aware scheduling, speculative decoding, and FP8 paths were utilized. StepFun claims that these methods accelerate the long-horizon Reinforcement Learning (RL) process by more than three times overall.

According to data from Artificial Analysis's composite AI index, it scores 44 points. The company places this rating among leading models focused on open access. Aggregator cards also list API prices: about $1.00 per input token and $2.70 per output token per million tokens, but comparisons of metrics and cost should be viewed as statements from third parties or providers.

Until October 15, Step 5 Preview should primarily be regarded as a functional API for long-horizon agents, having only a planned commitment for open weights, rather than a release of weights on Hugging Face or a drop-in replacement for Meituan LongCat, MiniMax Code Flash, or NaiveAI OSS MoE.

Similar stories

NaiveAI releases Naive-N0.5-Flash model: 309B MoE with 1M context under MIT license
Read more
pandaily.com

NaiveAI releases Naive-N0.5-Flash model: 309B MoE with 1M context under MIT license

NaiveAI has introduced the open-weight Naive-N0.5-Flash model on the Hugging Face platform. This model features a Mixture-of-Experts (MoE) architecture designed for software development tasks and artificial intelligence research. Information about the model was published in mid or late September 2026.

The official product name is NaiveAI / Naive-N0.5-Flash. According to technical documentation, the total number of parameters is approximately 309 billion, with 15.5 billion weights being active. The model is distributed under the MIT license and includes inference code, as well as native support for a one-million-token context window.

The context length is achieved through a hybrid combination of the Sliding-Window Attention (SWA) mechanism and lightweight DeepSeek Sparse Attention (DSA), which utilizes grouped attention. The architecture is described as having a predominant SWA–DSA ratio of approximately 5:1, with no fully attentive layers in the stack. The model is based on the open-source MiMo-V2.5 model, followed by fine-tuning and post-training, rather than full pre-training from scratch.

In terms of serving, NaiveAI emphasizes NaiveRT—an inference stack developed using AI-oriented research. This stack integrates mega-kernel fusion, Programmatic Dependent Launch, and speculative decoding. According to the vendor, processing speed reaches approximately 50 tokens per second per user in Standard mode and up to 2000 tokens per second in Ultra-Fast mode.

Tags on Hugging Face highlight that the model is suitable for code text generation, long-context handling, and AI research tasks. This aligns with the stated concept that fine-tuning and post-training on a strong open base can push boundaries for coding agents. For developers comparing Chinese MoE releases this week, Naive-N0.5-Flash is a specific example of a 309B / 15.5B MoE with a million-token hybrid context under the MIT license, representing an architectural announcement rather than a cost or product launch news.

Inspur unveiled MetaBrain SD200 Ultra supernode with 128 domestic chips for running Kimi K3 model
Read more
pandaily.com

Inspur unveiled MetaBrain SD200 Ultra supernode with 128 domestic chips for running Kimi K3 model

Inspur Information has introduced the MetaBrain SD200 Ultra artificial intelligence supernode at the AICC2026 Artificial Intelligence Computing conference. This node is positioned as a capability-class system for processing tasks involving trillion-parameter models and agents. The official English name of the system is Inspur / MetaBrain.

According to company data and information agencies, the node integrates 128 domestic AI chips, provides access to 8 TB of unified accelerator memory, and 64 TB of host memory. It is capable of running the Kimi K3 model from Moonshot AI, which has 2.8 trillion parameters, on a single machine with a token generation latency of less than 5.85 milliseconds, equating to approximately 170 tokens per second for a single user, according to Inspur's estimates.

The architectural features of the node include the 3D Hyper Mesh factory, which ensures native, semantics-sensitive memory communication with a claimed latency of 0.69 microseconds. Short-reach copper interconnects and symmetric memory are also utilized, allowing accelerators to directly access the memory of remote nodes. Inspur claims that the AllReduce operation execution time has been reduced by approximately 3.5 times compared to previous solutions.

Furthermore, integrated 'superoperators' for the KDA, gated MLA, and MoE stages of the Kimi K3 model reduce the number of operators by about ten times while increasing inference performance by more than threefold. These figures are vendor claims. Materials also indicate the potential to support models up to 10 trillion parameters on a single node.

Another product, MetaBrain HC2000, is a rack with multiple accelerators designed for capacity-class inference, boasting a claimed tenfold tokens-per-investment throughput under agreed SLA constraints. Although various publications describe the accelerators only as domestic or local AI chips, none of the primary sources analyzed for this review name the silicon manufacturer or the SKU within the SD200 Ultra. Therefore, this article does not attribute the node to Ascend, Cambricon, or any other specific supplier.

StepFun releases preliminary version of Step 5: a model with 600 billion parameters, sparse MoE, and 1 million token context
Read more
pandaily.com

StepFun releases preliminary version of Step 5: a model with 600 billion parameters, sparse MoE, and 1 million token context

StepFun has introduced the preliminary version of the Step 5 model, which is a base model featuring a sparse Mixture-of-Experts (MoE) architecture and approximately 600 billion total parameters, with about 27 billion activated per token. This release targets long-horizon agent workloads, including AI-assisted code development, software engineering, financial analysis, and professional knowledge work. The model supports a one-million-token context window and accepts both text and image inputs. StepFun announced that API access is already open, and the model weights are scheduled to be made publicly available on October 15th.

Instead of expanding the network, Step 5 Preview utilizes a narrow, deep Transformer architecture consisting of 92 layers. The company asserts that deeper stacks provide longer information pathways for implicit multi-step reasoning during long prefill, when agents perform searches, run code, and process tool usage results. To maintain practicality in the million-token sessions, the model incorporates Sparse Grouped-Query Attention with token block merging. According to StepFun, this reduces the cost of indexer selection and top-k by roughly one-eighth compared to a denser base model while consolidating overlapping adjacent selections.

The training process emphasizes on-policy long-horizon reinforcement learning, as well as bit-level alignment of training and inference in MoE routing. Furthermore, techniques such as load-aware scheduling, MTP-3 speculative decoding, FP8 MoE, and KV cache offloading are employed. StepFun reports more than a threefold acceleration of the end-to-end RL process for the long horizon and a sample registry loss metric below one percent. This model is also used within a human-managed data pipeline that generates verifiable complex tasks at scale in the millions, covering science, software development, and machine learning research.

Regarding AI analysis, Step 5 Preview scores around 44 points on an intelligence index that StepFun ranks among the best open-weights models in this rating system. The API cost, according to published pricing, is approximately $1 per million input tokens and $2.7 per million output tokens, with an output speed of about 100 tokens per second. The model's primary focus is on agent architecture and efficiency, distinguishing it from previous releases like Step 3.5 Flash and Step 3.7 Flash, emphasizing depth, sparse long context, and reliable multi-stage tool use before the weights become available in October.

Popular