MiniMax integrates M3.1-Flash-Preview model for coding into MiniMax Code product
Read more
Pandaily
pandaily.com

MiniMax integrates M3.1-Flash-Preview model for coding into MiniMax Code product

MiniMax announced on September 27, 2026, that the M3.1-Flash-Preview model is now available within its product for assisting with code writing in MiniMax Code. This information was confirmed by sources, including IT Home, Startup Fortune, and other developer reviews. The official English name of the product is MiniMax / MiniMax Code.

The main focus of this review is on the in-product coding model and the reasoning controls. However, MiniMax has not yet published a full public specification of parameters or an open API for M3.1-Flash-Preview, so these details cannot be confirmed.

The company's marketing message presents this Preview model as an ultra-fast text model designed for daily development. It is suitable for tasks such as debugging, generating function snippets, handling edge cases, regression testing, and checking the impact of changes, rather than being a new flagship for general chat.

Reviews describe a closed loop operation within MiniMax Code that covers problem localization and implementation, as well as verification through testing and subsequent delivery. Developers who noticed the model in the product list a day before the official announcement were already considering it as an SKU for coding workloads, coexisting with older M3 and M2.7 variants in one interface.

A practical difference noted in the English product reports is the five-level reasoning resource slider, which ranges from low to a new maximum level. This slider allows users to balance computation depth against latency when performing routine or more complex coding tasks. This control is located inside MiniMax Code; the Preview endpoint is described as limited to the company's own tool, not as a widely documented public API in this launch.

Parallel promotions include a Token Plan reset and a verification campaign from September 28 to October 7, which offers double points applicable to code models, including M3.1-Flash-Preview. Users should remember that the M3 architecture and SWE-bench metrics relate to the previous M3 flagship, not automatically to Flash-Preview, until MiniMax releases a specific card.

For coding tool users, the main news is that MiniMax Code has received the Flash-Preview coding model with an explicit reasoning effort slider—a component built into the development environment (IDE), not an open-weights release.

Similar stories

NaiveAI releases Naive-N0.5-Flash model: 309B MoE with 1M context under MIT license
Read more
pandaily.com

NaiveAI releases Naive-N0.5-Flash model: 309B MoE with 1M context under MIT license

NaiveAI has introduced the open-weight Naive-N0.5-Flash model on the Hugging Face platform. This model features a Mixture-of-Experts (MoE) architecture designed for software development tasks and artificial intelligence research. Information about the model was published in mid or late September 2026.

The official product name is NaiveAI / Naive-N0.5-Flash. According to technical documentation, the total number of parameters is approximately 309 billion, with 15.5 billion weights being active. The model is distributed under the MIT license and includes inference code, as well as native support for a one-million-token context window.

The context length is achieved through a hybrid combination of the Sliding-Window Attention (SWA) mechanism and lightweight DeepSeek Sparse Attention (DSA), which utilizes grouped attention. The architecture is described as having a predominant SWA–DSA ratio of approximately 5:1, with no fully attentive layers in the stack. The model is based on the open-source MiMo-V2.5 model, followed by fine-tuning and post-training, rather than full pre-training from scratch.

In terms of serving, NaiveAI emphasizes NaiveRT—an inference stack developed using AI-oriented research. This stack integrates mega-kernel fusion, Programmatic Dependent Launch, and speculative decoding. According to the vendor, processing speed reaches approximately 50 tokens per second per user in Standard mode and up to 2000 tokens per second in Ultra-Fast mode.

Tags on Hugging Face highlight that the model is suitable for code text generation, long-context handling, and AI research tasks. This aligns with the stated concept that fine-tuning and post-training on a strong open base can push boundaries for coding agents. For developers comparing Chinese MoE releases this week, Naive-N0.5-Flash is a specific example of a 309B / 15.5B MoE with a million-token hybrid context under the MIT license, representing an architectural announcement rather than a cost or product launch news.

Xiaomi releases weights and resources for MiMo-V2.6 Pro and MiMo-V2.6 Flash models using RL stack
Read more
pandaily.com

Xiaomi releases weights and resources for MiMo-V2.6 Pro and MiMo-V2.6 Flash models using RL stack

Xiaomi has made the weights, technical documentation, and reinforcement learning (RL) training resources for its MiMo-V2.6 series publicly available. According to English notes from Xiaomi dated September 22, the MiMo-V2.6-Pro and MiMo-V2.6-Flash repositories were published on the Hugging Face platform along with the RL code and over 7,000 test environments.

This release involves providing weights and code, which differs from previous coverage of the RL training process via live streaming. The branding remains unchanged: Xiaomi / MiMo.

The Pro and Flash models are positioned as native multimodal models with a claimed context window of one million tokens. Model cards and supplementary reports in English indicate that the Pro model features a sparse mixture-of-experts architecture with a total parameter count of approximately 1.02 trillion, actively utilizing about 42 billion parameters. The Flash model is rated at 309 billion total parameters with an activity level of around 15 billion.

Xiaomi reports that each model underwent 30 RL steps across approximately 750,000 trajectories in less than six days. The process utilized task mixing, including coding, general agent work, visual tasks, and cybersecurity. In each update, approximately 1,568 queries and 16 runs were used, with a data volume per step of 3.5–3.7 billion tokens.

According to the company, performance gains in training tasks were observed at approximately 25% for Flash and 12% for Pro. Furthermore, improvements were recorded in DeepSWE v1.1: from 48.8 to approximately 65.7 for Flash and from 58.4 to approximately 72.6 for Pro; however, these figures are vendor-provided data and require third-party verification.

The available assets go beyond mere checkpoints. Xiaomi has provided a comprehensive RL framework built on verl and related agent mechanisms, over 7,000 classified environments covering software development, vulnerability reproduction, intellectual labor, and web design. A Distill-Qwen-9B starting point is also available for community RL experiments, along with lightweight components for multichannel training. The API cost for hosted Pro and Flash versions is stated to be comparable to MiMo-V2.5, while the UltraSpeed tier promises up to 20 times higher throughput than standard Pro while maintaining quality.

Nevertheless, engineers still face high maintenance costs for large MoEs, and Xiaomi's claims regarding artificial intelligence analysis and agent benchmarks require external reproduction. The main news is the complete MiMo-V2.6 package with open weights, RL code, and environments that laboratories can study, rather than just a graphical representation in a live stream.

Popular