Huawei releases open-source training code for openPangu-2.0 (Pretrain, SFT, and RL) on the Ascend platform
Read more
Pandaily
pandaily.com

Huawei releases open-source training code for openPangu-2.0 (Pretrain, SFT, and RL) on the Ascend platform

Huawei has released the source code for the pretraining, Supervised Fine-Tuning (SFT), and Reinforcement Learning (RL) processes for the openPangu-2.0 model. This release occurred on September 28, 2026, according to information from TMT Post and coverage related to the training stack optimized for Ascend.

The official branding is used as Huawei / openPangu. Previously, the weights for openPangu-2.0-Pro and openPangu-2.0-Flash were released on the Hugging Face platform; however, the current release pertains specifically to the training code path that underpins these MoE models trained on Ascend, rather than releasing new parameters.

The model cards on Hugging Face describe openPangu as Huawei's brand of open artificial intelligence models designed for training and inference on Ascend. The openPangu-2.0-Pro model features approximately 505 billion total parameters with about 18 billion active per token, a 512K context window, and a training budget of around 34 trillion tokens. The openPangu-2.0-Flash model has about 92 billion total parameters, approximately 6 billion active, the same 512K context size, and a comparable budget of ~34 trillion tokens.

Additional details following the training of both models mention combined fast/slow SFT, multi-specialized RL, and Online Distillation (OPD). Architectural features common to Pro and Flash include multi-head latent attention, DSA-plus-SWA layered mix (approximately 1:2), an mHC residual topology with four branches, three-head multi-token prediction, and training using the Muon optimizer.

The September 28th release provides the pretraining, SFT, and RL components as Ascend-specific tools, allowing developers to reproduce and extend this stack instead of merely loading weights for inference. It is important to distinguish between open weights, inference code, and this training code set: the weights were already publicly available; the news is that Huawei is opening up the openPangu-2.0 training pipeline on the Ascend side for the pretrain/SFT/RL components.

For teams working on Ascend, the concrete outcome is an OSS training stack corresponding to the confirmed Pro family models (505B / ~18B active, 512K) and Flash (92B / ~6B active, 512K).

Similar stories

Xiaomi releases weights and resources for MiMo-V2.6 Pro and MiMo-V2.6 Flash models using RL stack
Read more
pandaily.com

Xiaomi releases weights and resources for MiMo-V2.6 Pro and MiMo-V2.6 Flash models using RL stack

Xiaomi has made the weights, technical documentation, and reinforcement learning (RL) training resources for its MiMo-V2.6 series publicly available. According to English notes from Xiaomi dated September 22, the MiMo-V2.6-Pro and MiMo-V2.6-Flash repositories were published on the Hugging Face platform along with the RL code and over 7,000 test environments.

This release involves providing weights and code, which differs from previous coverage of the RL training process via live streaming. The branding remains unchanged: Xiaomi / MiMo.

The Pro and Flash models are positioned as native multimodal models with a claimed context window of one million tokens. Model cards and supplementary reports in English indicate that the Pro model features a sparse mixture-of-experts architecture with a total parameter count of approximately 1.02 trillion, actively utilizing about 42 billion parameters. The Flash model is rated at 309 billion total parameters with an activity level of around 15 billion.

Xiaomi reports that each model underwent 30 RL steps across approximately 750,000 trajectories in less than six days. The process utilized task mixing, including coding, general agent work, visual tasks, and cybersecurity. In each update, approximately 1,568 queries and 16 runs were used, with a data volume per step of 3.5–3.7 billion tokens.

According to the company, performance gains in training tasks were observed at approximately 25% for Flash and 12% for Pro. Furthermore, improvements were recorded in DeepSWE v1.1: from 48.8 to approximately 65.7 for Flash and from 58.4 to approximately 72.6 for Pro; however, these figures are vendor-provided data and require third-party verification.

The available assets go beyond mere checkpoints. Xiaomi has provided a comprehensive RL framework built on verl and related agent mechanisms, over 7,000 classified environments covering software development, vulnerability reproduction, intellectual labor, and web design. A Distill-Qwen-9B starting point is also available for community RL experiments, along with lightweight components for multichannel training. The API cost for hosted Pro and Flash versions is stated to be comparable to MiMo-V2.5, while the UltraSpeed tier promises up to 20 times higher throughput than standard Pro while maintaining quality.

Nevertheless, engineers still face high maintenance costs for large MoEs, and Xiaomi's claims regarding artificial intelligence analysis and agent benchmarks require external reproduction. The main news is the complete MiMo-V2.6 package with open weights, RL code, and environments that laboratories can study, rather than just a graphical representation in a live stream.

Xiaomi releases embodied world foundation model and Robotics-U0 training kit
Read more
pandaily.com

Xiaomi releases embodied world foundation model and Robotics-U0 training kit

Xiaomi has made its product, Xiaomi-Robotics-U0—an autoregressive foundational model of the embodied world—available to the public. This model views robot-oriented synthesis not merely as a specialized process for trajectory generation, but as an extension of image and video generation based on foundational models.

According to official materials, the main line of the model contains approximately 38 billion parameters and is continuously trained to achieve embodied intelligence. Additional public weights, including variants Xiaomi-Robotics-U0-4B and FlashAR, are available on Hugging Face and ModelScope platforms, and inference code, demonstration via Gradio, and FSDP training releases dated early September are provided.

Within a single next-token framework, this model integrates several functions: text-to-image generation, image editing from any source, multi-view scene creation, embodied controlled transfer, and embodied video generation. Xiaomi asserts that this stack preserves fundamental visual semantics while adding robot-oriented reasoning capabilities through both single-step and sequential training stages. These stages include alternating goal-subtask streams and embodied video at frame rates of 1, 3, and 5 frames per second.

Architectural notes indicate the use of IBQ image tokenization and a multimodal vocabulary that supports autoregressive scaling across different modalities. This architecture is built upon components derived from the Qwen3-32B and EMU3.5 lines.

A second crucial aspect is inference efficiency. The FlashAR+ technology replaces sequential decoding of image tokens with anti-diagonal parallel generation and is used in conjunction with vLLM batch processing. Xiaomi reports a potential acceleration of 1024x1024 image generation up to nearly 83 times faster, reducing the latency for a single sample from approximately 450.8 seconds to 5.44 seconds on the stated hardware.

Capability pages claim that performance in multi-view scenes and transfer surpasses GPT-Image-2 when internally splitting data into complex and simple sets. Furthermore, UNIS (Xiaomi-Robotics-U0) ranks first among over 100 models on WorldArena with an EWMScore_P of 73.64.

In terms of application, style transfer data from the model is mixed with real demonstration datasets for subsequent policy training. Xiaomi reports an increase in intervention task progress from approximately 36.9% to 63.2% when using backgrounds and lighting not present in the training, with a smaller decrease compared to the baseline lab level.

Although video generation checkpoints are still catching up to the image and transfer releases, teams must verify licensing conditions, FlashAR engine selection, and third-party WorldArena reproducibility before using the open weights as a ready-made data factory for production robots.

Popular