Alibaba's XekRung Model Takes First Place in CyberGym Ranking with 88.9% Score
Read more
Pandaily
pandaily.com

Alibaba's XekRung Model Takes First Place in CyberGym Ranking with 88.9% Score

Alibaba's cybersecurity large language model, XekRung, has secured the top spot in the CyberGym ranking, which serves as a benchmark for vulnerability reproduction and is supported by researchers from the University of California, Berkeley. This version, designated as XekRung-1.5-27B-Preview and fine-tuned based on Alibaba's open-source model Qwen3.8-27B, demonstrated a success rate of 88.9% as of September 13th. Chinese tech media reported this result on September 28th, noting it is the first time a model from a Chinese developer has led this ranking.

The CyberGym test includes 1507 real vulnerabilities from 188 open-source projects that were initially discovered by Google's OSS-Fuzz program. In the first level of the task, the model receives a vulnerability description and unpatched code, and must then create an input file in the form of a proof-of-concept that triggers the error in the vulnerable version but not in the patched one.

In the current ranking, the XekRung model surpasses Google's Gemini 3.8 Flash Cyber with a score of 86.3%, OpenAI's GPT-5.5-Cyber with 85.6%, Zhipu AI's GLM-5.3 with 84.5%, and DeepSeek-V4-Pro with 83.3%. Supporters of this benchmark note that evaluations are provided by separate teams, runs are stochastic, and small differences in scores may not reflect significant differences in capabilities.

The model was developed in Alibaba's AGI Security Lab under the guidance of Huang Luntao. According to the team's report, the base model Qwen3.8-27B achieved 54.51% under the same conditions, representing an increase of 34.39 percentage points after additional training. The lab explains this success through supervised fine-tuning on data anonymized from Alibaba's own security operations, reinforcement learning of agents within a general instrumental framework, reward collection based on build, failure, and PoC validation, as well as reusing failed attempts as corrective pairs for training. It is emphasized that no CyberGym tasks, patches, or reference PoCs were used during the training process. The evaluation was conducted on locally deployed FP8 inference with a 256K context window, without pre-installed fuzzing frameworks, and with limited network access according to the list permitted in the benchmark.

Alibaba also emphasizes efficiency, stating that the 27-billion parameter model is 1/27th or 1/370th the size compared to comparable cyber models, which will reduce computational costs for automated vulnerability sorting. The weights of XekRung have not yet been published, and Alibaba has not provided a timeline for external access, although Huang stated that the lab's security technology will be offered more broadly, and future versions will target adversarial intelligence, self-evolution, and agent tasks. A technical report on the 8-billion parameter version built on Qwen was previously released in May.

Similar stories

Alibaba launches Zhenwu V900 AI chip, claiming it is the most powerful in China
Read more
tecnoblog.net

Alibaba launches Zhenwu V900 AI chip, claiming it is the most powerful in China

Alibaba officially announced the Zhenwu V900, its latest processor designed for training and inference tasks of artificial intelligence models. According to the company, this accelerator demonstrates three times the performance of the Zhenwu M890, its previous version, and features 216 GB of memory.

This launch took place during the Apsara 2026 conference, held in Hangzhou, China. Eddie Wu, CEO of Alibaba, stated that the V900 is currently the 'most powerful AI chip in China.' However, the company did not provide sufficient data to validate this claim or to conduct a performance comparison with its competitors' chips.

The Zhenwu V900 was designed to meet two primary demands of AI: model training and inference, which is the phase where a trained model executes the necessary calculations to produce responses. In addition to 216 GB of memory, the model offers a bandwidth of 1,200 GB/s in inter-chip communication; for comparison, the M890 has 144 GB of memory and 800 GB/s inter-chip communication.

The T-Head semiconductor division reported that more than a thousand units can operate together as a single system. On an even larger scale, Alibaba stated that the accelerator has the potential to integrate clusters containing up to 500,000 boards.

Mass production and commercial release of the device are scheduled for the first quarter of 2027. Despite the promises of great capacity, crucial information is still missing for a complete evaluation of the V900. Alibaba also omitted details on FLOPS performance, manufacturing method, the company responsible for component production, or its energy consumption, leaving the claim of being the most powerful chip in China as a statement from the manufacturer itself.

Additionally, this new accelerator integrates Alibaba's plans to develop larger AI models. The company plans to create a future Qwen with a scale of 5 to 10 trillion parameters, while the current Qwen3.8-Max has 2.4 trillion.

This model expansion aligns with the Chinese giant's projects aimed at increasing its computing infrastructure. Alibaba aims to exceed 20 gigawatts of capacity in its data centers by 2032. However, Wu mentioned that shortages in the global supply chain for AI data centers are restricting the speed of this expansion.

Popular