The Index LLM team at Bilibili has introduced Index-Translate—a family of multilingual translation models developed based on Alibaba's Qwen3.5. These text models support 150 languages, including Chinese and English, and are available under the Apache-2.0 license on Hugging Face and ModelScope platforms. They are released in sizes of 2B, 9B, and 35B-A3B (preliminary version), and are accompanied by a technical report, GitHub code, and an online demo.
The models' functionality goes beyond simple translation; they are capable of following instructions regarding terminology, formatting, and preserving unchanged content. For instance, in one example, the 9B model translated a game service notification in JSON format into Korean while maintaining its structure, stars, and hashtag as requested by the user.
The model family is expanded with three specialized branches. Index-Echo generates translated subtitles or voiceovers while preserving the characteristics of the original speaker's voice. Its batch speech-to-speech conversion package covers translations from Chinese to English, Spanish, and Japanese, as well as from English to Chinese, Spanish, and Japanese. The Index-Homura model customizes translation according to the target syllable count, which is critical for dubbing and subtitles. And Index-NativeLong, available under the identifier Index-Nailong, handles the translation of entire documents, ensuring consistency of references between them.
Examples presented by Bilibili are based on content from its own community. When the 9B model processed a gamer's slang expression about wanting to play Final Fantasy XIV, it translated it as 'When the FFXIV itch hits, just go play.'. In a fictional text of about 32,000 tokens, where a character's name could have been interpreted as a royal title, the long document model preserved this name consistently, whereas the same 9B model, working in chunks, allowed for other translation variations.
According to the team's internal benchmarks, the preliminary 35B-A3B version achieved a score of 0.8794 on FLORES (COMET-22) and 76.76 on the WMT26 judge evaluation. The 9B model reached scores of 0.8789 and 75.35. It is worth noting that some large general models demonstrate higher results on WMT26, including DeepSeek-V4.1-Flash with a score of 83.55. When handling low-resource instruction-following tasks, the 9B model showed the best instruction score—0.7725—and the lowest deviation from the target—3.47%—among all compared models. The Index-Homura-9B model falls within the specified syllable count within 10% in 81.92% of cases according to the team's SandGlass test.
The team has also developed proprietary benchmarks to monitor instruction following, gaming slang, and syllable control. The repository includes a browser extension that allows translating web pages using a locally deployed model, as well as a pipeline for video dubbing. Bilibili plans to release the official 35B-A3B model, open-source its benchmarks, add languages to Index-Echo, and publish larger models.
