Meituan releases LongCat-2.5-Preview model: multimodal agent with 1.6 trillion parameters and 1 million token context
Read more
Pandaily
pandaily.com

Meituan releases LongCat-2.5-Preview model: multimodal agent with 1.6 trillion parameters and 1 million token context

Meituan introduced LongCat-2.5-Preview on its LongCat API platform on September 25, 2026. According to information from iFeng Tech, AI Tool Lab, and related industry publications, this release is positioned as a long-term multimodal agent model, rather than just an update to chat functions. The official English name for the product is Meituan / LongCat.

This review presents the technical specifications of the Preview product and the intended use of the agent software. Meituan has not published official benchmarks for this release, so claims regarding capabilities should be viewed with caution.

Architecturally, the model retains the Mixture-of-Experts approach used in LongCat-2.0. The total parameter count is approximately 1.6 trillion, with about 48 billion parameters activated at each inference step. The model features a built-in context window of one million tokens, making it suitable for processing long documents, repositories, logs, and multi-turn tool cycles.

The main difference from version 2.0 is the integration of native multimodal capability directly into the base model. This includes image parsing for cross-modal question answering, as well as visual reasoning and summarization. Furthermore, there has been a clear shift from purely agentic coding to autonomous workflow management across terminals, browsers, user interfaces, spreadsheets, and design tools.

Access to the model is provided in two formats: via API (with endpoints compatible with OpenAI and Anthropic, as described in reviews) and through a web interface at longcat.chat. The maximum output text length is set at 128 thousand tokens.

Integration notes mention tools such as Codex, OpenCode, OpenClaw, CatPaw, Claude Code, Hermes, and Kilo Code, allowing existing agent systems to directly access the Preview version. Secondary reports cite promotional pricing indicating limited-time input usage at around $0.30 and output usage at around $1.20 per million tokens, with an additional tiered input level and free tokens for current users; commercial terms should be considered as temporary Preview offers.

Since LongCat-2.5-Preview does not come with a public rating in this cycle, readers are advised to view the improvements in agent software as a positioning claim awaiting independent evaluation. For developers tracking multimodal agents in China, the key takeaway is that Meituan has presented a 1.6T / ~48B MoE model with native 1M context, focused on long sequences of software operations rather than financial history or ranking changes.

Similar stories

StepFun releases preliminary version of Step 5 model: agent with 600 billion parameters and open weights expected on October 15
Read more
pandaily.com

StepFun releases preliminary version of Step 5 model: agent with 600 billion parameters and open weights expected on October 15

StepFun introduced the preliminary version of Step 5 on September 20, 2026, as its next flagship foundational model designed for long-horizon agents, according to information from Tencent Tech, DataLearner model cards, and company materials available at stepfun.com. The official English name of the model is StepFun / Step 5.

It is important to note that although the model identifier step-5-preview is already available via the product API and the open StepFun platform, the open weights in BF16 format are scheduled only for October 15, 2026, and were not available at launch. Therefore, availability via API and open weights should be considered as separate points.

The model architecture features a Mixture of Experts (MoE) sparse design with an approximate total of 600 billion parameters. Approximately 27 billion parameters are activated per token within a 92-layer 'narrow and deep' Transformer. The choice of depth is driven by agent requirements, as longer information paths through layers contribute to improved multi-step implicit reasoning and handling of long tool outputs.

The model supports a one-million-token context window, accepting both text and image input (though video input is mentioned on third-party cards). It is oriented towards applications in AI coding, software development, financial analysis, and professional agent usage. To maintain practicality with the million-token attention, sparse GQA was applied in combination with token block merging, reducing the cost of indexer and top-k selection by approximately one eighth.

During training, special attention was paid to bit-level alignment between training and inference to ensure stable MoE routing, and load-aware scheduling, speculative decoding, and FP8 paths were utilized. StepFun claims that these methods accelerate the long-horizon Reinforcement Learning (RL) process by more than three times overall.

According to data from Artificial Analysis's composite AI index, it scores 44 points. The company places this rating among leading models focused on open access. Aggregator cards also list API prices: about $1.00 per input token and $2.70 per output token per million tokens, but comparisons of metrics and cost should be viewed as statements from third parties or providers.

Until October 15, Step 5 Preview should primarily be regarded as a functional API for long-horizon agents, having only a planned commitment for open weights, rather than a release of weights on Hugging Face or a drop-in replacement for Meituan LongCat, MiniMax Code Flash, or NaiveAI OSS MoE.

StepFun releases preliminary version of Step 5: a model with 600 billion parameters, sparse MoE, and 1 million token context
Read more
pandaily.com

StepFun releases preliminary version of Step 5: a model with 600 billion parameters, sparse MoE, and 1 million token context

StepFun has introduced the preliminary version of the Step 5 model, which is a base model featuring a sparse Mixture-of-Experts (MoE) architecture and approximately 600 billion total parameters, with about 27 billion activated per token. This release targets long-horizon agent workloads, including AI-assisted code development, software engineering, financial analysis, and professional knowledge work. The model supports a one-million-token context window and accepts both text and image inputs. StepFun announced that API access is already open, and the model weights are scheduled to be made publicly available on October 15th.

Instead of expanding the network, Step 5 Preview utilizes a narrow, deep Transformer architecture consisting of 92 layers. The company asserts that deeper stacks provide longer information pathways for implicit multi-step reasoning during long prefill, when agents perform searches, run code, and process tool usage results. To maintain practicality in the million-token sessions, the model incorporates Sparse Grouped-Query Attention with token block merging. According to StepFun, this reduces the cost of indexer selection and top-k by roughly one-eighth compared to a denser base model while consolidating overlapping adjacent selections.

The training process emphasizes on-policy long-horizon reinforcement learning, as well as bit-level alignment of training and inference in MoE routing. Furthermore, techniques such as load-aware scheduling, MTP-3 speculative decoding, FP8 MoE, and KV cache offloading are employed. StepFun reports more than a threefold acceleration of the end-to-end RL process for the long horizon and a sample registry loss metric below one percent. This model is also used within a human-managed data pipeline that generates verifiable complex tasks at scale in the millions, covering science, software development, and machine learning research.

Regarding AI analysis, Step 5 Preview scores around 44 points on an intelligence index that StepFun ranks among the best open-weights models in this rating system. The API cost, according to published pricing, is approximately $1 per million input tokens and $2.7 per million output tokens, with an output speed of about 100 tokens per second. The model's primary focus is on agent architecture and efficiency, distinguishing it from previous releases like Step 3.5 Flash and Step 3.7 Flash, emphasizing depth, sparse long context, and reliable multi-stage tool use before the weights become available in October.

TaichuAI releases open multimodal model ZDTaichu5.0-9B for spatial understanding
Read more
pandaily.com

TaichuAI releases open multimodal model ZDTaichu5.0-9B for spatial understanding

TaichuAI has made the ZDTaichu5.0-9B model publicly available. This multimodal foundation model has approximately 9 billion parameters and is designed for general visual understanding, spatial reasoning, agent tool usage, and embodied AI research.

The model's architecture combines the Qwen3.5-9B language backbone with the C-RADIOv4-H vision encoder. The model accepts text input, as well as one or more images and videos of any resolution, supporting a context length of up to 128 thousand tokens. The release, summarized by TMTPost on September 15th, focused on real-world spatial perception, transformations between different views, and task planning for embodied systems.

According to the public model card, TaichuAI demonstrates state-of-the-art results among approximately 10-billion general VLMs across a wide range of visual tasks, while also expanding its capabilities in spatial, embodied, and agentic scenarios. Noteworthy metrics in space and embodiment include ViewSpatial 62.50, MMSI-Bench 47.20, MindCube-tiny 78.27, ERQA 48.00, and RoboSpatial 56.00.

Regarding agent and instruction sets listed in the evaluation card notes, TAU2-Bench achieves a score of 87.70, and the average Claw-Eval score is 71.40, while IFEval shows 93.70. The entropy-prioritized adaptive recursive reasoning mechanism is described as a mechanism that allocates additional refinement steps for latent variables for more complex tokens.

The model's capabilities cover Optical Character Recognition (OCR) and document understanding, math on images, fine-grained 2D relationships, multi-view association, 3D scene and perspective perception, as well as tracking multiple images and videos within a long context window, including multi-step tool use planning. Actual tool execution remains the responsibility of the host application.

The spatial learning topics listed on the card include relative relationships, dense counting and framing, camera motion and depth ordering, egocentric versus allocentric views, and high-level capability definition and action planning for VLA-style adaptation.

The model weights are hosted on Hugging Face at TaichuAI/ZDTaichu5.0-9B under the NVIDIA Open Model License, while retaining the Apache-2.0 notices from Qwen3.5. A custom branch of vLLM 0.26.0 and a Docker image from TaichuAI are used for serving to provide OpenAI-compatible endpoints, with recommended sampling settings for spatial grounding and general tasks.

As with other vendor-released models, external laboratories should view benchmark scores as reported with the specified prompts and judges until independent reproduction results emerge. However, the combination of a mid-sized open multimodal checkpoint, a focus on space and agents, and ready vLLM packaging provides researchers with a concrete artifact for evaluation.

Popular