Shanghai AI Lab releases Intern-Decision: compact models that output decisions instead of text
Read more
Pandaily
pandaily.com

Shanghai AI Lab releases Intern-Decision: compact models that output decisions instead of text

The Shanghai Artificial Intelligence Laboratory (Shanghai AI Lab) has made Intern-Decision available—a family of compact decision-making models. These models were released on September 28th in three size variants: 0.8B, 2B, and 4B parameters. Instead of generating long textual excerpts, Intern-Decision is designed to directly provide structured judgments along with probabilities for each possible answer.

The design concept is based on the principle of 'decision, not string.' Intern-Decision supports three types of questions. A choice question requires the model to select the best option from a predefined list and assign a probability to each option. A rating question requests a score on a user-defined scale. And a 'yes/no' question returns a binary answer with the corresponding probability. Furthermore, a single query can combine multiple types of questions about the same state, and all are processed in a single forward pass.

Integrating Visual Perception

According to the project repository on GitHub, the models are fine-tuned on the Qwen3.5 language core, while the vision tower and projector remain frozen. This allows Intern-Decision to accept images as part of the input data, meaning visual perception directly influences the decision-making stage. The team demonstrated this capability on tasks related to games and interfaces: the model analyzes raw game frames, including two-dimensional grid maps and side-scrolling levels, and independently selects the next action. Other demonstrations cover mouse control and browser usage.

Testing Results and Performance

The laboratory reports that the 4B model achieves an average accuracy of 90.02% across seven test sets, which is 1.28 percentage points higher than the 88.74% achieved by the commercial decision-making model Jev, which is used as the primary benchmark. These test sets include decision types, tool calling, news classification, and intrusion detection, totaling over 10,000 test strings. The repository also indicates an average local query latency of about 44 milliseconds when using a single consumer RTX 4090 GPU, and provides calibration results on a 96-case distribution benchmark. It should be noted that these figures are self-reported by the developers, and the repository emphasizes that protocols and evaluation scopes differ between models.

The release includes training code, two inference backends, tools for calculating benchmark results, temperature calibration presets, and a browser demo version; however, training data and some closed validation records are not included. Model weights are published on Hugging Face.

Local GPU manufacturer MetaX stated that it has completed the Day-0 adaptation, so Intern-Decision has been running on its hardware since launch. MetaX claims that starting in December 2025, it will provide Day-0 support for 40 major flagship models, and its MXMACA software stack supports over 40 AI frameworks and more than 1000 models.

For developers creating agents, routers, and content filters, small models that output calibrated probabilities instead of free text can make automated decisions more cost-effective, faster, and easier to audit.

Popular