Huawei's openJiuwen platform provides WorkSwarm with a full-duplex voice and video interface for operation on Ascend NPUs
Read more
Pandaily
pandaily.com

Huawei's openJiuwen platform provides WorkSwarm with a full-duplex voice and video interface for operation on Ascend NPUs

OpenJiuwen, an open-source platform for AI agents developed jointly by Huawei Labs 2012, Huawei Cloud, the company's device and computing teams, as well as universities and other developers, has integrated full-duplex multimodal interaction into WorkSwarm—its open desktop for office and coding tasks.

The update, announced on October 1st, allows users to communicate with the agent just like in a regular phone call while another agent continues to perform longer tasks in the background.

Principle of Full-Duplex Interaction

Full-duplex communication means that the assistant can listen and speak simultaneously. Unlike traditional voice assistants that operate like walkie-talkies, where the user must wait for the AI's response to finish, the user can now interrupt the response mid-way, for example, by asking it to skip the background and proceed to the conclusions.

When WorkSwarm detects a new actionable voice command, it stops the current response and processes the new instruction; the previously generated text is saved in the task record. To minimize false interruptions from background noise, voice activity detection and noise suppression features are used.

Task Separation Architecture

The design separates the work into two streams. The real-time model monitors the video stream, reacts quickly, and responds, while the Main Agent handles task decomposition, tool invocation, and advancing the process. A user can point the camera at a product, ask about its brand, and then instruct WorkSwarm to conduct research and compile a document. The search runs in the background while the conversation continues, and the interface displays the task status, tool call logs, received results, and created files.

After the background task is completed, WorkSwarm first publishes a written summary and waits for the current verbal dialogue to end before vocalizing only the key points. Subsequent requests are queued, and users can change the order of tasks, move one to the front, or stop pending or running tasks. Interrupting the conversation does not cancel instrumental tasks, which must be stopped separately, and closing the audio and video session does not halt the work already handed over to the Main Agent. Results are returned to the original dialogue.

Currently, WorkSwarm supports two model protocols: Qwen Omni in real-time via WebSocket, which can use an official endpoint or local deployment, and JoyAI, which connects separate speech recognition, language, and speech synthesis models via official APIs or other platforms. Multimodal models with full-duplex connectivity can also be deployed on Ascend NPUs through Huawei Cloud's ModelArts platform using the vLLM-Omni inference framework. WorkSwarm offers installers for Windows, Mac, and HarmonyOS, as well as pip packages, and its GitHub repository is licensed under Apache-2.0. The latest tagged release of the project, v0.2.7 from September 30th, added Proactive Context Service and inter-agent sessions.

Popular