tradingkey.logo

Alibaba unveils new flagship AI model: Qwen2.5-Omni

Investing.comMar 27, 2025 12:27 AM

Investing.com -- Alibaba Group Holdings Ltd ADR (NYSE:BABA) has introduced Qwen2.5-Omni, its new flagship model in the Qwen series. The end-to-end multimodal model is designed for extensive multimodal perception and can process a variety of inputs such as text, images, audio, and video. It provides real-time streaming responses through text generation and natural speech synthesis.

Key features of the model include its Thinker-Talker architecture, designed to perceive a range of modalities, including text, images, audio, and video. This architecture allows the model to generate text and natural speech responses simultaneously. It also includes a novel position embedding, dubbed TMRoPE (Time-aligned Multimodal RoPE), which synchronizes the timestamps of video inputs with audio.

The model is designed for fully real-time interactions, supporting chunked input and immediate output. It surpasses many existing streaming and non-streaming alternatives in terms of robustness and naturalness in speech generation. Qwen2.5-Omni showcases exceptional performance across all modalities and outperforms the similarly sized Qwen2-Audio in audio capabilities. It also matches the performance of Qwen2.5-VL-7B.

Qwen2.5-Omni employs the Thinker-Talker architecture, where the Thinker functions like a brain, processing and understanding inputs from text, audio, and video modalities. It generates high-level representations and corresponding text. The Talker operates like a human mouth, taking in the high-level representations and text produced by the Thinker and outputting discrete tokens of speech fluidly.

A comprehensive evaluation of Qwen2.5-Omni has been conducted, showing strong performance across all modalities when compared to similarly sized single-modality models and closed-source models like Qwen2.5-VL-7B, Qwen2-Audio, and Gemini-1.5-pro. In tasks requiring the integration of multiple modalities, such as OmniBench, Qwen2.5-Omni achieves state-of-the-art performance.

In the near future, Alibaba plans to enhance the model's ability to follow voice commands and improve audio-visual collaborative understanding. The company also aims to integrate more modalities towards an omni-model.

The Qwen2.5-Omni model is now publicly available on platforms like Hugging Face, ModelScope, DashScope, and GitHub. Users can experience the model's interactive features through a demo or join discussions on Discord.

Disclaimer: The information provided on this website is for educational and informational purposes only and should not be considered financial or investment advice.

Related Articles

tradingkey.logo
tradingkey.logo
Intraday Data provided by Refinitiv and subject to terms of use. Historical and current end-of-day data provided by Refinitiv. All quotes are in local exchange time. Real-time last sale data for U.S. stock quotes reflect trades reported through Nasdaq only. Intraday data delayed at least 15 minutes or per exchange requirements.
* References, analysis, and trading strategies are provided by the third-party provider, Trading Central, and the point of view is based on the independent assessment and judgement of the analyst, without considering the investment objectives and financial situation of the investors.
Risk Warning: Our Website and Mobile App provides only general information on certain investment products. Finsights does not provide, and the provision of such information must not be construed as Finsights providing, financial advice or recommendation for any investment product.
Investment products are subject to significant investment risks, including the possible loss of the principal amount invested and may not be suitable for everyone. Past performance of investment products is not indicative of their future performance.
Finsights may allow third party advertisers or affiliates to place or deliver advertisements on our Website or Mobile App or any part thereof and may be compensated by them based on your interaction with the advertisements.
© Copyright: FINSIGHTS MEDIA PTE. LTD. All Rights Reserved.