i
News
News · 2026-09-23

Alibaba cuts Qwen Audio prices by up to 95%

@neuronium_ai @neuronium_ai

Alibaba has released Qwen Audio 3.1, a five-model package covering speech recognition, speech synthesis and real-time conversation. The release combines broader audio handling with price cuts of up to 95%, making the economics at least as notable as the feature list. For developers building voice interfaces, lower inference costs can matter more than another incremental improvement in model quality.

Cover: Alibaba cuts Qwen Audio prices by up to 95%

What shipped

The recognition model is designed for multiple languages and dialects. It removes filler words and repetitions automatically, while ASR-Next adds:

speaker identification;
timestamps;
emotion recognition;
background-sound detection;
detection of noise from operating machinery.

The synthesis model supports multilingual generation and natural voice transfer between languages. Developers can control emotion, speed and style with plain-text prompts. For example, this prompt asks the model to use a sharp, commanding delivery:

Read this with a sharp, commanding tone, demanding respect.

TTS-Next combines a language model with a diffusion method. In a single pass, it generates speech, sound effects and background audio.

The real-time model can listen and speak at the same time, responding immediately when a user interrupts. Qwen says that when it detects a subdued mood, the system responds more slowly and with greater empathy.

The price cut is the real product story

Alibaba has reduced its prices across the three areas:

70%speech synthesis cut
85%real-time mode cut
95%maximum recognition cut

Those percentages change how the release should be read. The feature set is broad, but the commercial pitch is unusually direct: voice generation, recognition and live interaction are being positioned as cheaper building blocks rather than isolated demonstrations.

My read is that the 95% figure will draw more attention than any single audio capability. A model that identifies speakers or detects machinery noise is useful, but lower operating costs determine whether teams can place those features throughout a product instead of reserving them for a narrow workflow.

What the announcement leaves out

The percentages are striking, but they do not show the absolute prices behind them. That makes it difficult to judge whether the reduction changes the practical cost of running a voice product or simply makes an already inexpensive service cheaper.

What I’d want to know is how these models perform across the languages, dialects and noisy environments they are meant to handle. The release describes the functions, including emotional adaptation in real time, but gives no comparative results. That leaves the central trade-off unresolved: Alibaba has made audio AI cheaper on paper, while the evidence for how reliably it handles real conversations remains outside the announcement.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X