Alibaba's AI team Qwen released Qwen-Audio-3.1 on Tuesday. The release includes a lineup of five models designed for speech recognition, text-to-speech, and real-time interaction.
The automatic speech recognition model improves multilingual and dialect recognition. It also automatically cleans up filler words and repetitions.
A second model named ASR-Next adds multi-speaker identification with timestamps. This version also detects emotions, ambient sounds, and machine noise.
The text-to-speech model handles multilingual synthesis. It includes natural cross-language voice transfer capabilities.
Users can control emotion, speed, and style through simple text prompts. An example provided by the team is reading a text with a sharp, commanding tone that demands respect.
TTS-Next pairs a language model with a diffusion approach. This combination generates voice, sound effects, and background audio in a single pass.
The real-time model supports simultaneous speaking and listening with instant interruption. Qwen states that the model responds more slowly and with more empathy when it detects a low mood.
Alongside the model release, Alibaba is slashing prices significantly. Text-to-speech prices drop by about 70 percent.
Real-time model prices drop roughly 85 percent. Speech recognition prices face the steepest reduction at up to 95 percent.
More details are available on the official blog and on Qwen Cloud.



