Alibaba's Tongyi Qianwen has released the Qwen-Audio-3.1 series, introducing five new voice models that form a complete audio stack covering understanding, generation, interaction, and creation. The upgrade includes evolved core models for speech recognition, synthesis, and real-time interaction, alongside new TTS-Next and ASR-Next variants. Prices across the entire Qwen-Audio lineup have been significantly reduced to lower user costs. Text-to-speech pricing dropped by approximately 70%, real-time voice interaction fees fell by about 85%, and automatic speech recognition costs were slashed by as much as 95%.