Xiaomi released and open-sourced Xiaomi-CocktailASR-1, an industrial-grade target speaker speech recognition large model designed to address the “cocktail party” problem. The model uses an end-to-end LLM architecture and a reference audio clip of the target speaker as a voiceprint prompt, enabling it to extract and transcribe only the target user’s speech in environments where multiple people are speaking simultaneously.