Xiaomi released and open-sourced Xiaomi-CocktailASR-1, an industrial-grade target speaker speech recognition large model designed to address the “cocktail party” problem.
The model uses an end-to-end LLM architecture and a reference audio clip of the target speaker as a voiceprint prompt, enabling it to extract and transcribe only the target user’s speech in environments where multiple people are speaking simultaneously.
Xiaomi Open-Sources CocktailASR-1 Speech Recognition Model
Disclaimer: The content provided on Phemex News is for informational purposes only. We do not guarantee the quality, accuracy, or completeness of the information sourced from third-party articles. The content on this page does not constitute financial or investment advice. We strongly encourage you to conduct you own research and consult with a qualified financial advisor before making any investment decisions.
