Xiaomi launches CocktailASR-1, an open artificial intelligence model. It aims to solve one of the most difficult problems in audio transcription – isolating a specific person’s voice when multiple people are speaking at the same time.
Image source: Xiaomi
At Xiaomi, this problem is often referred to as the “cocktail party effect.” Most specialized AI models can handle speech recognition when a person is speaking, but things get more complicated when multiple voices overlap. Texts are distorted, sentences are merged, and remarks are attributed to the wrong interlocutors. Instead, the company solved this problem with its OmniVoice model – the CocktailASR-1 doesn’t generate speech, but selects and transcribes it. To use the model, you need to submit a short sample containing the desired person’s voice. She used it as a reference, and then after analyzing several participants’ recordings, she identified and transcribed only that person’s speech.
Xiaomi noted that in multiple speech recognition tests with multiple speech samples, it demonstrated leading results and outperformed existing similar products designed to solve the same problem. Its capabilities aren’t limited to complex conditions – it can decipher speech even if there’s only one person’s voice in the recording. It also knows how to avoid making unfounded assumptions: if the sample speech provided doesn’t appear in the suggested recording, it generates blank text and makes no attempt to decipher the other person’s words.
Finally, the Xiaomi CocktailASR-1 features a chain inference mode – which allows users to study the logic that performs decryption rather than just seeing the final text. New models available at GitHub and Face hugging.
If you find an error, select it with your mouse and press CTRL+ENTER.










