The most interesting point in the comments
Yuan✴ Launched the first artificial intelligence model designed specifically for real-time audio processing. Muse Voice Transcribe can transcribe speech for more than 20 speakers and handle multiple languages simultaneously.
Photo credit: Kate Bezzubets/unsplash.com
In the demo video, the model not only supports multiple speaker speech recognition and handles multiple languages - it can decipher phrases when a person switches from one language to another in the middle of a sentence. “The model decides when to listen. It waits slightly longer for complex words and reacts faster for simple words, using adaptive delays to predict each token and improve accuracy. It also handles low-quality real-world audio well – it was trained on 70 languages (of which 25 were tested at launch), handles mid-sentence language switching, and handles conversations between speakers and over 20 sessions.”” said Mark Zuckerberg, head of the company.
Last week, Google launched Gemini 3.5 Transcribe, which has similar functionality and will be built into Android and Chrome. About meta-intention✴ There has been no mention of integrating its development into its own flagship service. You can use Muse Voice Transcribe in the Meta app✴ Artificial intelligence for Apple macOS; on Muse Code and Meta platforms✴ Model API developers can connect to it for $3 per 1000 minutes of audio.
Tell Google to receive more frequent links to our news about artificial intelligence
If you find an error, select it with your mouse and press CTRL+ENTER.









