SpaceXAI doubles speech recognition accuracy in new Grok Voice Transcribe 2.0 model

SpaceXAI doubles speech recognition accuracy in new Grok Voice Transcribe 2.0 model

SpaceXAI has released a new model, Grok Voice Transcribe 2.0, for converting speech into text. According to the developers, it has half as many bugs as previous versions for the same cost. This model is designed for complex recordings, including static phone conversations, multiple people speaking simultaneously, and regional accents and dialects.

    Image source: Grok

Image source: Grok

According to reports, Grok Voice Transcribe 2.0 Mark Technology Postbuilt based on the audio model used in Grok Voice. The technology already handles tens of thousands of support calls every day, transcribes millions of hours of video footage, and is used in the voice assistant in Tesla vehicles. Train using multilingual recordings with noise and interference, then use post-training methods to further refine the model.

In the Artificial Analysis public ranking, Grok Voice Transcribe 2.0 ranked first among 32 streaming media models in terms of AA-WER Streaming. The company also tested it on four internal data sets: phone conversations; conversations with Grok; spoken phone numbers, regular addresses, and email addresses; and voice assistant phrases in 19 languages. In all four cases, version 2.0 performed better than version 1.0, but internal data is provided by SpaceXAI itself and has not been independently verified.

Developers say multi-language recognition is one of the major improvements. The model automatically detects language, supports dozens of languages, and is able to continue transcribing even if the language is changed directly during the conversation. On a set of short phrases, the WER indicator, which reflects the percentage of single-word recognition errors, dropped from 20.6% to 6.8%, which is a drop of approximately 67%. The document also states that automatic formatting of numbers, monetary amounts and units of measurement is supported in 25 languages.

This model is available in batch and streaming modes via the Speech to Text API. In streaming mode, the Opus audio codec supports data streams of approximately 4 KB/s. In comparison, uncompressed PCM audio at a 24 kHz sampling rate requires approximately 48 KB/s. The API supports timestamps and confidence scores for identifying each word, speech separation between different interlocutors (at no additional cost), up to 8 audio channels, and the ability to define up to 100 key terms. There’s also automatic formatting of numbers and dates, removal of filler words, and line-ending determination for voice assistants.

In batch mode, you can upload files in 12 audio formats up to 500 MB and stream them through WebSocket. In the API, the model identification code is grok-voice-transcribe-2.0. Mass transcription costs $0.10 per hour of audio and $0.20 per hour for streaming.

If you find an error, select it with your mouse and press CTRL+ENTER.

Exit mobile version