Mistral AI unveils Voxtral, an innovative suite of open source speech understanding models designed to transform human-machine interaction through voice. Considering voice as humanity's original interface, Voxtral aims to overcome the limitations of current systems, whether unreliable or proprietary, by offering robust, multilingual, and deeply intelligent voice tools. The models are available in two versions: a 24B variant built for production-scale applications, and a more compact 3B variant, Voxtral Mini, suited to local and edge deployments. Both are released under the permissive Apache 2.0 license and accessible via the Mistral AI API, with Voxtral Mini Transcribe, optimized for transcription, offering unmatched cost/latency efficiency.
Two-tier architecture and cost advantage
Voxtral stands out by bridging the gap between open source ASR systems with high error rates and costly proprietary APIs. It delivers state-of-the-art accuracy and native semantic understanding within an open framework, at less than half the price of comparable proprietary solutions. This cost efficiency makes high-quality voice intelligence accessible and controllable at scale for a wide range of applications.
Advanced capabilities beyond transcription
The models offer several advanced capabilities beyond simple transcription. They support long audio contexts, up to 30 minutes for transcription and 40 minutes for understanding, enabling full processing of extended conversations or recordings. A standout feature: built-in Q&A and summarization, which allow direct querying of audio content or generation of structured summaries without chaining separate ASR and language models. Voxtral is natively multilingual, with automatic language detection and state-of-the-art performance across many widely used languages: English, Spanish, French, Portuguese, Hindi, German, Dutch, Italian. It also enables function-calling directly from voice, translating users' spoken intentions into actionable system commands. The models retain the strong text understanding capabilities of their foundation, the Mistral Small 3.1 language model.
Competitive benchmarks
Benchmark results underscore Voxtral's superior performance. It outperforms Whisper large-v3 overall, the reference open source transcription model, and surpasses GPT-4o mini Transcribe and Gemini 2.5 Flash on various tasks. Voxtral Small achieves state-of-the-art results on short-form English and Mozilla Common Voice, demonstrating strong multilingual capabilities, and matches ElevenLabs Scribe for premium use cases at a significantly reduced cost. Voxtral Mini Transcribe outperforms OpenAI Whisper at less than half the price.
Roadmap and vision
Mistral AI plans to introduce speaker diarization, audio annotations (age, emotion), word-level timestamps, and non-speech audio recognition. The company is expanding its audio team and encourages developers to integrate Voxtral via local download on Hugging Face, via the API, or by trying it in Le Chat's voice mode, with advanced enterprise features: private deployment, domain-specific fine-tuning, and dedicated integration support.