Google’s newly released open weights model, Gemma 4 12B, supports transcription but is far from the frontier, scoring 8.8% on AA-WER (#58) Gemma 4 12B is the latest release from @GoogleDeepMind in the Gemma 4 family. With a score of 8.8% on AA-WER, it is able to capture a reasonable amount of conversation context, but underperforms compared to transcription-focused open weights models like Voxtral Mini Transcribe 2 (3.6% WER, with 4B parameters) and slightly larger open weights language models like Voxtral Small (2.8% WER, with 12B parameters). The new model launched alongside their local dictation app, Eloquent, available on MacOS and iOS. Gemma 4 12B is the largest in the Gemma 4 family to support transcription, alongside Gemma 4 E4B and Gemma 4 E2B, with Gemma 4 31B and Gemma 4 26B A4B supporting text, image and video input only. These models are available on a variety of platforms including Hugging Face, Ollama and LMStudio. We are currently running Gemma 4 12B through the full Artificial Analysis Intelligence Index and will share results soon.
Artificial Analysis Benchmarks Google's Gemma 4 12B Transcription at 8.8% WER
Artificial Analysis· Updated
Artificial Analysis benchmarked Google DeepMind's new open-weight Gemma 4 12B model for transcription, reporting an 8.8% Word Error Rate (WER). This places the model behind specialized open-weight transcription solutions, but it is available for local deployment alongside Google's new Eloquent dictation app.
- AA-WER (Overall)
- 8.8% (#58)
- Comparison (Voxtral Mini Transcribe 2)
- 3.6% WER (4B parameters)
- Comparison (Voxtral Small)
- 2.8% WER (12B parameters)
- Associated App
- Eloquent (MacOS, iOS)
- Availability Platforms
- Hugging Face, Ollama, LMStudio
Performance lags behind dedicated open-weight solutions. Mistral's Voxtral Mini Transcribe 2 (4B parameters) scores 3.6% WER, while Voxtral Small (12B parameters) reaches 2.8% WER. This gap highlights the trade-offs between general-purpose multimodal models and those optimized for single-task speech-to-text accuracy.
Gemma 4 12B is available on Hugging Face, Ollama, and LMStudio for local deployment. This supports Google's push for Gemma 4 models in efficient on-device AI. It launched alongside Eloquent, a local dictation app for MacOS and iOS, enabling users to run transcription workflows entirely offline.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →





