Artificial Analysis Benchmarks Google's Gemma 4 12B Transcription at 8.8% WER

Artificial AnalysisArtificial Analysis

· Updated

Artificial Analysis benchmarked Google DeepMind's new open-weight Gemma 4 12B model for transcription, reporting an 8.8% Word Error Rate (WER). This places the model behind specialized open-weight transcription solutions, but it is available for local deployment alongside Google's new Eloquent dictation app.

Artificial Analysis benchmarked Google DeepMind's Gemma 4 12B, an open-weight model supporting transcription. It scored 8.8% AA-WER, ranking #58 on the leaderboard. While Gemma 4 12B is the largest in its family to support transcription, other variants like 31B and 26B A4B are limited to text, image, and video. It captures reasonable context but trails specialized transcription models.
AA-WER (Overall)
8.8% (#58)
Comparison (Voxtral Mini Transcribe 2)
3.6% WER (4B parameters)
Comparison (Voxtral Small)
2.8% WER (12B parameters)
Associated App
Eloquent (MacOS, iOS)
Availability Platforms
Hugging Face, Ollama, LMStudio

Performance lags behind dedicated open-weight solutions. Mistral's Voxtral Mini Transcribe 2 (4B parameters) scores 3.6% WER, while Voxtral Small (12B parameters) reaches 2.8% WER. This gap highlights the trade-offs between general-purpose multimodal models and those optimized for single-task speech-to-text accuracy.

Gemma 4 12B is available on Hugging Face, Ollama, and LMStudio for local deployment. This supports Google's push for Gemma 4 models in efficient on-device AI. It launched alongside Eloquent, a local dictation app for MacOS and iOS, enabling users to run transcription workflows entirely offline.

Artificial Analysis
Artificial Analysis
@ArtificialAnlys
X

Google’s newly released open weights model, Gemma 4 12B, supports transcription but is far from the frontier, scoring 8.8% on AA-WER (#58) Gemma 4 12B is the latest release from @GoogleDeepMind in the Gemma 4 family. With a score of 8.8% on AA-WER, it is able to capture a reasonable amount of conversation context, but underperforms compared to transcription-focused open weights models like Voxtral Mini Transcribe 2 (3.6% WER, with 4B parameters) and slightly larger open weights language models like Voxtral Small (2.8% WER, with 12B parameters). The new model launched alongside their local dictation app, Eloquent, available on MacOS and iOS. Gemma 4 12B is the largest in the Gemma 4 family to support transcription, alongside Gemma 4 E4B and Gemma 4 E2B, with Gemma 4 31B and Gemma 4 26B A4B supporting text, image and video input only. These models are available on a variety of platforms including Hugging Face, Ollama and LMStudio. We are currently running Gemma 4 12B through the full Artificial Analysis Intelligence Index and will share results soon.

9retweets125likes
View on X

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update