Google Gemma

Google Gemma — Latest Updates & Releases20 Updates

The latest AI news and updates of Google Gemma — Google's family of lightweight, open-weight AI models. Covering Google Gemma's latest product updates, launches, and research from the past 90 days.

Google GemmaGoogle GemmaAug 22

Google Gemma 4 31B Matches Sonnet 5 Performance at 40x Lower Cost

Google reports that its Gemma 4 31B model matches Claude Sonnet 5 on answer quality in AlphaSense benchmarks, while operating at roughly 40x lower cost. This cost-efficiency and low latency enable high-volume AI use cases that are uneconomical with larger frontier models. The benchmark evaluated 245 multi-step finance questions using pre-registered rubrics.

Read more
Google GemmaGoogle GemmaAug 22

Google Developer Experts Build Real-Time AI Race Coach at Sonoma Raceway

Google Developer Experts built a real-time AI Race Coach at Sonoma Raceway, running Gemma 4 locally on a Pixel 10. The system integrates Google Antigravity for telemetry orchestration and Gemini for cloud reasoning. This architecture delivers zero-latency audio coaching directly to drivers, processing data at 40 tokens per second to provide split-second track advice during high-speed racing.

Read more
Google GemmaGoogle GemmaAug 12

Google Powers Offline Reachy Mini Robot With Gemma 4 and LiteRT

Google demonstrates Gemma 4 E2B running entirely locally on a Raspberry Pi 5 to power the Reachy Mini robot. The system handles voice, vision, and movement offline, achieving text generation speeds of 300 words per minute on 1.5 GB of memory. LiteRT splits workloads between the CPU and GPU to maintain real-time, low-latency performance.

Read more
Google GemmaGoogle GemmaAug 10

Google Gemma Publishes Technical Report for Experimental DiffusionGemma Model

Google Gemma published the technical report for DiffusionGemma, an experimental open-weight model using discrete diffusion for text generation. By refining blocks of 256 tokens in parallel, it avoids the sequential bottleneck of autoregressive models. It achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU while retaining support for multimodal inputs and long contexts.

Read more
Google GemmaGoogle GemmaJul 17

Google Gemma Releases gemma-trainer Skill for Agent-Led Model Fine-Tuning

Google Gemma released gemma-trainer, a new skill that automates local fine-tuning for Gemma 4 models. The skill enables AI agents to manage training configurations, execute supervised fine-tuning, and evaluate performance. It includes guardrails to prevent capability mismatches and supports exporting models for edge devices using LiteRT-LM.

Read more
Google GemmaGoogle GemmaJul 16

Google Gemma 4 31B Launches on LiveKit Inference for Voice Agents

Google Gemma makes the Gemma 4 31B model available on LiveKit Inference, optimized for real-time voice agents. The model delivers a 192ms time-to-first-token and 354ms time-to-first-audio. It achieves 76.9% accuracy on the tau2bench agentic tool-use evaluation, surpassing GPT-4.1. This integration provides a low-latency, cost-efficient option for production voice agent workloads.

Read more
Google GemmaGoogle GemmaJul 15

Google Gemma 4 Gains Flash Attention 4 and Vision Token Options

Google updated Gemma 4 with uniform Flash Attention 4 support on NVIDIA Hopper GPUs, increasing prefill throughput by 25-70% and reducing time-to-first-token by up to 31%. The release also adds manual vision token scaling up to 1120 for higher resolution, patches tool-calling consistency, and reduces model laziness to ensure more complete, accurate responses.

Read more
Google GemmaGoogle GemmaJul 10

Google Gemma Challenge Results: 5x Faster Inference via Agent Collaboration

Google Gemma and the community completed a six-day challenge where over 100 AI agents and humans collaborated to optimize Gemma 4 inference on NVIDIA A10G hardware. The effort achieved a 5x speedup, reaching 315 tokens per second in lossless mode. Agents demonstrated emergent coordination, including resource pooling and self-policing to prevent quality degradation during the optimization process.

Read more

Google Gemma Fine-Tunes Gemma 4 to Translate Classical Korean Literature

Google Gemma demonstrates how to fine-tune the Gemma 4 E2B model on a single NVIDIA T4 GPU to translate Classical Korean into modern text. By using LoRA, the model’s character-by-character similarity score improved from 4.85% to 85.71%. This project provides a tutorial for preserving cultural history by making ancient texts accessible through lightweight model fine-tuning.

Read more

Google Brings Gemma 4 On-Device Support to React Native Apps

Google enables Gemma 4 to run fully offline in React Native applications using the react-native-executorch library. The integration supports local hardware acceleration via Vulkan on Android and MLX on Apple Silicon. This update allows mobile apps to perform on-device vision and tool-use tasks, such as reading documents and scheduling events, without cloud infrastructure.

Read more

Google DeepMind Releases Gemma 4 Technical Report Detailing New Architectures

Google DeepMind published the Gemma 4 technical report, detailing the architecture and efficiency of its latest open-weight multimodal models. The report covers the new encoder-free 12B variant, quantization-aware training, and multi-token prediction drafters. It also introduces a thinking mode that enables models to generate explicit reasoning traces before responding to prompts.

Read more
Google GemmaGoogle GemmaJun 28

Google Gemma 4 31B Launches on Cerebras with Virtual Hackathon

Google DeepMind released Gemma 4 31B on Cerebras, marking the platform's first multimodal model support with inference speeds reaching 1,500 tokens per second. To celebrate, Cerebras and Google DeepMind are hosting a 24-hour virtual hackathon starting June 28 with $5,000 in prizes. Participants receive early access to the model for the duration of the event.

Read more
Google GemmaGoogle GemmaJun 23

Google Gemma 4 26B A4B Achieves High Concurrency on DGX Spark

Google Gemma 4 26B A4B, quantized with NVFP4, sustains 16 parallel instances on a single NVIDIA DGX Spark. This configuration delivers 18 tokens per second per instance, with an aggregate throughput of 300 tokens per second. The architecture supports scaling up to 32 parallel runs, demonstrating efficient concurrent serving for multi-user and agentic workflows.

Read more
Google GemmaGoogle GemmaJun 18

Google Gemma 3 Enables Autonomous Image Analysis on Orbiting Satellite

Google DeepMind’s Gemma 3 model is now running on Loft Orbital’s YAM-9 satellite, marking the first use of a vision-language model in orbit. The system, powered by NASA JPL software, autonomously identifies areas of interest from sensor data based on natural language queries. This capability allows satellites to triage data in space, reducing the need for ground-based human analysis.

Google GemmaGoogle GemmaJun 18

Google Gemma Demonstrates Local Multi-Agent Orchestration with Gemma 4 26B

Google Gemma published a cookbook demo showing Gemma 4 26B orchestrating 10 parallel sub-agents locally on macOS. The system uses an orchestrator to delegate tasks to concurrent instances, achieving speeds over 100 tokens per second. This architecture enables multi-agent workflows for SVG generation, code, and translation to run entirely on-device without cloud infrastructure.

Read more
Google GemmaGoogle GemmaJun 16

Google Gemma 4 E2B Now Runs on Intel AI PC NPUs

Google updated Gemma 4 E2B to support NPU acceleration on Intel Core Ultra processors via LiteRT and OpenVINO. This integration delivers 1.3x faster prefill performance and a 2.8x improvement in performance-per-watt compared to GPU execution. The update enables background LLM tasks on Intel AI PCs with zero thermal throttling or heavy battery drain.

Read more
Google GemmaGoogle GemmaJun 13

Google Releases DiffusionGemma for 4x Faster Parallel Text Generation

Google released DiffusionGemma, an experimental open model that generates text using diffusion instead of sequential token prediction. By generating 256 tokens in parallel, it delivers up to 4x faster inference on dedicated GPUs, exceeding 1000 tokens per second on an H100. This 26B Mixture of Experts model supports real-time self-correction for tasks like code infilling and in-line editing.

Read more

Google Magenta RealTime 2 Turns MacBooks into Live AI Music Instruments

Google's Magenta Project released Magenta RealTime 2 (MRT2), an open-weight, open-source live music model. It enables low-latency, real-time music synthesis natively on Apple Silicon MacBooks using MIDI, text, and audio inputs. This allows musicians to play AI-generated music as an instrument directly on their device, fostering new creative workflows.

Google Gemma releases gemma-skills to accelerate agentic workflows with multi-token prediction

Google launched the first iteration of gemma-skills, an open-source library of reusable capabilities for AI agents. By standardizing how agents select model sizes and use performance optimizations like MTP, Google is making it easier to build efficient, autonomous workflows on top of the Gemma ecosystem.

Read more
Google GemmaGoogle GemmaMay 29

Google Launches On-Device Agent Skills for Offline Gemma 4 Workflows

Google released the Google AI Edge Gallery app and LiteRT-LM framework to enable fully offline agentic workflows on mobile and IoT devices. By running Gemma 4 locally, developers can build multi-step agents that plan, use tools, and process multimodal data without cloud latency or privacy risks.

Read more

Frequently asked questions

Google Gemma is Google's family of lightweight, open-weight AI models. HeadsUpAI tracks Google Gemma across the AI ecosystem and curates every significant update — the latest being "Google Gemma 4 31B Matches Sonnet 5 Performance at 40x Lower Cost" (August 22, 2026) — so you get the whole story in a 30-second read.

The most recent Google Gemma update is "Google Gemma 4 31B Matches Sonnet 5 Performance at 40x Lower Cost" (August 22, 2026). HeadsUpAI curates every significant Google Gemma release as a 30-second read — what shipped and why it matters.

The latest Google Gemma updates: "Google Gemma 4 31B Matches Sonnet 5 Performance at 40x Lower Cost", "Google Developer Experts Build Real-Time AI Race Coach at Sonoma Raceway", "Google Powers Offline Reachy Mini Robot With Gemma 4 and LiteRT", "Google Gemma Publishes Technical Report for Experimental DiffusionGemma Model", and "Google Gemma Releases gemma-trainer Skill for Agent-Led Model Fine-Tuning". HeadsUpAI has curated 20 Google Gemma updates over the last 90 days, covering product updates, launches, and research — listed newest first, presented straight, no hype, no bias.

Google Gemma is Google's family of lightweight, open-weight AI models. On this page you'll find every significant Google Gemma development HeadsUpAI has tracked recently — product updates, launches, and research — so you can keep up with where Google Gemma is heading without reading a dozen sources.

Continuously. HeadsUpAI adds new Google Gemma updates as they're announced — usually within hours — and the 20 updates currently shown cover the past 90 days, newest first.