Ollama

Ollama AI News & Updates23 Updates

The latest AI news and updates of Ollama — Open-source tool for running, managing, and serving large language models locally. Covering Ollama's latest product updates and company news from the past 90 days.

OllamaOllamaSep 11

Ollama Rolls Out DeepSeek-V4.1-Flash to Cloud Service Subscribers

Ollama has begun rolling out DeepSeek-V4.1-Flash on its cloud platform, starting with Max and Team account holders. The model features a 1-million-token context window and native visual understanding. Ollama is currently adding capacity to extend access to all subscribers.

Read more
OllamaOllamaSep 11

Ollama 0.34 Adds Local and Cloud Model Support to ChatGPT Desktop

Ollama 0.34 enables ChatGPT Desktop to run on local and cloud-hosted Ollama models. The integration supports plugins, web search, and computer use out of the box. Model selection is customizable within the app’s settings, and the setup process runs via the command line using the ollama launch codex-app command.

Read more
OllamaOllamaSep 6

Ollama Introduces Off-Peak Token Pricing for DeepSeek Cloud Models

Ollama introduced off-peak token rates for DeepSeek-V4-Flash and Pro models on its cloud service. These models are now half price outside the 12:00 to 18:00 UTC weekday peak window and all day on weekends. The service, hosted in the US and Europe with zero data retention, plans to expand off-peak pricing to additional models soon.

Read more
OllamaOllamaSep 1

Ollama Transitions Cloud Plans to Transparent Per-Token Pricing Model

Ollama transitioned its Pro, Max, and Team cloud plans to transparent per-token pricing, replacing previous GPU-time billing. Each plan now includes a monthly pool of usage credits, such as $60 for the $20 Pro tier and $1,000 for the $500 Team plan. The Team plan is now generally available for unlimited users with shared credits and no service fees.

Read more
OllamaOllamaAug 29

Ollama Adds Z.ai GLM 5.3 and GLM 5.3 Flash to Cloud

Ollama has fully rolled out Z.ai's GLM 5.3 and GLM 5.3 Flash models on its cloud service. Hosted in the US and Europe with zero data retention, these models support coding harnesses like Claude Code and OpenCode. Both models are accessible via API key, with adjustable reasoning effort settings available for the flagship GLM 5.3.

Read more
OllamaOllamaAug 26

Ollama v0.33 Adds Claude Desktop Integration as Gateway Provider

Ollama v0.33 adds Claude Desktop support, connecting the application as a third-party gateway provider. This integration brings local models and Ollama cloud models into the Claude Desktop interface. The update includes support for auto mode and web search, while maintaining Ollama’s Zero Data Retention policy for all requests.

Read more
OllamaOllamaAug 25

Ollama Adds IBM Granite 4.2 Enterprise-Ready Open Models

Ollama now hosts IBM Granite 4.2, a family of open foundation models in 3B, 8B, and 30B parameter sizes. These models feature a 128K context window, multilingual support, and a built-in thinking mode for reasoning. Licensed under Apache 2.0, the models are trained with governance, risk, and compliance evaluations specifically for enterprise agent applications.

Read more
OllamaOllamaAug 15

Ollama Deploys DeepSeek-V4-Pro Frontier Model to Cloud Service

Ollama has fully rolled out DeepSeek-V4-Pro (0813) on its cloud platform for Pro and Max subscribers. This 1.6T-parameter Mixture-of-Experts model features a 1-million-token context window and three distinct reasoning modes for logical analysis. The service is hosted in the US with a Zero Data Retention policy, accessible via the ollama run deepseek-v4-pro:cloud command.

Read more
OllamaOllamaAug 14

Ollama to Add Support for Z.ai's GLM-5.3 Coding Model

Ollama will support Z.ai's upcoming GLM-5.3 model upon its open release. The model, built on a 743B base, features specialized capabilities for coding, agentic workflows, and cybersecurity. This integration brings the model's post-trained performance to local environments, supporting autonomous code generation and defense tasks.

Read more
OllamaOllamaAug 12

Ollama Integrates as a Model Provider for GitHub Copilot in JetBrains

Ollama now functions as a provider for GitHub Copilot in JetBrains IDEs. This integration enables provider configuration and model selection directly within the JetBrains experience. Local models managed by Ollama now power Copilot chat and code suggestions, offering an alternative to cloud-based model options.

Read more
OllamaOllamaAug 11

Ollama Partners With NVIDIA to Launch Nemotron 3.5 Lightning Model

Ollama partners with NVIDIA to launch Nemotron 3.5 Lightning, a 30B-parameter mixture-of-experts model with 3.6B active parameters. The model targets high-volume agentic workflows, providing efficient reasoning for long-running tasks. This integration makes the open model available for local execution, supporting autonomous agent development.

Read more
OllamaOllamaAug 10

Ollama Adds Support for Meta’s Muse Glimmer 30B Model

Ollama adds support for Meta’s Muse Glimmer 30B model, enabling local agent workflows on Apple Silicon. The integration uses Ollama’s MLX engine with DFlash for 1.5x–1.8x faster performance and adds native image input support for coding agents. The model integrates with tools like Claude Code and Codex, featuring controllable reasoning strength settings.

Read more
OllamaOllamaAug 8

Ollama Sets DeepSeek-V4-Flash-0731 as Default Cloud Model

Ollama rolled out DeepSeek-V4-Flash-0731 as the new default model for its cloud inference service. The update delivers over 120 output tokens per second with zero data retention in the US and Europe. Pro and Max plan subscribers receive generous usage for long-running sessions with coding agents like Claude Code, OpenCode, and Hermes Agent.

Read more
OllamaOllamaAug 5

Ollama Scales Cloud Capacity for DeepSeek-V4-Flash-0731 Following Record Usage

Ollama is scaling its cloud capacity in the US and Europe following record-breaking token usage for the DeepSeek-V4-Flash-0731 model. The model now runs on Ollama’s cloud with performance exceeding 100 tokens per second. Ollama maintains zero data retention for all requests, ensuring user data remains private during inference.

Read more
OllamaOllamaJul 27

Ollama Adds Moonshot AI’s Kimi K3 to Cloud Model Library

Ollama added Moonshot AI’s Kimi K3, a 2.8-trillion-parameter multimodal model, to its cloud platform. The model features a 1-million-token context window and native visual understanding. It is available to Pro and Max subscribers via extra usage credits and integrates with agentic coding tools including Claude Code, OpenCode, Hermes Agent, and OpenClaw.

Read more
OllamaOllamaJul 17

Ollama 0.32.1 Improves Gemma 4 Tool Calling Reliability

Ollama 0.32.1 updates Gemma 4 with improved tool calling reliability for coding agents. The release addresses previous consistency issues, ensuring more dependable performance in autonomous workflows. The 26B model runs via the Pi agent, with an optional MLX engine available for maximum performance on Apple Silicon.

Read more
OllamaOllamaJul 15

Ollama Adds Support for OpenCode Desktop Coding Agent

Ollama now integrates with OpenCode Desktop, enabling the terminal-based coding agent to run with local or cloud models. The integration launches via the ollama launch opencode command and supports file editing, command execution, web fetching, and vision capabilities. Models require a context window of 64k tokens or higher for repository-wide tasks.

Read more
OllamaOllamaJul 8

Ollama Expands Cloud Capacity for GLM-5.2 in US and Europe

Ollama increased cloud capacity for Z.ai’s GLM-5.2 model across US and European regions. The infrastructure update delivers 80 to 120 output tokens per second, significantly outpacing the 30 to 40 tokens per second reported on other providers. The model remains accessible for agentic coding workflows via Claude Code, VS Code, Codex, and Hermes using the ollama launch command.

Read more
OllamaOllamaJul 1

Ollama Accelerates Gemma 4 on Apple Silicon with Multi-Token Prediction

Ollama 0.31 enables multi-token prediction by default for Gemma 4 on Apple Silicon, increasing generation speed by nearly 90%. The engine uses a small draft model to propose tokens, which the main model verifies in a single pass. This optimization, measured at 95.0 tokens per second on an M5 Max, accelerates agentic coding tasks without requiring manual configuration.

Read more
OllamaOllamaJun 28

Ollama Adds Ornith-1.0 Family for Agentic Coding Tasks

Ollama adds the Ornith-1.0 family of open-source agentic coding models to its library. The models, ranging from 9B to 397B parameters, feature a 256K context window and use reinforcement learning for self-improving code generation. The integration supports direct execution via ollama run and agentic workflows through ollama launch for Claude Code and Pi.

Read more
OllamaOllamaJun 16

Ollama Hosts Z.ai's GLM-5.2 Coding Model on NVIDIA Blackwell GPUs

Ollama is now hosting Z.ai's GLM-5.2 model on its US-based cloud, powered by NVIDIA Blackwell GPUs. The model features a 1-million-token context window and two reasoning effort levels for long-horizon coding tasks. It is available for immediate use within Claude Code, Codex App, and Hermes Agent via the ollama launch command, with a zero-data-retention privacy policy.

Read more
OllamaOllamaJun 15

Ollama Adds Support for Cline CLI and Parallel Kanban Tasks

Ollama now supports Cline CLI, enabling the autonomous coding agent to launch directly from the terminal. The integration auto-configures local model providers and supports the Kanban feature, which runs multiple coding agents in parallel to handle separate project tasks simultaneously. Sessions launch with a chosen local or cloud model via the ollama launch cline command.

Read more
OllamaOllamaJun 15

Ollama Adds Moonshot AI Kimi K2.7 Code to Cloud Platform

Ollama added Moonshot AI's Kimi K2.7 Code model to its US-hosted cloud on NVIDIA B300 GPUs. The model supports text and image input with a 256K context window. It is available for immediate use within Claude Code, Codex App, and OpenCode via the ollama launch command, with no user data retention or training.

Read more

Frequently asked questions

Ollama is Open-source tool for running, managing, and serving large language models locally. HeadsUpAI tracks Ollama across the AI ecosystem and curates every significant update — the latest being "Ollama Rolls Out DeepSeek-V4.1-Flash to Cloud Service Subscribers" (September 11, 2026) — so you get the whole story in a 30-second read.

The most recent Ollama update is "Ollama Rolls Out DeepSeek-V4.1-Flash to Cloud Service Subscribers" (September 11, 2026). HeadsUpAI curates every significant Ollama release as a 30-second read — what shipped and why it matters.

The latest Ollama updates: "Ollama Rolls Out DeepSeek-V4.1-Flash to Cloud Service Subscribers", "Ollama 0.34 Adds Local and Cloud Model Support to ChatGPT Desktop", "Ollama Introduces Off-Peak Token Pricing for DeepSeek Cloud Models", "Ollama Transitions Cloud Plans to Transparent Per-Token Pricing Model", and "Ollama Adds Z.ai GLM 5.3 and GLM 5.3 Flash to Cloud". HeadsUpAI has curated 23 Ollama updates over the last 90 days, covering product updates and company news — listed newest first, presented straight, no hype, no bias.

Ollama is Open-source tool for running, managing, and serving large language models locally. On this page you'll find every significant Ollama development HeadsUpAI has tracked recently — product updates and company news — so you can keep up with where Ollama is heading without reading a dozen sources.

Continuously. HeadsUpAI adds new Ollama updates as they're announced — usually within hours — and the 23 updates currently shown cover the past 90 days, newest first.