Together AI

Together AI AI News & Updates30 Updates

The latest AI news and updates of Together AI — AI-native cloud platform for inference, fine-tuning, and training open-source and custom models. Covering Together AI's latest product updates, company news, and analysis from the past 90 days.

Together AITogether AIJul 25

Together AI Benchmarks Kimi K3 Max and GPT 5.6 Sol Performance

Together AI analyzed Kimi K3 Max and GPT 5.6 Sol Max on software engineering tasks using DeepSWE. Kimi K3 Max matches GPT 5.6 Sol Max performance at approximately 55% of the price. Combining both models delivers a 16% performance lift. Together AI launches Kimi K3 on its inference platform this Monday.

Read more
Together AITogether AIJul 25

The Washington Post Scales AI Journalism Using Together AI Infrastructure

The Washington Post processes 1.79 billion input tokens monthly through Together AI to power its Ask The Post AI service. By running open models like Llama and Mistral, the publication maintains full model ownership and predictable fixed pricing. The infrastructure delivers consistent 2-second response times for real-time reader interactions without relying on proprietary model APIs.

Read more
Together AITogether AIJul 23

Together AI Adds Moonshot AI's Kimi K3 to Inference Platform

Together AI will add Moonshot AI's Kimi K3 model to its serverless inference platform on July 27. The 2.8-trillion-parameter Mixture-of-Experts model features a 1-million-token context window and native vision capabilities. Together AI will power the model for coding, agentic tasks, and production workloads starting on its launch day.

Read more
Together AITogether AIJul 23

Together AI Launches Production-Grade Inference Platform with Advanced Deployment Controls

Together AI released a next-generation inference platform for open-weight models, featuring live-traffic testing, signal-based autoscaling, and stable endpoints. The platform includes pre-optimized deployment profiles and a closed beta for custom training, allowing direct deployment of fine-tuned checkpoints. These tools enable production-grade rollouts, including canary and blue-green updates, without requiring custom infrastructure management.

Read more
Together AITogether AIJul 20

Together AI and Y Combinator Launch Dedicated GPU Cluster for Startups

Together AI and Y Combinator launched the first dedicated GPU cluster for YC portfolio startups. The partnership replaces standard 24-month compute contracts with commitments as short as a few weeks. Startups manage GPU provisioning and billing through a self-service portal, providing flexible access to compute without the prohibitive upfront costs of long-term agreements.

Read more
Together AITogether AIJul 17

Together AI Adds Provisioned Throughput and Workload Calculator for MiniMax M3

Together AI now offers Provisioned Throughput for the MiniMax M3 model, featuring guaranteed capacity, production SLAs, and token-based pricing. This service provides costs over 80% lower than closed models. A new PTU calculator sizes production workloads, while MiniMax has released a breakdown of the economics for agentic tasks.

Read more
Together AITogether AIJul 17

Together AI Adds Inkling Multimodal Reasoning Model to Serverless Inference

Together AI now serves Inkling, a multimodal Mixture-of-Experts model from Thinking Machines Lab, on its serverless inference platform. The model features 975 billion total parameters, 41 billion active parameters, and a 1-million-token context window. It supports native text, image, and audio inputs with controllable inference effort, optimized via Together’s FlashAttention-4-based kernel.

Read more
Together AITogether AIJul 16

Together AI Powers Production Inference for Cursor and Grok 4.5

Together AI powers production inference for Cursor, supporting the platform's real-time coding agents with low-latency responses. The infrastructure runs on NVIDIA Blackwell GPUs with custom kernel optimizations, delivering the performance required for Cursor's in-editor intelligence at scale. This partnership confirms Together AI as a key inference layer behind one of the fastest-growing agentic coding tools.

Read more
Together AITogether AIJul 11

Together AI Launches Phone Interface for Real-Time AI Assistant Calls

Together AI launched a phone interface that connects callers directly to an AI assistant. Users dial +1 (415) 723-8167 and enter extension 676 to initiate a voice conversation. This interface provides a direct, text-free way to interact with the platform’s AI models, bypassing traditional chat interfaces for real-time voice communication.

Read more

Together AI Launches Provisioned Throughput for Frontier Open Models

Together AI launched Provisioned Throughput, a reserved inference capacity service for frontier open models. The service offers token-based pricing and a 99% uptime SLA, starting with MiniMax M3 and GLM-5.2. It provides guaranteed capacity for production workloads at costs up to 90% lower than Claude Opus 4.8 list prices, without requiring manual infrastructure management.

Read more

Together AI Ranks #1 for GLM 5.2 Speed and Latency

Together AI now holds the top position on Artificial Analysis for GLM 5.2 inference, achieving 446.1 tokens per second. This performance lead across 15 providers delivers the fastest output speed and lowest latency for the model. The ranking reflects Together AI’s inference stack optimizations for the 1-million-token context model.

Read more

Together AI Uses Blackwell Megakernels to Power Cursor Coding Agents

Together AI delivers sub-100ms inference for Cursor coding agents by integrating NVIDIA Blackwell GPUs with custom megakernels. This stack fuses an entire model forward pass into a single launch using CUDA, TensorRT-LLM, and the Dynamo inference framework. These optimizations enable real-time feedback loops for in-editor coding tasks.

Read more

Together AI Reports Open Model Token Usage Tripled in One Year

Together AI reports that open model usage has grown from 10% to 30% of all AI tokens over the past year. This shift toward open, modular AI is driven by significant cost efficiencies, with open models delivering 5x cheaper inference and 7x lower training costs compared to closed alternatives.

Read more
Together AITogether AIJun 25

Together AI Optimizes NVIDIA Parakeet for Record Speech-to-Text Throughput

Together AI engineered an inference stack that makes NVIDIA's Parakeet TDT 0.6B V3 the fastest speech-to-text model available. The system achieves a speed factor of ~302 seconds of audio per second of processing time, according to Artificial Analysis benchmarks. A Together AI voice engineer published a technical breakdown of the systems optimizations behind this performance.

Read more
Together AITogether AIJun 23

Together AI Releases ParallelKernelBench to Evaluate Multi-GPU Kernel Generation

Together AI released ParallelKernelBench, a benchmark evaluating how well LLMs generate multi-GPU CUDA kernels across 87 real-world problems. Frontier models struggle, solving under a third of tasks correctly, with only 31% of attempts outperforming standard PyTorch and NCCL baselines. While agentic loops improve performance, models still fail to reason about complex rank coordination and optimal data transfer mechanisms.

Together AITogether AIJun 22

Together AI and 5C Deploy NVIDIA GB300 NVL72 for AI Inference

Together AI and 5C are deploying NVIDIA GB300 NVL72 systems to power large-scale inference and reasoning workloads. The infrastructure integrates high-density compute, advanced cooling, and AI-optimized storage provided by Pegatron, Vertiv, and VAST Data. This deployment establishes purpose-built infrastructure designed for the next generation of AI inference tasks.

Read more
Together AITogether AIJun 19

Together AI Adds OpenAI GPT Image 2 for High-Fidelity Image Generation

Together AI now hosts OpenAI’s GPT Image 2 on its Serverless Inference platform for $0.053 per image. The model supports 1K to 4K resolutions, up to 16 reference images per call, and achieves 95%+ multilingual text rendering accuracy. These capabilities provide layout control and precise typography for design, marketing, and e-commerce workflows.

Read more
Together AITogether AIJun 18

Together AI Adds Z.ai GLM-5.2 for Long-Horizon Agentic Coding

Together AI has added Z.ai’s GLM-5.2 model to its serverless inference platform. The model features a 1-million-token context window, a 131,072-token output cap, and two configurable thinking-effort levels for complex agentic coding tasks. It is available via OpenAI-compatible and Anthropic Messages APIs, with input pricing starting at $1.40 per million tokens.

Read more
Together AITogether AIJun 16

Together AI Benchmark Finds Open Models Significantly Cheaper for Game Building

Together AI benchmarked closed and open models on building playable browser games. Open models, including MiniMax M3 and Nemotron Ultra, delivered comparable quality to Opus 4.8 and GPT-5.5 at 7x to 15x lower costs. These results demonstrate that open-weight models now provide competitive performance with significantly better tokenomics across a growing range of workloads.

Read more
Together AITogether AIJun 16

Decagon Cuts Voice Agent Costs 6x Using Together AI Infrastructure

Decagon reduced voice agent costs per turn nearly 6x by migrating from closed models to fine-tuned open-weight models on Together AI. The production stack achieves p95 model latency under 400ms using NVIDIA Blackwell GPUs, custom speculative decoders, and prompt caching. This infrastructure supports weekly model deployment cycles for Decagon’s conversational AI concierge agents.

Read more
Together AITogether AIJun 16

Together AI Adds Cartesia Sonic 3.5 TTS Model and Voice Finder

Together AI added the Cartesia Sonic 3.5 text-to-speech model to its platform, featuring 150+ voices in a new Voice Finder for comparison. The model delivers sub-90ms latency and native support for 42 languages, including alphanumeric handling for IDs and phone numbers. It deploys on Together AI's serverless or dedicated infrastructure.

Read more
Together AITogether AIJun 15

Together AI Optimizes MiniMax M3 Inference with New Systems Kernels

Together AI implemented custom engineering optimizations to serve MiniMax M3 at production scale. The team built a KV-block-major sparse attention kernel, integrated paged attention for MSA, and optimized decode index scoring. These changes, alongside a Rust-based multimodal preprocessing gateway, delivered 81–125% throughput improvements across varying concurrency levels for the 1-million-token context model.

Read more
Together AITogether AIJun 15

Together AI DeepSeek V4 Pro Deployment Tops Industry Speed Benchmarks

Together AI now ranks first on Artificial Analysis for DeepSeek V4 Pro inference, delivering 211.9 tokens per second. This performance lead across 11 providers stems from inference systems optimizations, including custom KV cache management, prefix reuse, and kernel tuning on NVIDIA HGX B200 hardware. The deployment achieves the lowest latency and highest output speed for the model.

Read more
Together AITogether AIJun 15

Together AI Adds Moonshot AI Kimi-K2.7-Code for Long-Horizon Coding

Together AI has made Moonshot AI’s Kimi-K2.7-Code model available on its inference platform. This coding-focused agentic model features a 256K context window and interleaved thinking for multi-step tool calling. It achieves 81.1% on MCP Mark Verified benchmarks, with pricing starting at $0.19 per million cached tokens and $0.95 per million input tokens.

Read more
Together AITogether AIJun 15

Together AI Delivers 31% Faster Coding Agent Inference on Blackwell

Together AI published coding agent benchmarks showing its inference engine achieves 31% more tokens per second than the next-fastest open-source engine on NVIDIA Blackwell hardware. These performance gains result from custom kernels targeting Blackwell Tensor Core instructions. Cursor now runs its real-time coding agents on this production stack to maintain low-latency feedback loops during development.

Read more
Together AITogether AIJun 15

Together AI Presents Untied Ulysses for Memory-Efficient Long-Context Training

Together AI researcher Max Ryabinin introduced Untied Ulysses, a context parallelism technique that optimizes GPU memory usage during transformer training. By chunking attention heads and reusing buffers across iterations, the method enables training 8B and 32B scale models on a single 8xH100 node with 25% longer sequences than prior implementations, overcoming memory limits that previously stalled 3M-token context training.

Together AITogether AIJun 15

Together AI Delivers Real-Time Blackwell Inference Infrastructure for Cursor Agents

Together AI built a real-time inference stack for Cursor’s in-editor coding agents using NVIDIA Blackwell GB200 NVL72 and B200 GPUs. The infrastructure features custom kernels for Blackwell Tensor Core instructions, ARM host optimization, and a quantization pipeline that moves internally trained model weights to production test endpoints within days, ensuring predictable latency for real-time code refactoring.

Read more

Together AI Adds Ideogram 4 for Design-First 2K Image Generation

Together AI has made Ideogram 4, an open image model from Ideogram AI, available on its Serverless Inference platform. This integration provides designers with a model focused on precise text rendering, layout control, and native 2K image output for production creative workflows. Its capabilities address specific needs in brand design and marketing, offering a specialized tool for high-quality visual content.

Read more

Together AI Adds NVIDIA Nemotron Models for Agentic AI and Real-Time Voice

Together AI has made NVIDIA's Nemotron 3 Ultra and Nemotron 3.5 ASR models available on its AI Native Cloud. This integration provides developers with specialized capabilities for building high-throughput AI agents and low-latency multilingual voice systems. The move expands access to advanced models for autonomous workflows and real-time conversational AI.

Read more

Together AI powers MiniMax M3 with 1M context and sparse attention

Together AI is now powering inference for MiniMax M3, a multimodal model featuring a 1-million-token context window. The model uses a new sparse attention architecture to process massive datasets with significantly lower computational overhead than previous-generation models.

Read more

Frequently asked questions

Together AI is AI-native cloud platform for inference, fine-tuning, and training open-source and custom models. HeadsUpAI tracks Together AI across the AI ecosystem and curates every significant update — the latest being "Together AI Benchmarks Kimi K3 Max and GPT 5.6 Sol Performance" (July 25, 2026) — so you get the whole story in a 30-second read.

The most recent Together AI update is "Together AI Benchmarks Kimi K3 Max and GPT 5.6 Sol Performance" (July 25, 2026). HeadsUpAI curates every significant Together AI release as a 30-second read — what shipped and why it matters.

The latest Together AI updates: "Together AI Benchmarks Kimi K3 Max and GPT 5.6 Sol Performance", "The Washington Post Scales AI Journalism Using Together AI Infrastructure", "Together AI Adds Moonshot AI's Kimi K3 to Inference Platform", "Together AI Launches Production-Grade Inference Platform with Advanced Deployment Controls", and "Together AI and Y Combinator Launch Dedicated GPU Cluster for Startups". HeadsUpAI has curated 30 Together AI updates over the last 90 days, covering product updates, company news, and analysis — listed newest first, presented straight, no hype, no bias.

Together AI is AI-native cloud platform for inference, fine-tuning, and training open-source and custom models. On this page you'll find every significant Together AI development HeadsUpAI has tracked recently — product updates, company news, and analysis — so you can keep up with where Together AI is heading without reading a dozen sources.

Continuously. HeadsUpAI adds new Together AI updates as they're announced — usually within hours — and the 30 updates currently shown cover the past 90 days, newest first.