Together AI

Together AI AI News & Updates54 Updates

The latest AI news and updates of Together AI — AI-native cloud platform for inference, fine-tuning, and training open-source and custom models. Covering Together AI's latest product updates, company news, and analysis from the past 90 days.

Together AITogether AI21h ago

Together AI Expands Fine-Tuning with New Models, Metrics, and Controls

Together AI expanded its fine-tuning service with support for GLM-5.3, Kimi K2.7-Code, and other open-weight models. The update adds live experiment tracking, early stopping, dataset previews, and finer training controls like Expert LoRA. Training prices for selected models are now 30–70% lower, alongside automated pre-flight validation for uploaded datasets.

Read more
Together AITogether AISep 11

Together AI Launches Preemptible Compute for GPU Clusters

Together AI introduced preemptible compute for its GPU clusters, offering the same NVIDIA infrastructure at a flat 50% of the on-demand rate. The service supports interruption-tolerant workloads like fine-tuning and batch inference, providing a five-minute drain window for checkpointing before node reclamation. Preemptible nodes automatically refill toward target capacity and are billed sub-hourly.

Read more
Together AITogether AISep 10

Together AI Ports ThunderKittens to NVIDIA Vera Rubin NVL72

Together AI ported its ThunderKittens kernel framework to the NVIDIA Vera Rubin NVL72 platform. The team reworked kernels to leverage the new ISA, achieving over 22 PFLOPS for NVFP4 and 12 PFLOPS for FP8 GEMMs. These results are competitive with cuBLAS and CuTe DSL, with the 16k NVFP4 GEMM reaching 22.4 PFLOPS through deepened pipelines and increased data reuse.

Read more
Together AITogether AISep 10

Together AI Reports Kimi K3 Outperforms Fable 5.1 on Legal Benchmark

Together AI reports that Kimi K3 scored 60% higher than Fable 5.1 on Harvey LAB-AA’s hard autonomous legal tasks. This benchmark measures the ability of models to complete complex, end-to-end legal work without human intervention.

Read more

Together AI Partners with Equinix and NVIDIA for Inference Exchange

Together AI is partnering with Equinix and NVIDIA to launch Equinix Inference Exchange, a distributed platform for enterprise AI inference. The service runs on NVIDIA Enterprise Reference Architectures across Equinix data centers, supporting over 200 open-source models. It offers both multitenant and dedicated single-tenant deployment modes, with general availability scheduled for Q1 2027.

Read more

Together AI Reduces Dedicated H100 Inference Price to $3.99 Hourly

Together AI reduced the price of Dedicated Inference on NVIDIA H100 GPUs from $5.49 to $3.99 per hour for September 2026. This lower rate applies automatically to both new and existing deployments. The service supports various open-weight models, including Llama 3.3, Qwen3.5, and Nemotron 3.5 Lightning, as well as custom fine-tuned LoRA checkpoints.

Read more
Together AITogether AIAug 31

Together AI and HUMAIN Build 250MW Data Center in Saudi Arabia

Together AI signed a deal with HUMAIN to construct a 250MW data center in Saudi Arabia. The project is projected to generate over $5 billion in annualized revenue. This partnership secures large-scale compute capacity for open-source AI, addressing the power constraints currently limiting infrastructure development.

Read more
Together AITogether AIAug 31

Together AI Reports GLM-5.3 Hallucination Rates Below Claude and GPT Models

Together AI reports that Z.ai’s GLM-5.3 model exhibits a lower hallucination rate than frontier competitors. The company found Claude Fable 5’s hallucination rate is over 2x higher, while GPT-5.6 Luna’s rate is over 3x higher. These findings establish the model as a reliable option for production workflows requiring factual accuracy.

Read more
Together AITogether AIAug 27

Together AI Announces GLM-5.3 Flash Multimodal Model

Together AI introduced GLM-5.3 Flash, a natively multimodal model from Z.ai featuring 320B total parameters, 18B active parameters, and a 1-million-token context window. The model utilizes hybrid attention to match Luna’s performance on the DeepSWE benchmark while delivering twice the throughput for the same budget. It is coming soon to Together AI’s serverless API.

Read more
Together AITogether AIAug 25

Together AI Adds Alibaba Qwen3.8 27B for Fine-Tuning and Inference

Together AI adds Alibaba's Qwen3.8 27B model to its platform for fine-tuning and Dedicated Model Inference. This integration supports training on custom data and deployment on dedicated production infrastructure within a single stack, removing the need to manage separate training and serving environments.

Read more
Together AITogether AIAug 23

Together AI DeepSWE Analysis: GLM-5.3 Matches Claude Fable 5 Coding

Together AI analyzed GLM-5.3 and Claude Fable 5 on the DeepSWE benchmark, finding a statistical tie in single-shot accuracy. GLM-5.3 cost 5.4x less per task and achieved an 87.6% solve rate across four attempts for approximately $16, compared to Fable 5’s 84.1% at $21.63. The results highlight GLM-5.3 as a cost-effective default for coding agents.

Read more
Together AITogether AIAug 14

Together AI and L&T Build India's Largest NVIDIA B300 AI Factory

Together AI and Larsen & Toubro are building India's largest AI Factory in Chennai, featuring 10,000 NVIDIA B300 GPUs. The project, valued at up to $1.57 billion, provides dedicated infrastructure for large-scale open-source inference, fine-tuning, and training. This deployment marks the country's largest single-cluster AI infrastructure to date.

Read more
Together AITogether AIAug 13

Together AI Adds ByteDance Seedance 2.5 for 30-Second Video Generation

Together AI now hosts ByteDance’s Seedance 2.5, a video generation model capable of producing 30-second audio-video clips in a single pass. The model supports multi-round extensions for consistent long-form storytelling and offers timestamp-level editing for narrative and camera control. It accepts up to 50 multimodal reference inputs and is available on Together AI’s serverless infrastructure for $0.115 per video.

Read more
Together AITogether AIAug 13

Together AI Adds Qwen3.8-2.4T-A95B Flagship Model for Agentic Workflows

Together AI now hosts Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter sparse Mixture-of-Experts model. The model features a 256K-token context window and native multimodal input, with adjustable reasoning effort levels for complex coding and long-horizon agentic tasks. It supports function calling and structured outputs for production-grade agent workflows.

Read more
Together AITogether AIAug 11

Together AI, IBM, and NVIDIA Partner for B300 Inference Cluster

Together AI, IBM, and NVIDIA signed a 240 million dollar multi-year agreement to deploy a dedicated NVIDIA HGX B300 cluster on IBM Cloud. Expected in Q1 2027, the cluster uses Spectrum-X networking to provide enterprise-grade inference for open-source models. This deployment marks the first large-scale inference cluster of its kind on IBM Cloud, designed for high-performance production workloads.

Read more
Together AITogether AIAug 11

Together AI Adds NVIDIA Nemotron 3.5 Lightning for High-Volume Agents

Together AI now hosts NVIDIA Nemotron 3.5 Lightning, a 30B-parameter hybrid Mixture-of-Experts model with 3B active parameters. Available through Dedicated Model Inference, the model supports 1M-token context windows for long-running agent sessions. It delivers up to 4x higher throughput and 30% faster task completion for specialized, high-volume agentic workflows.

Read more
Together AITogether AIAug 10

Together AI Powers DeepSeek-V4-Flash Inference on Ollama Cloud

Together AI now powers DeepSeek-V4-Flash-0731 on Ollama's cloud, providing the inference infrastructure for the 284B-parameter model. The deployment delivers over 120 output tokens per second with zero data retention and hosting in the US and Europe. This setup serves as the default for Ollama cloud users, supporting long-running sessions for coding agents.

Read more

Together AI Benchmarks DeepSeek-V4 Flash Against GPT-5.6 Luna

Together AI analyzed DeepSeek-V4 Flash-0731 and GPT-5.6 Luna on DeepSWE software engineering tasks. The analysis shows DeepSeek delivers 80% of Luna’s performance at one-sixth the cost. A cascade strategy, running DeepSeek first and escalating to Luna on failure, solves 78.9% of tasks at $0.385 each, outperforming Luna alone in both accuracy and cost-efficiency.

Read more

Together AI Integrates as Inference Provider for Roomote Agent Platform

Together AI now serves as an inference provider for Roomote, an agentic platform for multi-step workflows. The integration supports model mapping for task-specific execution, assigning distinct open models to coding, planning, vision, and review roles. Roomote provides direct selection of models like GLM 5.2 and MiniMax M3 within the platform interface.

Read more

Together AI Launches Free No-Setup Chat Access for Kimi K3

Together AI now provides free, no-setup chat access to Moonshot AI’s Kimi K3 model through Together Chat. This interface allows direct prompting without requiring an API key or configuration. The service runs on Together AI’s secure North American infrastructure, providing a testing environment for the sparse mixture-of-experts model.

Read more

Together AI Benchmarks Kimi K3 and GPT-5.6 Sol on DeepSWE

Together AI analyzed Moonshot AI’s Kimi K3 and OpenAI’s GPT-5.6 Sol on the DeepSWE benchmark. While Sol leads in single-shot accuracy, Kimi K3 wins on pass@4 at a lower cost. A Kimi-first cascade with test-suite verification reaches 85.6% accuracy, outperforming either model alone while reducing costs per solved task.

Read more

Together AI Now Serves DeepSeek V4 Flash for Agentic Workloads

Together AI now serves DeepSeek V4 Flash on its inference platform. The company says the model delivers a major jump in coding, tool use, and long-running agent performance, with low, high, and max reasoning effort and DSpark speculative decoding for faster generation. Pricing starts at $0.14 per million input tokens and $0.28 per million output tokens.

Read more

Together AI Benchmarks Kimi K3 Max Against Claude Fable 5

Together AI analyzed Kimi K3 Max and Claude Fable 5 on the DeepSWE benchmark. While Fable 5 leads on pass@1 at 69.9% versus 68.5%, Kimi K3 Max wins on pass@4 at 89.4% versus 88.5%. Kimi K3 Max costs one-third as much per rollout, delivering 2.8 times more solved tasks per dollar for high-volume coding agent workloads.

Read more

Together AI Monthly Inference Volume Surges to 400 Trillion Tokens

Together AI now processes over 400 trillion inference tokens per month, marking a 10,000x increase from 30 billion tokens one year ago. This growth reflects a shift as enterprises and AI-native companies move production workloads to open-source models. The platform supports this scale through its production-grade inference infrastructure and provisioned throughput services.

Read more

Together AI Details Effective Autoscaling Metrics for Dedicated Model Inference

Together AI published a deep-dive on autoscaling dedicated inference, noting that traditional CPU-style utilisation metrics often fail to capture real queue pressure. The guide recommends using in-flight requests as the default scaling metric, details how to configure scaling windows to avoid sawtooth oscillations, and provides measured cold-start costs for H100 GPUs alongside experimental results on scaling policies.

Read more
Together AITogether AIJul 31

Together AI Details Capacity-Aware Routing for Dedicated Model Inference Deployments

Together AI detailed its capacity-aware routing model for Dedicated Model Inference, which distributes traffic based on weight multiplied by ready replicas rather than fixed percentages. This architecture automatically scales traffic shares as deployments grow or shrink. It also supports advanced operations like rollouts, A/B tests, and shadow traffic by treating them as deployments with specific routing weights.

Read more
Together AITogether AIJul 30

Together AI Releases ThunderAgent for Faster, Program-Aware Agentic Inference

Together AI released ThunderAgent, a program-aware scheduler that eliminates KV cache thrashing in agentic workflows. By treating workflows as schedulable programs, it achieves 2.5x higher single-node throughput and 10x lower P50 latency at high concurrency. The open-source system supports near-linear multi-node scaling and integrates as a drop-in via a single program_id field.

Read more
Together AITogether AIJul 25

Together AI Benchmarks Kimi K3 Max and GPT 5.6 Sol Performance

Together AI analyzed Kimi K3 Max and GPT 5.6 Sol Max on software engineering tasks using DeepSWE. Kimi K3 Max matches GPT 5.6 Sol Max performance at approximately 55% of the price. Combining both models delivers a 16% performance lift. Together AI launches Kimi K3 on its inference platform this Monday.

Read more
Together AITogether AIJul 25

The Washington Post Scales AI Journalism Using Together AI Infrastructure

The Washington Post processes 1.79 billion input tokens monthly through Together AI to power its Ask The Post AI service. By running open models like Llama and Mistral, the publication maintains full model ownership and predictable fixed pricing. The infrastructure delivers consistent 2-second response times for real-time reader interactions without relying on proprietary model APIs.

Read more
Together AITogether AIJul 23

Together AI Adds Moonshot AI's Kimi K3 to Inference Platform

Together AI will add Moonshot AI's Kimi K3 model to its serverless inference platform on July 27. The 2.8-trillion-parameter Mixture-of-Experts model features a 1-million-token context window and native vision capabilities. Together AI will power the model for coding, agentic tasks, and production workloads starting on its launch day.

Read more
Together AITogether AIJul 23

Together AI Launches Production-Grade Inference Platform with Advanced Deployment Controls

Together AI released a next-generation inference platform for open-weight models, featuring live-traffic testing, signal-based autoscaling, and stable endpoints. The platform includes pre-optimized deployment profiles and a closed beta for custom training, allowing direct deployment of fine-tuned checkpoints. These tools enable production-grade rollouts, including canary and blue-green updates, without requiring custom infrastructure management.

Read more
Together AITogether AIJul 20

Together AI and Y Combinator Launch Dedicated GPU Cluster for Startups

Together AI and Y Combinator launched the first dedicated GPU cluster for YC portfolio startups. The partnership replaces standard 24-month compute contracts with commitments as short as a few weeks. Startups manage GPU provisioning and billing through a self-service portal, providing flexible access to compute without the prohibitive upfront costs of long-term agreements.

Read more
Together AITogether AIJul 17

Together AI Adds Provisioned Throughput and Workload Calculator for MiniMax M3

Together AI now offers Provisioned Throughput for the MiniMax M3 model, featuring guaranteed capacity, production SLAs, and token-based pricing. This service provides costs over 80% lower than closed models. A new PTU calculator sizes production workloads, while MiniMax has released a breakdown of the economics for agentic tasks.

Read more
Together AITogether AIJul 17

Together AI Adds Inkling Multimodal Reasoning Model to Serverless Inference

Together AI now serves Inkling, a multimodal Mixture-of-Experts model from Thinking Machines Lab, on its serverless inference platform. The model features 975 billion total parameters, 41 billion active parameters, and a 1-million-token context window. It supports native text, image, and audio inputs with controllable inference effort, optimized via Together’s FlashAttention-4-based kernel.

Read more
Together AITogether AIJul 16

Together AI Powers Production Inference for Cursor and Grok 4.5

Together AI powers production inference for Cursor, supporting the platform's real-time coding agents with low-latency responses. The infrastructure runs on NVIDIA Blackwell GPUs with custom kernel optimizations, delivering the performance required for Cursor's in-editor intelligence at scale. This partnership confirms Together AI as a key inference layer behind one of the fastest-growing agentic coding tools.

Read more
Together AITogether AIJul 11

Together AI Launches Phone Interface for Real-Time AI Assistant Calls

Together AI launched a phone interface that connects callers directly to an AI assistant. Users dial +1 (415) 723-8167 and enter extension 676 to initiate a voice conversation. This interface provides a direct, text-free way to interact with the platform’s AI models, bypassing traditional chat interfaces for real-time voice communication.

Read more

Together AI Launches Provisioned Throughput for Frontier Open Models

Together AI launched Provisioned Throughput, a reserved inference capacity service for frontier open models. The service offers token-based pricing and a 99% uptime SLA, starting with MiniMax M3 and GLM-5.2. It provides guaranteed capacity for production workloads at costs up to 90% lower than Claude Opus 4.8 list prices, without requiring manual infrastructure management.

Read more

Together AI Ranks #1 for GLM 5.2 Speed and Latency

Together AI now holds the top position on Artificial Analysis for GLM 5.2 inference, achieving 446.1 tokens per second. This performance lead across 15 providers delivers the fastest output speed and lowest latency for the model. The ranking reflects Together AI’s inference stack optimizations for the 1-million-token context model.

Read more

Together AI Uses Blackwell Megakernels to Power Cursor Coding Agents

Together AI delivers sub-100ms inference for Cursor coding agents by integrating NVIDIA Blackwell GPUs with custom megakernels. This stack fuses an entire model forward pass into a single launch using CUDA, TensorRT-LLM, and the Dynamo inference framework. These optimizations enable real-time feedback loops for in-editor coding tasks.

Read more

Together AI Reports Open Model Token Usage Tripled in One Year

Together AI reports that open model usage has grown from 10% to 30% of all AI tokens over the past year. This shift toward open, modular AI is driven by significant cost efficiencies, with open models delivering 5x cheaper inference and 7x lower training costs compared to closed alternatives.

Read more
Together AITogether AIJun 25

Together AI Optimizes NVIDIA Parakeet for Record Speech-to-Text Throughput

Together AI engineered an inference stack that makes NVIDIA's Parakeet TDT 0.6B V3 the fastest speech-to-text model available. The system achieves a speed factor of ~302 seconds of audio per second of processing time, according to Artificial Analysis benchmarks. A Together AI voice engineer published a technical breakdown of the systems optimizations behind this performance.

Read more
Together AITogether AIJun 23

Together AI Releases ParallelKernelBench to Evaluate Multi-GPU Kernel Generation

Together AI released ParallelKernelBench, a benchmark evaluating how well LLMs generate multi-GPU CUDA kernels across 87 real-world problems. Frontier models struggle, solving under a third of tasks correctly, with only 31% of attempts outperforming standard PyTorch and NCCL baselines. While agentic loops improve performance, models still fail to reason about complex rank coordination and optimal data transfer mechanisms.

Together AITogether AIJun 22

Together AI and 5C Deploy NVIDIA GB300 NVL72 for AI Inference

Together AI and 5C are deploying NVIDIA GB300 NVL72 systems to power large-scale inference and reasoning workloads. The infrastructure integrates high-density compute, advanced cooling, and AI-optimized storage provided by Pegatron, Vertiv, and VAST Data. This deployment establishes purpose-built infrastructure designed for the next generation of AI inference tasks.

Read more
Together AITogether AIJun 19

Together AI Adds OpenAI GPT Image 2 for High-Fidelity Image Generation

Together AI now hosts OpenAI’s GPT Image 2 on its Serverless Inference platform for $0.053 per image. The model supports 1K to 4K resolutions, up to 16 reference images per call, and achieves 95%+ multilingual text rendering accuracy. These capabilities provide layout control and precise typography for design, marketing, and e-commerce workflows.

Read more
Together AITogether AIJun 18

Together AI Adds Z.ai GLM-5.2 for Long-Horizon Agentic Coding

Together AI has added Z.ai’s GLM-5.2 model to its serverless inference platform. The model features a 1-million-token context window, a 131,072-token output cap, and two configurable thinking-effort levels for complex agentic coding tasks. It is available via OpenAI-compatible and Anthropic Messages APIs, with input pricing starting at $1.40 per million tokens.

Read more
Together AITogether AIJun 16

Together AI Benchmark Finds Open Models Significantly Cheaper for Game Building

Together AI benchmarked closed and open models on building playable browser games. Open models, including MiniMax M3 and Nemotron Ultra, delivered comparable quality to Opus 4.8 and GPT-5.5 at 7x to 15x lower costs. These results demonstrate that open-weight models now provide competitive performance with significantly better tokenomics across a growing range of workloads.

Read more
Together AITogether AIJun 16

Decagon Cuts Voice Agent Costs 6x Using Together AI Infrastructure

Decagon reduced voice agent costs per turn nearly 6x by migrating from closed models to fine-tuned open-weight models on Together AI. The production stack achieves p95 model latency under 400ms using NVIDIA Blackwell GPUs, custom speculative decoders, and prompt caching. This infrastructure supports weekly model deployment cycles for Decagon’s conversational AI concierge agents.

Read more
Together AITogether AIJun 16

Together AI Adds Cartesia Sonic 3.5 TTS Model and Voice Finder

Together AI added the Cartesia Sonic 3.5 text-to-speech model to its platform, featuring 150+ voices in a new Voice Finder for comparison. The model delivers sub-90ms latency and native support for 42 languages, including alphanumeric handling for IDs and phone numbers. It deploys on Together AI's serverless or dedicated infrastructure.

Read more
Together AITogether AIJun 15

Together AI Optimizes MiniMax M3 Inference with New Systems Kernels

Together AI implemented custom engineering optimizations to serve MiniMax M3 at production scale. The team built a KV-block-major sparse attention kernel, integrated paged attention for MSA, and optimized decode index scoring. These changes, alongside a Rust-based multimodal preprocessing gateway, delivered 81–125% throughput improvements across varying concurrency levels for the 1-million-token context model.

Read more
Together AITogether AIJun 15

Together AI DeepSeek V4 Pro Deployment Tops Industry Speed Benchmarks

Together AI now ranks first on Artificial Analysis for DeepSeek V4 Pro inference, delivering 211.9 tokens per second. This performance lead across 11 providers stems from inference systems optimizations, including custom KV cache management, prefix reuse, and kernel tuning on NVIDIA HGX B200 hardware. The deployment achieves the lowest latency and highest output speed for the model.

Read more
Together AITogether AIJun 15

Together AI Adds Moonshot AI Kimi-K2.7-Code for Long-Horizon Coding

Together AI has made Moonshot AI’s Kimi-K2.7-Code model available on its inference platform. This coding-focused agentic model features a 256K context window and interleaved thinking for multi-step tool calling. It achieves 81.1% on MCP Mark Verified benchmarks, with pricing starting at $0.19 per million cached tokens and $0.95 per million input tokens.

Read more
Together AITogether AIJun 15

Together AI Delivers 31% Faster Coding Agent Inference on Blackwell

Together AI published coding agent benchmarks showing its inference engine achieves 31% more tokens per second than the next-fastest open-source engine on NVIDIA Blackwell hardware. These performance gains result from custom kernels targeting Blackwell Tensor Core instructions. Cursor now runs its real-time coding agents on this production stack to maintain low-latency feedback loops during development.

Read more
Together AITogether AIJun 15

Together AI Presents Untied Ulysses for Memory-Efficient Long-Context Training

Together AI researcher Max Ryabinin introduced Untied Ulysses, a context parallelism technique that optimizes GPU memory usage during transformer training. By chunking attention heads and reusing buffers across iterations, the method enables training 8B and 32B scale models on a single 8xH100 node with 25% longer sequences than prior implementations, overcoming memory limits that previously stalled 3M-token context training.

Together AITogether AIJun 15

Together AI Delivers Real-Time Blackwell Inference Infrastructure for Cursor Agents

Together AI built a real-time inference stack for Cursor’s in-editor coding agents using NVIDIA Blackwell GB200 NVL72 and B200 GPUs. The infrastructure features custom kernels for Blackwell Tensor Core instructions, ARM host optimization, and a quantization pipeline that moves internally trained model weights to production test endpoints within days, ensuring predictable latency for real-time code refactoring.

Read more

Frequently asked questions

Together AI is AI-native cloud platform for inference, fine-tuning, and training open-source and custom models. HeadsUpAI tracks Together AI across the AI ecosystem and curates every significant update — the latest being "Together AI Expands Fine-Tuning with New Models, Metrics, and Controls" (September 11, 2026) — so you get the whole story in a 30-second read.

The most recent Together AI update is "Together AI Expands Fine-Tuning with New Models, Metrics, and Controls" (September 11, 2026). HeadsUpAI curates every significant Together AI release as a 30-second read — what shipped and why it matters.

The latest Together AI updates: "Together AI Expands Fine-Tuning with New Models, Metrics, and Controls", "Together AI Launches Preemptible Compute for GPU Clusters", "Together AI Ports ThunderKittens to NVIDIA Vera Rubin NVL72", "Together AI Reports Kimi K3 Outperforms Fable 5.1 on Legal Benchmark", and "Together AI Partners with Equinix and NVIDIA for Inference Exchange". HeadsUpAI has curated 54 Together AI updates over the last 90 days, covering product updates, company news, and analysis — listed newest first, presented straight, no hype, no bias.

Together AI is AI-native cloud platform for inference, fine-tuning, and training open-source and custom models. On this page you'll find every significant Together AI development HeadsUpAI has tracked recently — product updates, company news, and analysis — so you can keep up with where Together AI is heading without reading a dozen sources.

Continuously. HeadsUpAI adds new Together AI updates as they're announced — usually within hours — and the 54 updates currently shown cover the past 90 days, newest first.