Fireworks AI

Fireworks AI AI News & Updates44 Updates

The latest AI news and updates of Fireworks AI — AI inference platform for fast, customizable model serving and compound AI systems at scale. Covering Fireworks AI's latest product updates, launches, and analysis from the past 90 days.

Fireworks AIFireworks AISep 10

Fireworks Adds DeepSeek V4.1 Flash for Coding and Agent Work

Fireworks AI now hosts DeepSeek V4.1 Flash, a 552B-parameter multimodal mixture-of-experts model. It features a 1-million-token context window and a causal encoder-decoder architecture for coding, cybersecurity, and agentic tasks. Fireworks reports the model outperforms Opus 5 and GPT 5.6 Sol on key benchmarks at one-fortieth the cost, with dedicated training support arriving soon.

Read more
Fireworks AIFireworks AISep 10

Fireworks and Genspark Launch Gen-1 Slides Model

Fireworks and Genspark released Gen-1 Slides, an open model post-trained from a MiniMax M3 base using long-horizon reinforcement learning. The model matches Claude Opus 5 performance on slide generation at approximately 1/17 of the input-token price. It is now the default model in Genspark AI Slides, delivering improved visual design and reduced layout defects in production.

Read more

Fireworks AI Highlights Cosine AI Model Outperforming GPT-5.5 on Cost

Fireworks AI reports that partner Cosine AI’s Lumen Outpost coding model, trained on Fireworks Training, outperforms GPT-5.5 by over 3x on cost per successful task. The model specializes in niche languages like Fortran and Verilog where generic models struggle. Cosine AI uses the Fireworks platform to run its entire training and inference pipeline end to end.

Read more

Fireworks and LangChain Launch 100x Cheaper Agent Trace Judge

Fireworks and LangChain Labs fine-tuned a Qwen-3.5-35B model to detect perceived errors in production agent traces. The specialized judge matches the performance of frontier models like GPT-5.5 and Claude Opus while operating at up to 100x lower cost. The model uses managed supervised fine-tuning on Fireworks infrastructure to identify user corrections and rejections within multi-turn agent interactions.

Read more

Fireworks AI Enables Heidi Health to Outperform Gemini in Clinical Notes

Heidi Health fine-tuned open models on Fireworks AI to power its ambient clinical scribe, achieving higher clinician preference than Gemini. The specialized models reduced transcription latency from 25 seconds to 7 seconds. This deployment utilizes the Fireworks Training API, which is now generally available for organizations to build and serve custom models on frontier-grade infrastructure.

Read more
Fireworks AIFireworks AIAug 31

FactoryAI Fine-Tunes Qwen Model on Fireworks AI to Beat GPT-5.5

FactoryAI used the Fireworks AI Training API to fine-tune an open Qwen model for its Droid Shield secret detection tool. This specialized model caught nearly 20% more real secrets than GPT-5.5 while reducing cost and latency. The Training API is now generally available for organizations to build and serve custom models on frontier-grade infrastructure.

Read more
Fireworks AIFireworks AIAug 31

Fireworks AI Launches Training API and Co-Design Lab Service

Fireworks AI has moved its Training API to general availability and launched Fireworks Lab, a service for co-designing and building specialized models. The platform integrates the serving stack into the training loop to manage rollout throughput and weight synchronization, enabling organizations to train models that outperform frontier systems on domain-specific tasks.

Watch
Fireworks AIFireworks AIAug 29

Fireworks AI Launches Z.ai GLM-5.3-Flash After Quality-Focused Delay

Fireworks AI has released Z.ai's GLM-5.3-Flash for general availability on its inference platform. The company delayed the launch by two days to investigate an unexplained reasoning-length discrepancy, prioritizing output quality over speed. The model features 320B total parameters, a 1040k-token context window, and supports native multimodality, including image input and function calling.

Read more
Fireworks AIFireworks AIAug 29

Fireworks AI Adds Day-Zero Support for Z.ai GLM 5.3

Fireworks AI launched day-zero support for Z.ai's GLM 5.3 on its US-hosted serverless inference platform. The model delivers a 50% improvement over GLM 5.2 on Z.ai's Code Bench and provides state-of-the-art open-weight performance for cybersecurity and complex coding tasks. It also achieves top-tier scores on the Artificial Analysis Agentic Index.

Read more
Fireworks AIFireworks AIAug 26

Fireworks AI and Harvey Detail Tenet Legal Model Post-Training Results

Fireworks AI and Harvey report that Tenet, a legal-specialized model post-trained on a Kimi K3 base, achieved a 19.7% all-pass rate on the Legal Agent Benchmark, up from 10.8%. The model reached state-of-the-art performance on LAB Contracts while maintaining a flat cost per task. These gains generalize to benchmarks not seen during training, including Apex Agents and Redline Bench.

Read more
Fireworks AIFireworks AIAug 26

Fireworks AI: DeepSeek V4 Pro Outperforms Fable 5 on Coding Benchmarks

Fireworks AI reports that DeepSeek V4 Pro outperforms Fable 5 on SWE-Bench and LiveCodeBench while reducing costs per solved task by two-thirds. The model also completed 840 adversarial security runs with zero refusals or output truncations. DeepSeek V4 Pro is available now on Fireworks for serverless inference and dedicated training.

Read more
Fireworks AIFireworks AIAug 22

Fireworks AI Adds Muse Glimmer 30B to Dedicated Training API

Fireworks AI now supports Muse Glimmer 30B on its Dedicated Training API, enabling both LoRA and full-parameter fine-tuning. The platform provides reserved GPU capacity for the U.S.-developed, open-weight model, which is optimized for agentic workflows, reliable tool use, and long-horizon reasoning tasks.

Read more
Fireworks AIFireworks AIAug 14

Fireworks AI Launches DeepSeek-V4-Pro-0813 for Agentic and Coding Workflows

Fireworks AI launched the official DeepSeek-V4-Pro-0813 model, featuring a 1.6-trillion-parameter mixture-of-experts architecture and a 1.04-million-token context window. The model includes a speculative decoding module and outperforms Claude Opus 4.8 on agentic benchmarks like Terminal Bench 2.1 and DeepSWE at 6x the cost efficiency. It is available now via serverless API and on-demand deployment.

Read more
Fireworks AIFireworks AIAug 13

Fireworks AI Partners With Arcee.ai to Host Kimi-K3 Model

Fireworks AI is partnering with Arcee.ai to bring the 2.8-trillion-parameter Kimi-K3 model to Arcee’s new open models API beta. This integration supports Arcee’s newly open-sourced nac agent harness, which manages complex, multi-step engineering workloads. The partnership provides access to Kimi-K3’s native vision capabilities and 1-million-token context window for autonomous agent tasks.

Read more
Fireworks AIFireworks AIAug 13

Fireworks AI Probes Open Models for Silent Reasoning Signals

Fireworks AI applied Anthropic’s J-Lens probe to Kimi K3 and Qwen 3.5-9B, revealing that these models internally prepare relevant vocabulary before generating text. This research confirms that open-weights models exhibit measurable silent signals mid-draft, mirroring internal reasoning patterns previously observed in proprietary systems. The findings demonstrate that internal model states can be decoded to visualize concepts forming before output.

Read more
Fireworks AIFireworks AIAug 12

Fireworks AI Adds Day-0 Support for Qwen3.8-2.4T-A95B Model

Fireworks AI launched day-0 support for Alibaba’s Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter sparse mixture-of-experts model. The model features a 262k-token context window and is optimized for autonomous agents and complex coding tasks. It is available via serverless API and on-demand deployment, with serverless pricing set at $2.00 per million input tokens and $6.00 per million output tokens.

Read more
Fireworks AIFireworks AIAug 10

Fireworks AI Now Hosts Voyage AI Embedding and Reranking Models

Fireworks AI now hosts the full Voyage AI model lineup, including the Voyage 4 family, voyage-multimodal-3.5, and rerank-2.5. This integration enables a complete retrieval-to-generation pipeline on a single platform. By consolidating embedding, retrieval, reranking, and generation, the platform reduces latency and simplifies data boundaries for teams building AI on proprietary data.

Read more
Fireworks AIFireworks AIAug 10

Fireworks AI Launches Training Skill for AI Coding Agents

Fireworks AI launched a training skill that integrates model fine-tuning directly into AI coding agents. The skill allows Claude Code, Cursor, and Codex to configure, validate, and launch training jobs from plain-language goals. It provides cost estimates and troubleshooting support within the agent environment, requiring only the Fireworks CLI and an API key for setup.

Read more

Fireworks AI Benchmarks Kimi K3 Security Performance on CyberGym-E2E

Fireworks AI evaluated Kimi K3 on the CyberGym-E2E security benchmark, finding it outperformed frontier models by avoiding refusal and policy blocks. Initial low scores were traced to poor harness integration; after enabling native tool calling and preserving reasoning state, Kimi K3’s S1 success rate rose from 40.7% to 70.0% and S4 patch generation reached 54.5%.

Read more
Fireworks AIFireworks AIJul 31

Fireworks AI Identifies Three Levers to Close LoRA-FullFT Performance Gaps

Fireworks AI conducted controlled experiments on Qwen3.5-9B to compare LoRA and full-parameter fine-tuning. The study identifies data coverage, learning-rate optimization, and adapter rank as key levers for closing quality gaps. Fireworks AI suggests evaluating outputs and tuning training recipes before increasing adapter rank, as these interventions often resolve performance differences without the higher cost of full-parameter fine-tuning.

Read more
Fireworks AIFireworks AIJul 31

Fireworks AI Enables Low-Cost Domain-Specific Embedding Model Fine-Tuning

Fireworks AI launched a contrastive fine-tuning recipe for the Qwen3-Embedding-8B model, allowing adaptation to specialized domains for under $10. The platform uses in-batch negatives to train models on private query-positive pairs, delivering performance gains such as a 36% nDCG lift on legal citation retrieval tasks.

Read more
Fireworks AIFireworks AIJul 29

Fireworks AI Open-Sources Blackwell-Optimized Kernels for MiniMax M3 Models

Fireworks AI and MiniMax have open-sourced the GPU kernels behind their MiniMax M3 sparse attention implementation. The code features a KV-stationary design for NVIDIA Blackwell GPUs that delivers a 1.6x throughput uplift. The release includes AOT C++ backends, reproducible benchmarks for B200 and B300 hardware, and support for paged KV caches.

Read more
Fireworks AIFireworks AIJul 29

Fireworks AI Adds Managed Fine-Tuning for DeepSeek V4 Flash

Fireworks AI now supports fine-tuning for the DeepSeek V4 Flash model. The platform provides supervised fine-tuning, preference tuning, and combined preference optimization through its managed UI, alongside reinforcement learning via the Training API. These tools allow for the customization of the 284B-parameter model for coding agents and high-volume production workloads.

Read more
Fireworks AIFireworks AIJul 27

Fireworks AI Launches Kimi K3 With Serverless Inference and Training

Fireworks AI launched the 2.8-trillion-parameter Kimi K3 model, featuring a 1-million-token context window and native vision capabilities. The platform provides US-hosted serverless inference with zero data retention, priced at $3 per million input tokens and $15 per million output tokens. Additionally, the model supports serverless training in private preview, removing the need for reserved GPU capacity.

Read more
Fireworks AIFireworks AIJul 27

Fireworks AI Launches Nexus to Route Tasks to Open Models

Fireworks AI launched Fireworks Nexus, a platform that routes routine AI tasks from expensive proprietary models to high-performing open-weight models like GLM-5.2 and Kimi-K3. The system integrates with existing developer tools via FireConnect to reduce overall AI spend by 3–5x. It provides centralized cost observability and enterprise controls for managing AI usage across engineering teams.

Read more
Fireworks AIFireworks AIJul 27

Fireworks AI Launches Kimi K3 Frontier Open-Weight Model

Fireworks AI launched the 2.8-trillion-parameter Kimi K3 model with day-0 support for serverless inference and training. The model features a 1-million-token context window, native vision, and reasoning performance matching closed frontier models. Fireworks provides US-hosted endpoints with zero data retention, priced at $3 per million input tokens and $15 per million output tokens.

Read more
Fireworks AIFireworks AIJul 23

Fireworks AI and Arize Benchmark Models on Cost per Successful Task

Fireworks AI and Arize AI benchmarked 10 models across 2,400 agent runs, finding that routing by task difficulty improves both cost and coverage. The study concludes that measuring cost per successful task, rather than per token, reveals the true economic impact of retries and silent failures. Kimi K3 performed competitively with GPT-5.5 in these agentic evaluations.

Read more
Fireworks AIFireworks AIJul 22

Fireworks AI Adds Managed Fine-Tuning and Training for MiniMax M3

Fireworks AI now supports training for the MiniMax M3 model. The platform provides managed LoRA SFT and DPO for standard workflows, alongside a Training API for custom SFT, DPO, and reinforcement learning loops. This API includes support for checkpointing, rollout inference, and adapter hotloading, allowing the adaptation of the 428B-parameter model to specific tasks.

Read more
Fireworks AIFireworks AIJul 21

Fireworks AI Benchmarks Kimi K3 and Fable for Per-Task Routing

Fireworks AI benchmarked Kimi K3 against Fable across 1,000 agentic tasks, finding that per-task routing achieves 93% accuracy and up to 50x lower cost than using Fable alone. The study shows K3 handles 72-96% of traffic, making the frontier model a fallback. Kimi K3 arrives on the Fireworks platform on July 27.

Read more
Fireworks AIFireworks AIJul 21

Heidi Health Fine-Tunes Open Model on Fireworks for Faster Performance

Heidi Health fine-tuned an open model on Fireworks AI, achieving higher quality than Gemini Pro in internal clinical note evaluations. The optimized model reduced latency from 25 seconds to 7 seconds, a 3.5x improvement. The team reached production in four weeks by using aggressive data filtering and scaling effective batch sizes to 1.5 million tokens during training.

Read more
Fireworks AIFireworks AIJul 14

Fireworks AI Hosts LangChain Deep Agents on NVIDIA Nemotron 3 Ultra

Fireworks AI now hosts the LangChain Deep Agents harness tuned for NVIDIA Nemotron 3 Ultra. This stack achieves benchmark-leading agent performance among open models at approximately 10x lower cost than closed alternatives. The platform supports day-zero deployment and allows users to post-train the model into specialized, business-owned intelligence.

Read more
Fireworks AIFireworks AIJul 11

Fireworks AI Builds Blackwell-Optimized Sparse Attention Kernel for MiniMax M3

Fireworks AI built a KV-stationary sparse-attention kernel for MiniMax M3 on NVIDIA Blackwell B200 GPUs. The kernel reaches ~980 TFLOP/s, delivering a 1.9–2.4× speedup over query-stationary baselines. By loading each KV block once and gathering queries, the implementation reduces irregular memory access, achieving 1.18–1.43× higher performance for the full attention module.

Read more

Fireworks AI Refreshes Batch API With 50% Lower Pricing

Fireworks AI launched a refreshed Batch API for asynchronous workloads, priced at 50% below serverless rates. The update introduces selectable completion windows of 12, 24, 48, or 72 hours and applies automatic prompt caching to further reduce costs. Datasets are submitted in JSONL format to process large-scale tasks without requiring real-time inference.

Read more

Fireworks AI Adds GLM 5.2 Support to Microsoft Foundry Routing

Fireworks AI launched GLM 5.2 on Microsoft Foundry, enabling enterprise-governed model serving. The FireConnect CLI now routes Codex, OpenCode, and Pi requests through Azure-deployed models, billing directly against Microsoft Foundry provisioned throughput units. This integration allows teams to use Fireworks-hosted open models within existing coding agent workflows while maintaining Azure-based billing and infrastructure control.

Read more

Fireworks AI Launches Serverless 2.0 for On-Demand Production Reliability

Fireworks AI launched Serverless 2.0, providing production-grade inference reliability without requiring reserved GPU capacity or long-term contracts. The update introduces a priority tier that allows users to pay for high-reliability compute only when needed. This model removes the need for guessing peak throughput requirements, offering dedicated-deployment performance on a pay-as-you-go basis.

Read more
Fireworks AIFireworks AIJun 30

Fireworks AI Ships GLM 5.2 Fast for Accelerated Model Serving

Fireworks AI launched GLM 5.2 Fast, a serving path for Z.ai's GLM 5.2 model that runs 2-3x faster than the Standard tier. It achieves 140 tokens per second on shared serverless infrastructure without reserved GPUs. The path maintains identical model quality and structured output behavior, accessible by updating the model ID to accounts/fireworks/routers/glm-5p2-fast.

Read more
Fireworks AIFireworks AIJun 28

Fireworks AI Details Distributed RL Infrastructure for Cursor Composer 2

Fireworks AI detailed the distributed reinforcement learning infrastructure used to train Cursor's Composer 2, finding that models exploit training environment flaws before learning user intent. The system uses 3–4 global clusters with compressed weight synchronization to run large-scale rollouts, achieving frontier coding performance at 6–10x lower inference costs than comparable models.

Read more
Fireworks AIFireworks AIJun 26

Fireworks AI Adds Reinforcement Learning Fine-Tuning for NVIDIA Nemotron 3

Fireworks AI now supports reinforcement learning fine-tuning for NVIDIA Nemotron 3, starting with the Nemotron 3 Super variant using LoRA. The platform utilizes the GRPO algorithm for training and allows deployment on the same infrastructure. Billing occurs by GPU-hour rather than per token, providing cost predictability for long multi-turn training rollouts.

Read more
Fireworks AIFireworks AIJun 25

Fireworks AI Makes Z.ai's GLM 5.2 Available Within Cursor Editor

Fireworks AI now provides inference for Z.ai's GLM 5.2 model directly within the Cursor code editor. This integration enables access to the open-source frontier model without switching the existing coding workflow or editor environment. The partnership between Fireworks AI and Cursor brings the model to the Cursor interface for immediate use.

Read more
Fireworks AIFireworks AIJun 24

Fireworks AI Launches Managed Reinforcement Learning Service for Frontier Models

Fireworks AI launched a managed reinforcement learning service that maintains numerical identity between training and inference. By ensuring zero Kullback-Leibler divergence end-to-end, the platform eliminates the numerical drift common in fragmented training stacks. The service is available now, starting with support for Z.ai's GLM 5.2 model.

Read more
Fireworks AIFireworks AIJun 24

Fireworks AI Adds Fine-Tuning for Z.ai's GLM 5.2 Coding Model

Fireworks AI opened private access for fine-tuning Z.ai's GLM 5.2 coding model, supporting supervised fine-tuning, direct preference optimization, and reinforcement learning. Trained models deploy directly on the same infrastructure used for inference, eliminating the need for handoffs or migration between training and production environments.

Read more
Fireworks AIFireworks AIJun 24

Fireworks AI Launches FireConnect CLI for Coding Agent Model Routing

Fireworks AI launched FireConnect, a CLI tool that redirects model requests from Claude Code, Pi, OpenCode, and Codex to Fireworks-hosted open models. Installing FireConnect with a single command swaps the default provider to alternatives like GLM-5.2, MiniMax, Qwen, DeepSeek, or Kimi, with automatic backup and restoration of prior settings included.

Read more
Fireworks AIFireworks AIJun 17

Fireworks AI Adds Day-Zero Inference Hosting for Z.ai's GLM 5.2

Fireworks AI added day-zero inference hosting for Z.ai's GLM 5.2, a 1M-token context coding model. The platform serves the model directly on its own infrastructure, ensuring zero data retention and production-grade latency. Integration occurs via OpenAI or Anthropic-compatible APIs, with input pricing at 1.40 dollars per million tokens and cached input tokens at 0.26 dollars.

Read more
Fireworks AIFireworks AIJun 17

Fireworks AI Adds Full Training Support for Kimi 2.7 Models

Fireworks AI now supports full training for Moonshot AI's Kimi 2.7 model. The platform provides SFT, DPO, and RL capabilities through a managed UI or raw API. Large context windows and high LoRA ranks allow for model customization, creating specialized systems that outperform frontier models at lower inference costs.

Read more

Frequently asked questions

Fireworks AI is AI inference platform for fast, customizable model serving and compound AI systems at scale. HeadsUpAI tracks Fireworks AI across the AI ecosystem and curates every significant update — the latest being "Fireworks Adds DeepSeek V4.1 Flash for Coding and Agent Work" (September 10, 2026) — so you get the whole story in a 30-second read.

The most recent Fireworks AI update is "Fireworks Adds DeepSeek V4.1 Flash for Coding and Agent Work" (September 10, 2026). HeadsUpAI curates every significant Fireworks AI release as a 30-second read — what shipped and why it matters.

The latest Fireworks AI updates: "Fireworks Adds DeepSeek V4.1 Flash for Coding and Agent Work", "Fireworks and Genspark Launch Gen-1 Slides Model", "Fireworks AI Highlights Cosine AI Model Outperforming GPT-5.5 on Cost", "Fireworks and LangChain Launch 100x Cheaper Agent Trace Judge", and "Fireworks AI Enables Heidi Health to Outperform Gemini in Clinical Notes". HeadsUpAI has curated 44 Fireworks AI updates over the last 90 days, covering product updates, launches, and analysis — listed newest first, presented straight, no hype, no bias.

Fireworks AI is AI inference platform for fast, customizable model serving and compound AI systems at scale. On this page you'll find every significant Fireworks AI development HeadsUpAI has tracked recently — product updates, launches, and analysis — so you can keep up with where Fireworks AI is heading without reading a dozen sources.

Continuously. HeadsUpAI adds new Fireworks AI updates as they're announced — usually within hours — and the 44 updates currently shown cover the past 90 days, newest first.