Artificial Analysis

Artificial Analysis AI News & Updates60 Updates

The latest AI news and updates of Artificial Analysis — Independent AI benchmarking and analysis company evaluating AI models and API providers across quality, price, and performance. Covering Artificial Analysis's latest analysis, launches, and product updates from the past 90 days.

Artificial Analysis Ranks Ant Group's Ling-3.0-flash-VL on Intelligence Index

Artificial Analysis reports that Ant Group's Ling-3.0-flash-VL scores 25 on its Intelligence Index, landing on the Pareto frontier for intelligence versus active parameters. The 124B-parameter model, which activates 5.5B parameters per token, supports text, image, and video inputs. While efficient, the model shows limited factual recall and struggles with complex agentic tasks like terminal use.

Artificial Analysis: OpenAI GPT Image 2.5 Models Take Top Spots

Artificial Analysis reports that OpenAI’s new GPT Image 2.5 models, Flare and Sunburst, hold the top two spots on its image leaderboards. Flare reduced median generation time by 63% to 52.7 seconds compared to GPT Image 2. Both models maintain the $30 per million output token price, with Sunburst showing the largest performance gains in image editing capabilities.

Read more

Octen Search Debuts Third on Artificial Analysis Search API Index

Artificial Analysis benchmarked Octen Search, which debuts third on its Search Index with a score of 77. The tool achieves the fastest Time per Task at 16.9 seconds and a cost of $0.058 per task. It performs strongly on BrowseComp and DeepSearchQA benchmarks, placing just behind Perplexity Search variants on the leaderboard.

Read more

Artificial Analysis Updates Intelligence Index with New Efficient Model Benchmarks

Artificial Analysis reports that Claude Fable 5.1, Muse Spark 1.3, and GPT-6 Astra have pushed the Intelligence Index vs. Cost Pareto frontier forward. These models establish new benchmarks for efficient intelligence, offering improved performance relative to their task costs. This update reflects the latest shifts in the intelligence-to-cost landscape for frontier AI models.

Read more

Artificial Analysis Launches Model Release Pages for Effort Level Comparison

Artificial Analysis launched Model Release pages to compare intelligence, cost per task, and speed across different reasoning effort levels. Frontier models now ship with up to six effort variants, producing distinct performance profiles. The pages feature side-by-side comparisons and Capability Index scores, illustrated by GPT-6 Astra and Claude Fable 5.1 performance ranges across their respective effort configurations.

Read more

Artificial Analysis Intelligence Index v4.3 Updates Benchmarks and Model Rankings

Artificial Analysis released Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and replacing 𝜏³-Banking with AutomationBench-AA, a Zapier-collaboration benchmark using a 657-task private test set. The update increases private-evaluation weighting to 45%. Claude Fable 5.1 and GPT-6 Astra tie for the lead with a score of 53, while GLM-5.3 and Kimi K3 lead open-weights models at 44.

Read more

Artificial Analysis Benchmarks OpenBMB MiniCPM5-2B Reasoning Model Performance

Artificial Analysis benchmarked OpenBMB’s MiniCPM5-2B, a 2.6B dense reasoning model, scoring it 15 on the Intelligence Index v4.2. This result marks the highest score for any open-weights model under 4B parameters. The model demonstrates strong agentic performance and token efficiency, though it trails on knowledge and coding evaluations compared to larger models.

Read more

Artificial Analysis Updates Intelligence Index with New Agentic and Reasoning Evals

Artificial Analysis released Intelligence Index v4.2, an interim update adding the AA-Briefcase agentic knowledge-work evaluation and Surge AI’s GDP.pdf long-document reasoning test. The update drops the saturated GPQA Diamond benchmark, doubles private held-out test weighting to 40%, and re-ranks the frontier with Claude Fable 5.1 leading and GPT-6 Astra second.

Read more

Microsoft Releases MAI-Image-2.6-Flash and MAI-Image-2.6 on Foundry

Microsoft released MAI-Image-2.6-Flash and MAI-Image-2.6 on Microsoft Foundry today. Artificial Analysis ranks the Flash variant #3 in Image Editing and #8 in Text to Image, marking a significant performance gain over the previous generation at the same $19.50 per 1,000 images price point. It joins the flagship model on the quality-vs-price Pareto frontier for editing.

Read more

Artificial Analysis Ranks Meta Muse Image on Image Leaderboards

Artificial Analysis added Meta's Muse Image to its Image Arena, ranking it #4 in Image Editing and #5 in Text to Image. The model sits on the quality-vs-price Pareto frontier, costing $0.01 per image. It performs strongest in lighting and reasoning tasks, with particular effectiveness for social media and creator content workflows.

Read more

Artificial Analysis Benchmarks GPT-6 Astra Performance and Pricing

Artificial Analysis benchmarked OpenAI's GPT-6 Astra, finding it scores 67 on the Coding Agent Index—matching Claude Opus 5 and Fable 5—with 70% higher token efficiency than GPT-5.6 Sol. On the Intelligence Index, it ties its predecessor at 61, but a 2.5x price increase makes it 75% more expensive per task, despite halving the hallucination rate to 51%.

Read more

Artificial Analysis Benchmarks MBZUAI's K2 Horizon 375B A23B Model

Artificial Analysis benchmarked MBZUAI’s K2 Horizon 375B A23B, an open-weights Mixture-of-Experts model scoring 47 on its Intelligence Index. The model features 375B total parameters with 23B active and a 512K-token context window. It demonstrates strong agentic performance and a 26% hallucination rate, driven by abstaining from 60% of questions rather than guessing.

Read more

Artificial Analysis Ranks Alibaba Wan 3.0 #1 in Video Editing

Artificial Analysis ranks Alibaba's Wan 3.0 #1 on its Video Editing Leaderboard and #2 in Text to Video with Audio. The model, which supports multimodal references and instruction-led editing for 30-second 1080p clips, marks a generational improvement over Wan 2.7. It is available in public preview on Alibaba Cloud Model Studio, with pricing starting at $0.05 per second.

Read more

Artificial Analysis Benchmarks Meta Muse Spark 1.3 Coding Agent Performance

Artificial Analysis added Meta's Muse Spark 1.3 to its Coding Agent Index, with the limited-preview (max) variant scoring 68, ranking second only to Claude Opus 5. The generally available (xhigh) variant scores 64 and costs $1.72 per task, making it the most cost-efficient agent among those scoring above 60 on the index.

Read more

Artificial Analysis Ranks Inworld Realtime TTS-2 #1 in Controlled Voice

Artificial Analysis ranks Inworld’s new Realtime TTS-2 model #1 on its Controlled Voice Arena and #2 on its Provider Voice Arena. The model supports over 100 languages with on-the-fly switching and plain-text delivery instructions. It processes 106 characters per second and costs $20.83 per 1 million characters, based on blind user preference tests.

Read more

Artificial Analysis Expands Image Editing Arena with New Task Taxonomy

Artificial Analysis updated its Image Editing Arena to evaluate models across seven editing actions and ten real-world use cases. The benchmark now tests complex, multi-step edit instructions to reflect modern workflows. Microsoft’s MAI-Image-2.6-Preview currently leads the overall leaderboard, while OpenAI’s GPT Image 2 and ByteDance’s Seedream 5.0 Pro rank highest in specific categories like object-level and identity-preserving edits.

Read more

Artificial Analysis Benchmarks Google Gemini 3.8 Flash Intelligence and Cost

Google released Gemini 3.8 Flash, its fourth Flash model in four months. Artificial Analysis benchmarks the model at 59 on its Intelligence Index, a 3-point gain over Gemini 3.7 Flash. It reaches the Intelligence vs. Cost per Task Pareto frontier at $0.58 per task, with discounted pricing of $0.75 per million input tokens through the end of the year.

Read more

Artificial Analysis Benchmarks Apodex 1.1 Agentic Capabilities and Reliability

Artificial Analysis benchmarked Apodex 1.1, a new model scoring 44 on its Intelligence Index. The model demonstrates strong agentic performance, achieving an Elo of 1348 on GDPval-AA v2 and 70% on Terminal-Bench v2.1. However, it shows significant reliability trade-offs, with a 78.4% hallucination rate and 32% accuracy on knowledge evaluations. It costs approximately $0.05 per task.

Read more

Artificial Analysis Ranks Perplexity Search API Top on Search Index

Artificial Analysis benchmarked the Perplexity Search API, with all three context variants taking the top three leaderboard positions. Perplexity Search (medium) leads with a score of 80, outperforming previous leaders at 75. It also achieves the lowest model inference cost per task tested, extending the quality-cost Pareto frontier for search-enabled AI agents.

Read more

Artificial Analysis Benchmarks Agnes 2.5 Pro Beta Intelligence and Agentic Gains

Artificial Analysis benchmarked Agnes AI's Agnes 2.5 Pro Beta, which scores 49 on its Intelligence Index. The model shows a 9-point gain driven by agentic performance, though it requires double the output tokens of the previous version. Its knowledge reliability improvement stems from increased abstention, which reduced the hallucination rate but also halved factual accuracy.

Read more

Artificial Analysis Benchmarks Z AI's New GLM-5.3-Flash Model

Artificial Analysis benchmarked Z AI's GLM-5.3-Flash, a 320B-parameter model that scores 57 on the Intelligence Index. At $0.09 per task, the model reaches the Pareto frontier for intelligence and cost, costing roughly 7.5 times less than the larger GLM-5.3. It supports a 1-million-token context window and is available via Z AI's API.

Read more

Artificial Analysis Updates Coding Agent Index With Reward-Hacking Corrections

Artificial Analysis has updated its Coding Agent Index to v1.4, introducing reward-hacking corrections to the Terminal-Bench v2.1 benchmark. Any passing attempt found to be reward hacking—such as an agent fetching published solutions online—now receives a zero score. Measured reward-hacking rates vary widely by agent and model, ranging from 0.0% to 27.3% across 267 trials.

Read more

Artificial Analysis Ranks Breeze TTS 2 Top Open Weights Model

Artificial Analysis ranks BreezeBlue's Breeze TTS 2 as the leading open-weights text-to-speech model in its Provider Voices Speech Arena with an Elo of 1,215. While it outperforms the previous leader, Fish Audio S2 Pro, by 90 points, it trails on speed at 45 characters per second and costs $34 per 1 million characters on hosted endpoints.

Read more

Artificial Analysis Ranks South Korean AI Models in Global Top Tier

Artificial Analysis reports South Korea holds the global #3 position in AI development, with four domestic labs scoring above 30 on its Intelligence Index. Motif Technologies’ Motif 3 scored 47 and Upstage’s Solar Pro 4 scored 42, ranking as the highest-scoring models developed outside the United States and China. The domestic ecosystem continues to broaden through government-backed initiatives.

Read more

Artificial Analysis Ranks Microsoft MAI-Image-2.6-Preview #1 in Image Editing

Artificial Analysis ranks Microsoft's MAI-Image-2.6-Preview #1 on its Image Editing Leaderboard, displacing the previous leader, MAI-Image-2.5-Pro. The model also takes the #2 spot in Text to Image, leading 5 of 19 taxonomy categories including Material and Frontier. It is currently available in the MAI Playground and in Private Preview on Microsoft Foundry.

Read more

Artificial Analysis Benchmarks NVIDIA Groq 3 LPX at 3,431 Tokens/Second

Artificial Analysis measured 3,431 tokens per second on NVIDIA's new Groq 3 LPX inference rack using the Gemma 4 31B model at 100k context. NVIDIA announced the rack is now in full production and will enter operation later this year. The benchmark shows output speed remains consistent across 10k and 100k input lengths, indicating robust long-context inference performance.

Read more

Artificial Analysis Reports Results of South Korea's Sovereign AI Competition

Artificial Analysis, as official evaluation partner for South Korea’s Sovereign AI Foundation Model project, reports that Round 2 narrowed the field to three teams: Upstage (score 37), SK Telecom (35), and LG AI Research (31). Each team receives access to roughly 1,000 NVIDIA B200 GPUs for six months. The competition plans to select two final teams in early 2027.

Read more

Artificial Analysis and Liquid AI Benchmark Small Models on Mobile Devices

Artificial Analysis and Liquid AI launched intelligence and inference benchmarking for small models on mobile devices. Testing on the iPhone 17 Pro, Nanbeige4.2-3B and LFM2.5-2.6B tied for the top intelligence score of 63. LFM2.5-2.6B leads in efficiency, answering prompts in 8.0 seconds using 2.3 GB of memory, while six models define the speed-intelligence Pareto frontier.

Read more

Artificial Analysis Launches MLCR-AA Leaderboard for Medical Reasoning Models

Artificial Analysis launched the MLCR-AA leaderboard, evaluating AI models on Wisedocs' medical long-context reasoning benchmark. Claude Fable 5 leads with a 64.4% score. The evaluation reveals that while models are largely accurate, they struggle with completeness, often omitting essential details. Leading models cost between $0.30 and $1 per task, while the median model scores below 15%.

Read more

Artificial Analysis Launches Speech Agent Arena for Voice Model Evaluation

Artificial Analysis launched the Speech Agent Arena to evaluate speech-to-speech models through blind human conversations. Gemini 3.1 Flash Live Preview (Minimal) leads preference at 1,046 Elo, while Grok Voice Think Fast 2.0 High leads task success at 94.7%. Results show a divergence: the preference leader completes only 74.6% of tasks, proving natural conversation does not guarantee successful tool use.

Read more

Artificial Analysis Ranks Alibaba Qwen-Image-3.0 Models on Leaderboards

Artificial Analysis ranks Alibaba's Qwen-Image-3.0-Pro at #6 for image editing and #9 for text-to-image generation. The flagship model gains 83 and 48 Elo points over its predecessor, while the faster Qwen-Image-3.0 lands at #11 and #15. Both models are available via Alibaba Cloud Model Studio, with pricing starting at $0.03 per 1,000 images.

Read more

Artificial Analysis Launches Endpoint Accuracy Index to Benchmark Provider Model Quality

Artificial Analysis launched the Endpoint Accuracy Index to measure how much accuracy API providers preserve for open-weight models. By comparing provider endpoints against a self-hosted reference, the index reveals quality variations from 73% to 100%. These differences stem from quantization, KV-cache compression, and context limits, which often remain hidden on standard pricing pages.

Artificial Analysis Ranks Microsoft MAI-Image-2.5-Pro #1 in Image Editing

Artificial Analysis ranks Microsoft’s MAI-Image-2.5-Pro first on its Image Editing Leaderboard and seventh in Text to Image. The model, which launched in preview on Microsoft Foundry on July 23, costs approximately $108.50 per 1,000 1024x1024 images. It currently holds an Elo score of 1,272 based on blind user votes in the Artificial Analysis Image Arena.

Read more

Artificial Analysis Benchmarks Z AI's GLM-5.3 at 60 Intelligence Score

Artificial Analysis reports that Z AI's GLM-5.3 scores 60 on the Intelligence Index, tying Kimi K3 for the open-weights lead. The model achieves a 1770 Elo on the GDPval-AA v2 agentic evaluation, a 246-point jump. While token usage increased 20% over its predecessor, the model remains 19% cheaper per task than Kimi K3. Weights are expected within a week.

Read more

Artificial Analysis Launches Search Index to Benchmark AI Agent Search APIs

Artificial Analysis launched the Search Index, a benchmark measuring search API performance for AI agents on quality, cost, and speed. Testing 11 results across seven providers using GPT-5.6 Luna, the index shows Parallel Search (advanced), Exa Search (auto), and Firecrawl Search leading with scores of 75, 74, and 73, significantly outperforming the 33-point model-only baseline.

Read more

Artificial Analysis Ranks Cartesia Sonic 3.6 #1 on Speech Leaderboards

Artificial Analysis ranks Cartesia's Sonic 3.6 as the top model on both its Provider Voice and Controlled Voice Speech Arena leaderboards. The model achieved an Elo of 1,286 on the Provider leaderboard and 1,144 on the Controlled leaderboard. It generates 136.1 characters per second at a price of $49 per 1 million characters.

Read more

Artificial Analysis Benchmarks DeepSeek V4 Pro 0813 Intelligence and Pricing

Artificial Analysis reports that DeepSeek V4 Pro 0813 scores 53 on its Intelligence Index, an 8-point gain over the April preview. The model improves agentic performance and token efficiency by 30%, but a 264% blended price increase narrows its lead on the cost-performance frontier. It remains the second-most intelligent open-weights model, trailing only Moonshot AI’s Kimi K3.

Read more

Artificial Analysis Benchmarks Gemini 3.7 Flash at Pareto Frontier for Speed and Intelligence

Google released Gemini 3.7 Flash, which Artificial Analysis benchmarks show delivers a 4-point intelligence gain over version 3.6 and reaches the Pareto frontier for speed-to-intelligence efficiency. The model produces 340 output tokens per second and leads agentic benchmarks like AA-AnalystAgent with a 60% pass rate. Google offers discounted pricing of $0.75 per 1M input tokens through year-end.

Read more

Artificial Analysis Launches Optima for Custom AI Workload Benchmarking

Artificial Analysis launched Optima, a platform for building custom benchmarks based on specific tasks and datasets. The tool tracks performance, cost, and time efficiency across frontier models. Evaluation options include objective rubric grading at 0.25 dollars per criterion or pairwise judging at 0.75 dollars per match to identify the most efficient model for a specific use case.

Read more

Artificial Analysis: xAI Grok 4.6 Joins the Intelligence Frontier

Artificial Analysis reports that xAI's Grok 4.6 scores 61 on its Intelligence Index, placing it alongside GPT-5.6 Sol at the frontier. The model maintains its $2/$6 per 1M token pricing while delivering strong agentic performance, including an Elo of 1753 on GDPval-AA v2. It features a 500k-token context window and a $0.50 per 1M token cache hit rate.

Read more

Artificial Analysis Launches AA-AnalystAgent Benchmark for Quantitative AI Agents

Artificial Analysis launched AA-AnalystAgent, a benchmark testing AI agents on quantitative analysis using real-world spreadsheets and documents. Claude Opus 5 leads with a 54% pass⁵ score, which measures consistency across five independent attempts. The results show that reliability, rather than raw capability, separates top models, with early interpretation errors causing 57% of failures across the leaderboard.

Read more

NVIDIA Releases Nemotron 3.5 Lightning Reasoning Model

NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter Mixture-of-Experts model with 3.6B active parameters. Artificial Analysis benchmarks show an Intelligence Index score of 24, tying gpt-oss-120b, alongside significant agentic performance gains. The model delivers output speeds of nearly 670 tokens per second and is available now under an OpenMDW-1.1 license for high-volume agentic workflows.

Read more

Artificial Analysis Refreshes Text to Image Arena With Expanded Evaluation Metrics

Artificial Analysis updated its Text to Image Arena to evaluate models across 10 real-world use cases and 9 capabilities. The new methodology uses human-curated, monthly-refreshed prompts to track the frontier. GPT Image 2 leads all categories, while Nano Banana 2 and Nano Banana Pro emerge as cost-effective specialists for UI/UX, retail, and film production.

Read more

Artificial Analysis: Meta Muse Spark 1.2 Hits Intelligence-Cost Pareto Frontier

Artificial Analysis reports that Meta’s Muse Spark 1.2 lands on the Intelligence Index vs. Cost per Task Pareto frontier. The model scores 57, placing it six points below Claude Opus 5 while costing roughly one-sixth as much per task. While per-token pricing remains unchanged, heavier token usage on agentic tasks increased the total cost per task compared to version 1.1.

Read more

Artificial Analysis Benchmarks Alibaba Qwen3.8 Max Intelligence and Cost

Artificial Analysis reports that Alibaba’s Qwen3.8 Max scores 56 on its Intelligence Index at $1.14 per task. The 2.4-trillion-parameter model shows gains in agentic and coding benchmarks but trails the open-weights leader Kimi K3 by one point. Alibaba plans to release the model weights next week, marking a shift for its Max-class models.

Read more

Artificial Analysis Maps Intelligence Against Model Time per Task

Artificial Analysis benchmarked the tradeoff between model intelligence and task completion time. All frontier models finishing tasks in under two minutes are from OpenAI, while four of the six models taking longer are from Anthropic. Moonshot AI’s Kimi K3 (max) is an outlier, requiring about 11 minutes per task due to slower output speeds.

Read more

Artificial Analysis Launches Endpoint Accuracy Index for Open Weights Models

Artificial Analysis launched an Endpoint Accuracy Index to measure how serverless API providers preserve model accuracy. Benchmarks for GLM-5.2, gpt-oss-120b, and DeepSeek V4 Pro reveal that serving choices like output token limits and tool-call parsing often degrade performance. While DeepSeek V4 Pro endpoints mostly match reference parity, others show significant accuracy losses compared to self-hosted deployments.

Artificial Analysis Ranks DeepSeek V4 Flash 0731 Top Three Open Weights

Artificial Analysis reports that DeepSeek released the weights for DeepSeek V4 Flash 0731 under an MIT license. The 284B-parameter model scores 50 on the Artificial Analysis Intelligence Index, placing it among the top three open-weights models. It features 13B active parameters and is now available via DeepSeek’s API, Fireworks AI, and Ollama.

Read more

Artificial Analysis Ranks Mureka V9 Second on Music Leaderboards

Artificial Analysis ranks Mureka V9 second on its Instrumental and Vocals music leaderboards. In the Instrumental category, the model debuts within one Elo point of the top-ranked Suno V5.5. Mureka V9 generates full songs with multi-language lyrics and supports splitting tracks into 12 stems, featuring section-by-section direction of emotion, instrumentation, and BPM.

Read more

Artificial Analysis Benchmarks DeepSeek V4 Flash 0731 Intelligence and Efficiency

Artificial Analysis benchmarked DeepSeek V4 Flash 0731, which scores 50 on its Intelligence Index—a 10-point increase over the previous version. The model achieves significant agentic performance gains and a 12-point reduction in hallucination rates. With unchanged pricing and a 98% cache-hit discount, the model delivers per-task costs approximately 60% lower than comparable frontier models.

Read more

OpenAI Cuts GPT-5.6 Luna and Terra Pricing by 20-80%

Artificial Analysis reports that OpenAI reduced GPT-5.6 Luna pricing by 80% to $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Terra pricing fell 20% to $2.00 per million input tokens and $12.00 per million output tokens. Despite these cuts, Terra remains behind Luna and Sol on the price-performance frontier.

Read more

Artificial Analysis Ranks MiniMax H3 #1 in Video Editing

Artificial Analysis ranks MiniMax H3 first on its Video Editing Leaderboard with an Elo of 1,130. The model also places in the top three for text-to-video and image-to-video generation. MiniMax plans to release the model weights under a community license, which would establish it as the leading open-weights video model.

Read more

Artificial Analysis Details Hardware Requirements for Local Kimi K3 Deployment

Artificial Analysis reports that Moonshot AI’s 2.8-trillion-parameter Kimi K3 model requires approximately 1.56 TB of memory for its weights alone. The analysis finds that only 8x NVIDIA B300/GB300 or 8x AMD MI350X/MI355X nodes can support local deployment, as all earlier single-node configurations lack sufficient memory. Running the full 1M-token context requires roughly 1.59 TB of VRAM.

Read more

Artificial Analysis Benchmarks Thinking Machines' New Inkling Small Reasoning Model

Artificial Analysis benchmarked Thinking Machines' Inkling Small, which scores 40 on the Intelligence Index, nearly matching its flagship sibling while using less than one-third of the active parameters. The model outperforms the flagship on coding and frontier reasoning benchmarks but trails on agentic tasks and factual knowledge. It demonstrates high token efficiency, averaging 24,000 output tokens per task.

Read more

Artificial Analysis Benchmarks Agnes AI Agnes 2.5 Pro Alpha Model

Artificial Analysis benchmarked Agnes AI’s new Agnes 2.5 Pro Alpha, which debuts at 39 on the Intelligence Index. The reasoning model features a 1M-token context window and supports text, image, and video inputs. At $0.45 per 1M input tokens and $0.90 per 1M output tokens, it offers competitive coding performance for its price tier.

Read more

Artificial Analysis Benchmarks SpaceXAI Grok Voice Think Fast 2.0

Artificial Analysis benchmarked SpaceXAI’s Grok Voice Think Fast 2.0, which debuts at #2 on its Speech to Speech Index with an 82.9% score. The model achieves a 0.70-second Time to First Audio, the fastest among top-five models, and leads agentic performance at 56.5%. It is priced at $4.80 per hour of input audio.

Read more

Artificial Analysis Benchmarks New OpenAI GPT Transcribe and Live Models

Artificial Analysis reports OpenAI released GPT Transcribe, a batch speech-to-text model scoring 3.31% on AA-WER. The model improves accuracy by 0.7 percentage points over its predecessor while reducing price by 25% to $4.50 per 1,000 minutes. It now accepts contextual prompts, keywords, and language hints. OpenAI also launched GPT-Live-Transcribe, a streaming model currently undergoing benchmarking.

Read more

Artificial Analysis Ranks Alibaba Qwen Audio 3.0 Realtime #1

Artificial Analysis benchmarked Alibaba’s Qwen Audio 3.0 Realtime, finding the Plus variant leads its Speech to Speech Index at 84.1%. The model outperforms GPT-Realtime-2.1 High on speech reasoning, conversational dynamics, and agentic performance. However, it records a 4.02-second time to first audio, significantly slower than the ~1.1-second latency achieved by OpenAI’s GPT-Realtime-2 series.

Read more

Artificial Analysis: Moonshot AI Releases Kimi K3 Open Weights

Artificial Analysis reports that Moonshot AI released the weights for Kimi K3, a 2.6T parameter model. It now leads the open-weight category with a score of 57 on the Artificial Analysis Intelligence Index. The release includes a custom license requiring separate agreements for high-revenue businesses and mandatory UI attribution for products exceeding 100 million monthly active users.

Read more

Artificial Analysis: Claude Opus 5 Leads Agentic Knowledge Work Benchmark

Artificial Analysis benchmarked Anthropic’s Claude Opus 5 on its AA-Briefcase agentic knowledge work test, where it took the top spot with a 1720 Elo score. The model outperforms Claude Fable 5 in analytical quality while reducing task costs by 20%. While Opus 5 leads in analytical rigor, it trails GPT-5.6 Sol in presentation quality and requires longer task times.

Read more

Frequently asked questions

Artificial Analysis is Independent AI benchmarking and analysis company evaluating AI models and API providers across quality, price, and performance. HeadsUpAI tracks Artificial Analysis across the AI ecosystem and curates every significant update — the latest being "Artificial Analysis Ranks Ant Group's Ling-3.0-flash-VL on Intelligence Index" (September 11, 2026) — so you get the whole story in a 30-second read.

The most recent Artificial Analysis update is "Artificial Analysis Ranks Ant Group's Ling-3.0-flash-VL on Intelligence Index" (September 11, 2026). HeadsUpAI curates every significant Artificial Analysis release as a 30-second read — what shipped and why it matters.

The latest Artificial Analysis updates: "Artificial Analysis Ranks Ant Group's Ling-3.0-flash-VL on Intelligence Index", "Artificial Analysis: OpenAI GPT Image 2.5 Models Take Top Spots", "Octen Search Debuts Third on Artificial Analysis Search API Index", "Artificial Analysis Updates Intelligence Index with New Efficient Model Benchmarks", and "Artificial Analysis Launches Model Release Pages for Effort Level Comparison". HeadsUpAI has curated 106 Artificial Analysis updates over the last 90 days, covering analysis, launches, and product updates — listed newest first, presented straight, no hype, no bias.

Artificial Analysis is Independent AI benchmarking and analysis company evaluating AI models and API providers across quality, price, and performance. On this page you'll find every significant Artificial Analysis development HeadsUpAI has tracked recently — analysis, launches, and product updates — so you can keep up with where Artificial Analysis is heading without reading a dozen sources.

Continuously. HeadsUpAI adds new Artificial Analysis updates as they're announced — usually within hours — and the 106 updates currently shown cover the past 90 days, newest first.