Arena

Arena AI News & Updates60 Updates

The latest AI news and updates of Arena — Community-driven AI model evaluation platform with Arena leaderboards spanning code, text, vision, and search. Covering Arena's latest analysis, product updates, and research from the past 90 days.

ArenaArena18h ago

Arena Study Finds Models Share 43% of Ideas on Average

Arena analyzed 30,086 model battles from May to September 2026, finding that models share 43.1% of their ideas on average. While adjacent generations show higher overlap, frontier models often draw on the same conceptual ground regardless of lab or country. Creative writing tasks produced the most diverse responses, while healthcare and legal topics showed the highest conceptual convergence.

ArenaArena21h ago

Arena Ranks Meta Muse Spark 1.3 Thirteenth in Agent Arena

Arena ranked Meta's Muse Spark 1.3 (Max) thirteenth on its Agent Arena leaderboard, achieving a 4.2% net improvement across 8,770 sessions. This result marks a 17-rank climb from the previous 1.2 (xHigh) version, with performance gains across code, work, and chat categories. The model operates at a median cost of $0.30 per task.

Read more
ArenaArenaSep 10

Arena Ranks Tencent Hy4 Preview #2 Among Open Agent Models

Arena ranks Tencent's Hy4 preview second among open models on its Agent Arena leaderboard, achieving a +5.4% net improvement across 14,500+ sessions. The model delivers performance at a median cost of $0.26 per task, roughly 68% lower than Kimi K3, and secures a position on the Agent Arena Pareto frontier.

Read more
ArenaArenaSep 10

Arena Ranks DeepSeek-V4.1-Flash Fourteenth on WebDev Code Leaderboard

Arena ranked DeepSeek-V4.1-Flash fourteenth overall on its Code Arena: WebDev leaderboard with an early AutoEval score of 1,620 points. The model ranks fourth among open-weight entries, outperforming previous DeepSeek-V4 variants by up to 40 points. It is priced at $0.30 per million input tokens and $1.20 per million output tokens, placing it among the most cost-efficient options.

Read more
ArenaArenaSep 10

Arena Reports Proprietary Models Reopen Web Development Frontier Gap

Arena reports that the proprietary-vs-open-source frontier gap on its Code Arena: WebDev leaderboard has widened to +120 points. After narrowing to roughly +10 points in early Q3 2026, the gap snapped back open following the releases of Claude Fable 5.1 and GPT-6 Astra, which pushed the proprietary frontier to nearly 1,800 points while open source held near 1,675.

Read more
ArenaArenaSep 10

Arena Offers Free Direct Access to Top-Ranked GPT-Image-2.5 Sunburst

Arena is providing free 72-hour access to OpenAI's GPT-Image-2.5 Sunburst in Direct Mode until September 13 at 8 a.m. PT. The model currently holds the number one ranking across Arena's Text-to-Image, Image Edit, and Multi-Image Edit leaderboards. After the access window closes, the model remains available for anonymous use in Battle and Agent modes.

Read more
ArenaArenaSep 9

Arena Analyzes Writing Style Shifts in Claude Fable 5.1

Arena analyzed tens of thousands of high-reasoning Text Arena outputs to measure how Claude Fable 5.1’s writing style differs from Fable 5. The model uses 58% fewer agreement openers, 45% less honesty-focused phrasing, and 32% fewer em dashes per 1,000 words. Meanwhile, median response length increased 30% to 414 words, while semicolon usage rose 63%.

Read more
ArenaArenaSep 9

Arena Ranks Meta Muse Spark 1.3 Max Eighth on WebDev Pareto Frontier

Arena ranks Meta’s Muse Spark 1.3 Max eighth on its Code Arena: WebDev leaderboard with 1,650 points. The model reshapes the Pareto frontier by filling the price gap between Qwen3.8-max and Qwen3.8-Flash-Next at $3.50 per million tokens. It performs within 20 to 24 points of higher-priced models while offering 30% to 70% cost savings.

Read more
ArenaArenaSep 8

Arena Ranks OpenAI GPT-Image-2.5 Models #1 and #2

Arena.ai reports that OpenAI's new GPT-Image-2.5-Sunburst and GPT-Image-2.5-Flare models have debuted at #1 and #2 across its Text-to-Image, Image Edit, and Multi-Image Edit leaderboards. Sunburst leads with a 40-to-81 point increase over the previous GPT-Image-2 (medium) across categories, while Flare secures the second spot with faster generation speeds and consistent point gains.

Read more
ArenaArenaSep 8

Arena Ranks OpenAI GPT-6 Astra (Max) Second on Agent Leaderboard

Arena.ai reports OpenAI's GPT-6 Astra (Max) debuted at #2 on its Agent Arena leaderboard, achieving a +12.5% net improvement across 8,900 real-world agentic sessions. The model scored +40.6% in user sentiment and +21.9% in confirmed task success. At a median cost of $4.01 per task, the release reshapes the leaderboard's cost-performance Pareto frontier.

Read more
ArenaArenaSep 6

Arena Ranks Anthropic Claude Fable 5.1 (Max) #1 on Agent Arena

Arena.ai ranks Anthropic's Claude Fable 5.1 (Max) first on its Agent Arena leaderboard, achieving a +15.8% net improvement across 6.7k+ real-world agentic sessions. The model also leads the price-performance frontier at a median cost of $4.14 per task. It shows a +42.5% lead in user praise and strong confirmed success rates, with no tool hallucinations.

Read more
ArenaArenaSep 5

Arena Ranks Grok Imagine Video 1.5 Agent Fifth on Leaderboard

Arena.ai ranked SpaceXAI's Grok Imagine Video 1.5 Agent fifth on its Text-to-Video leaderboard with 1,491 points. The model performs on par with Wan-3.0 and FLUX 3 Video, trailing by three points, while outperforming Dreamina Seedance-2.5 and MiniMax-H3. This debut places SpaceXAI among the top five labs in the arena.

Read more
ArenaArenaSep 5

Arena Ranks OpenAI GPT-6 Astra #1 on WebDev Leaderboard

Arena.ai ranks OpenAI's GPT-6 Astra (Max) first on its Code Arena: WebDev leaderboard with 1,797 points. The model leads the pack with a 35-point margin over Claude Fable 5.1 (Max) and a 180-point improvement over OpenAI's previous flagship. It also reshapes the Pareto frontier at $40 per million tokens, matching the latest Claude model pricing.

Read more
ArenaArenaSep 4

Microsoft MAI-Image-2.6 Exits Preview and Ranks Second on Arena

Microsoft released MAI-Image-2.6 on Foundry, where it now ties for second place on the Image Edit Arena with 1,439 points. The model generates images at an average cost of $0.048 per unit, placing it on the Pareto frontier. It also exits preview ranked second on the Text-to-Image leaderboard with 1,331 points and holds top-three category rankings across art and design.

Read more
ArenaArenaSep 3

Arena Ranks Alibaba Wan 3.0 Third on Image-to-Video Leaderboard

Arena.ai ranked Alibaba's Wan 3.0 third on its Image-to-Video leaderboard with 1,481 points. This result marks a 53-point improvement over the previous Wan 2.7 generation, which currently holds the tenth spot. The model achieved a 57% win-rate in community-driven blind evaluations, placing it 7 points behind Google's Gemini Omni 1.1 Flash and 16 points behind leader MiniMax-H3.

Read more
ArenaArenaSep 3

Arena Ranks Gemini 3.8 Flash (High) on Agent Pareto Frontier

Arena.ai reports that Google DeepMind's Gemini 3.8 Flash (High) has debuted on the Agent Arena Pareto frontier. The model achieves a +5.94% net improvement at a median cost of $0.22 per task, delivering performance comparable to Grok 4.5 and GLM 5.2 (Max) at a 44–50% lower price point. It is priced at $0.75/$3.75 per million input/output tokens.

Read more
ArenaArenaSep 2

Arena Ranks Anthropic Claude Fable 5.1 Max #1 on WebDev Leaderboard

Arena.ai ranks Anthropic's Claude Fable 5.1 (Max) first on its Code Arena: WebDev leaderboard with a score of 1,765. The model leads the pack with a 77-point margin over Qwen3.8-Max-0902 and 78 points over Claude Opus 5 (Max). At a blended $40 per million tokens, the release shifts the performance bar of the leaderboard's higher-priced Pareto frontier.

Read more
ArenaArenaSep 2

Arena Ranks Alibaba Qwen3.8-Max-0902 #1 on Code Arena WebDev

Arena’s Code Arena: WebDev leaderboard now ranks Alibaba’s Qwen3.8-Max-0902 at #1 with 1,691 points. The upgraded 2.4T-parameter model, priced at $2 per 1M input tokens and $6 per 1M output tokens on QwenCloud, scored 3 points above Claude Opus 5 (Max) and 17 points above Kimi K3 (Max).

Read more
ArenaArenaSep 1

Arena Ranks Alibaba Qwen3.8-Flash-Next Seventh Among Open Agentic Models

Arena ranked Alibaba’s Qwen3.8-Flash-Next seventh among open models on its Agent Arena leaderboard, achieving a +2.4% net improvement across 8,777 agentic sessions. The 125B-parameter MoE model, featuring 6B active parameters and a 262K native context window, also ranked fifth among open models for confirmed task success at +12.3%. It is priced at $0.16 per million input tokens.

Read more
ArenaArenaAug 31

Z.ai GLM-5.3-Flash Ranks 19th Overall on Agent Arena Leaderboard

Z.ai’s GLM-5.3-Flash joined the Agent Arena leaderboard, ranking #19 overall and #4 among open-source models. The model achieved a +4.56% net improvement at a $0.12 median cost per task, establishing a new Pareto frontier for agentic performance. It also ranks #4 in confirmed success with a +15.3% score.

Read more
ArenaArenaAug 29

Arena Ranks Tencent Hy4 Preview Fifth on WebDev Leaderboard

Arena ranked Tencent's Hy4 preview model fifth on its Code Arena: WebDev leaderboard with an early AutoEval score of 1,633. This result marks a 115-point improvement over the previous Hy3 generation, placing the 770B-parameter model third among open-weight entries. Arena notes this ranking is based on automated preference votes and will be updated as live human data accumulates.

Read more
ArenaArenaAug 28

Arena Ranks Alibaba Wan 3.0 First on Video Edit Leaderboard

Arena.ai ranked Alibaba's Wan 3.0 as the top model on its Video Edit Arena leaderboard. The model achieved a score of 1,414 points in community-driven blind evaluations, leading the category by 4 points. This debut marks the model's first appearance on the leaderboard, which currently tracks 10 models from 8 different labs.

Read more
ArenaArenaAug 27

Arena Ranks Grok-4.6 (xHigh) #15 Overall in Agent Arena

Arena ranked xAI's Grok-4.6 (xHigh) fifteenth overall in its Agent Arena, based on 4.5K agentic code sessions. The model achieved a #12 ranking in the Code category with a 7.7% net improvement and secured the #6 spot for Confirmed Success at 13.2%. It operates at a median cost of $1.12 per task.

Read more
ArenaArenaAug 27

Arena Ranks Google Gemini Omni 1.1 Flash #1 in Video Arena

Arena.ai ranked Google’s Gemini Omni 1.1 Flash first on its Text-to-Video leaderboard with 1,515 points, outperforming FLUX 3 Video. The model also debuted at second place on the Image-to-Video Arena with 1,488 points. These rankings reflect community-driven blind evaluations of the model’s video generation and editing capabilities.

Read more
ArenaArenaAug 26

Arena Integrates GitHub into Agent Mode for End-to-End Coding Workflows

Arena overhauled its Agent Mode architecture to support deep GitHub integration, transforming the interface into a full-featured development environment. The update adds a GitHub OAuth connector, sandbox-based cloning, a real-time diff review panel, and full Git lifecycle support. The agent now clones repositories, executes code, and pushes commits or pull requests directly to GitHub.

Read more
ArenaArenaAug 26

Arena Ranks Alibaba Qwen3.8-Flash-Next Eighth on WebDev Leaderboard

Arena ranked Alibaba’s Qwen3.8-Flash-Next eighth overall on its Code Arena: WebDev leaderboard with an AutoEval score of 1617. The 125B parameter model, which features 6B active parameters, ranks third among open-weight models. It is priced at $0.16 per million input tokens and $0.47 per million output tokens, serving as an early preview of the upcoming Qwen4 architecture.

Read more
ArenaArenaAug 26

Arena Ranks Z.ai GLM-5.3-Flash #5 on WebDev Leaderboard

Arena reports that Z.ai’s GLM-5.3-Flash, a 320B-parameter model with 18B active, has debuted at #5 on its Code Arena: WebDev leaderboard with a 1634 AutoEval score. Priced at $0.15/$0.5 per million tokens, the model ranks second among open-source entries, outperforming the larger GLM-5.3-Max variant and reshaping the leaderboard’s cost-performance Pareto frontier.

ArenaArenaAug 25

Arena Reports GPT-5.6 Sol and Luna Reshape Agent Pareto Frontier

Arena reports that OpenAI’s over-20% price cut on GPT-5.6 Sol has reshaped the Agent Arena Pareto frontier. Sol now ranks between Claude Fable 5 and Kimi K3 in Code and Work categories. Meanwhile, GPT-5.6 Luna (xHigh) joined the frontier overall, delivering performance gains across Code, Chat, and Work at costs ranging from $0.04 to $0.08 per task.

Read more
ArenaArenaAug 25

Arena Ranks Alibaba Qwen3.8-27B Ninth on WebDev Code Leaderboard

Arena ranked Alibaba’s Qwen3.8-27B ninth overall on its Code Arena: WebDev leaderboard with a score of 1595. The model is the only entry in its size class to reach the top 10, ranking six spots behind the larger Qwen3.8-Max. Arena reports the model reshapes the leaderboard’s cost-performance Pareto frontier at a price of $0.50 per million input tokens.

Read more
ArenaArenaAug 22

Arena Ranks Thinking Machines Inkling-Small #12 Among Open Models

Arena added Thinking Machines' Inkling-Small to its Agent Arena leaderboard, where it ranks #12 among open models with a -7.0% net-improvement score. The model features 12B active parameters and costs $0.09 per task. It leads the open-model category in Bash Recovery with an 11.0% net-improvement, outperforming larger models in CLI error recovery.

Read more
ArenaArenaAug 22

Arena Ranks Microsoft MAI-Image-2.6-Preview Third in Single Image Edit

Arena.ai ranked Microsoft’s MAI-Image-2.6-Preview third on its Single Image Edit leaderboard with a score of 1,420 points. This release improves upon the previous MAI-Image-2.5 model by 19 points, with notable performance gains in text rendering and commercial design. The model currently sits behind GPT Image 2 and Grok Imagine Image 2.0.

Read more
ArenaArenaAug 22

Agent Arena Adds New Categories for Code, Work, and Chat

Agent Arena launched new categories to measure model performance on long-horizon agentic tasks. GPT 5.6 Sol (xHigh) currently leads the Code category, while Claude Opus 5 (High) and (Max) rank first for Work and Chat, respectively. These rankings are based on millions of real-world agent sessions, reflecting performance across diverse, multi-step task environments.

Read more
ArenaArenaAug 22

Arena Ranks Dreamina Seedance-2.5 First on New Video Edit Leaderboard

Arena reports ByteDance’s Dreamina Seedance-2.5 now leads its new Video Edit leaderboard with 1,411 points, holding a 23-point margin over the next model. The model also ranks second in Image-to-Video with 1,484 points and fourth in Text-to-Video with 1,477 points. This release introduces native 1080p output with 10-bit color and enhanced texture rendering.

Read more
ArenaArenaAug 22

Code Arena Reports GLM-5.3 (Max) Shifts WebDev Pareto Frontier

Code Arena reports that Z.ai’s GLM-5.3 (Max) has shifted the Pareto frontier on its WebDev leaderboard, scoring 1597 points. At $3.65 per million tokens, the model ranks #2 among open models and #8 overall. It performs on par with Qwen3.8 (Max) while outperforming Gemini-3.7-flash-high and DeepSeek-v4-flash-high in cost-efficiency for web development tasks.

Read more
ArenaArenaAug 22

Arena Launches Pareto Frontier View to Compare Agent Model Cost-Efficiency

Arena launched a Pareto frontier view in Agent Arena, ranking models by net improvement against median cost per task. The frontier highlights models like Claude Opus 5 (High) at +12.34% for $1.78, Kimi K3 (Max) at +10.53% for $0.62, and GPT 5.5 (High) at +7.75% for $0.44. The view identifies models delivering the most performance for their price.

Read more
ArenaArenaAug 13

Arena Ranks Gemini 3.7 Flash High #8 in WebDev Arena

Arena.ai ranked Google DeepMind’s Gemini 3.7 Flash (High) eighth in its Code Arena: WebDev and ninth in the Text Arena. The model achieved scores of 1,588 and 1,490 respectively. It is priced at $0.75 per million input tokens and $3.75 per million output tokens.

Read more
ArenaArenaAug 13

Arena Ranks Gemini 3.7 Flash (High) #20 on Agent Leaderboard

Arena ranked Google DeepMind's Gemini 3.7 Flash (High) at #20 on its Agent Arena leaderboard. The model achieved a +3.4% net improvement score and secured the #5 spot for confirmed task success at +10%.

Read more
ArenaArenaAug 13

Gemini 3.7 Flash High Reshapes Arena Pareto Frontier for WebDev and Text

Arena.ai ranked Google DeepMind’s Gemini 3.7 Flash (High) on its leaderboards, where it reshapes the cost-performance Pareto frontier. The model achieved 1,588 points in the Code Arena: WebDev and 1,490 points in the Text Arena. It is priced at $0.75 per million input tokens and $3.75 per million output tokens.

Read more
ArenaArenaAug 13

Arena.ai Ranks MiniMax-H3 First Overall on Video Edit Leaderboard

Arena.ai ranked MiniMax-H3 first on its Video Edit Arena leaderboard with 1,390 points. This release leads both open and proprietary models, surpassing ByteDance’s Dreamina Seedance 2.0 and Google’s Gemini Omni Flash, which both scored 1,358 points. The model currently holds a 32-point lead over the next two best-performing systems in the category.

Read more
ArenaArenaAug 13

Arena Ranks Black Forest Labs FLUX 3 Video Fifth on Leaderboard

Arena.ai ranked Black Forest Labs' FLUX 3 Video fifth on its Image-to-Video leaderboard with 1,453 points. The model debuted in a tight cluster, trailing Google's Gemini Omni Flash and xAI's Grok Imagine Video 1.5. This ranking is based on community-driven blind battles measuring model performance in video generation.

Read more
ArenaArenaAug 13

Arena.ai Ranks DeepSeek-V4-Pro (Max) Using New AutoEval Methodology

Arena.ai ranked DeepSeek-V4-Pro (Max) eighth overall on its Code Arena: WebDev leaderboard with 1,607 points using its new AutoEval methodology. The model also placed fifth among open models in the Text Arena with 1,465 points. At $0.435 per million input tokens and $0.87 per million output tokens, the model demonstrates high price-to-performance efficiency.

Read more
ArenaArenaAug 12

Arena Analyzes Writing Style Shifts in Claude Opus 5 Models

Arena analyzed thousands of Claude Opus 5 outputs, finding responses are three times longer and more structurally complex than version 4.5. While sentence length and clause frequency increased, vocabulary became lexically simpler with fewer abstract nouns. Additionally, Opus 5 uses more em dashes and honesty-focused phrasing, while Fable 5 remains 38 percent more concise than Opus 5.

Read more
ArenaArenaAug 12

Arena Ranks Grok 4.6 High at #7 on WebDev Leaderboard

Arena ranked xAI's Grok 4.6 (High) seventh on its Code Arena: WebDev leaderboard with 1,618 points. This debut marks a significant performance jump from Grok 4.5, which holds the thirteenth spot with 1,553 points. The model now sits in a tight cluster with GPT-5.6 Sol xHigh and Claude Fable 5, separated by fewer than ten points.

Read more
ArenaArenaAug 12

Arena Ranks Upstage AI Solar Pro 4 Across Three Leaderboards

Arena ranked Upstage AI’s Solar Pro 4 on its Agent, Code: WebDev, and Text leaderboards, marking the first time a Korean lab has placed across all three. In the Agent Arena, the model debuted at #44 with a -12.10% net improvement score, showing a +0.3% tool hallucination rate while trailing in steerability and user sentiment signals.

Read more
ArenaArenaAug 11

Microsoft MAI-Image-2.6 Debuts at #2 on Arena Text-to-Image Leaderboard

Arena ranked Microsoft’s MAI-Image-2.6 second on its Text-to-Image leaderboard with 1,336 points, marking a significant performance jump from the previous MAI-Image-2.5 model. The release shows gains across all categories, including a #1 ranking in 3D Imaging and Modeling. Microsoft will make the model available on Playground and via early API access on Foundry next week.

Read more
ArenaArenaAug 10

Arena Reports Coding Model Performance Gap Narrows to 10 Points

Arena.ai reports that the performance gap between proprietary and open-source models on its Code Arena: WebDev leaderboard has narrowed to ~10 points, down from ~150 points in late 2025. The upcoming release of Muse Spark 1.2 weights and the arrival of the Glimmer model are expected to further impact the frontier, with official scores coming soon.

Read more
ArenaArenaAug 8

Arena Introduces Trace-and-Amplify to Detect Training-Time Reward Hacking

Arena introduces Trace-and-Amplify, a framework for collecting training-time reward-hacking trajectories in code generation. Monitors trained on this data achieve 90.16% detection accuracy on real hacks, significantly outperforming monitors trained on prompt-elicited data, which drop to 28% accuracy. The method uses unit-test tracers to identify evaluation-gaming during RL training, improving monitor generalization to unseen hacking types.

Read more
ArenaArenaAug 8

Arena Ranks Grok Imagine Image 2.0 (Low) Second on Leaderboards

Arena ranked xAI's Grok Imagine Image 2.0 (Low) second on its Text-to-Image and Image Edit leaderboards, with scores of 1,320 and 1,439 points respectively. This release marks a significant improvement over the previous Grok Imagine Image Quality model. The model is currently available only through the Grok app and is not accessible via API.

Read more
ArenaArenaAug 8

Arena Reports Diverging Token Efficiency Trends in Agent Arena Models

Arena reports that Anthropic’s Opus series increased token usage from 8.5k to 21k per task while improving performance. OpenAI’s GPT-5.6-Sol model improved performance by 1.5 percentage points while reducing token usage from 10k to 8k. These findings track the relationship between reasoning and output tokens across real-world agentic tasks.

Read more
ArenaArenaAug 8

Arena Ranks Meta Muse Spark 1.2 Fourth on Text Leaderboard

Arena ranked Meta’s Muse Spark 1.2 (xHigh) fourth on its Text Arena leaderboard with 1498 points. The model reshaped the cost-performance Pareto frontier at $1.25 per million input tokens and $4.25 per million output tokens. It also debuted at #11 on the Vision Arena, showing significant gains in multi-turn reasoning and hard prompt performance over the previous version.

Read more
ArenaArenaAug 5

Arena Adds Factuality-Weighted Leaderboard Toggle and Releases Model Truthfulness Analysis

Arena launched a factuality-weighted leaderboard toggle for its Text and Search Arenas, accompanied by an analysis of 2 million model claims. The findings show that while OpenAI models consistently improve in factuality over time, other providers often see scores decline when factuality is prioritized. The data reveals that human preference and factual accuracy are largely orthogonal signals.

ArenaArenaAug 5

Arena Ranks Claude Opus 5 (Max) First on Fullstack Leaderboard

Arena.ai ranked Anthropic's Claude Opus 5 (Max) first on its Fullstack Code Arena leaderboard with a score of 1,699. The model leads in multi-step reasoning, tool use, and end-to-end app generation based on 574 community-driven blind battles. Kimi K3 (Max) follows in second place with 1,660 points, while Claude Opus 5 (High) ranks third with 1,648 points.

Read more
ArenaArenaAug 4

Arena Ranks Qwen-Image-3.0-Pro Fifth on Text-to-Image Leaderboard

Arena.ai ranked Alibaba’s Qwen-Image-3.0-Pro fifth on its Text-to-Image leaderboard with a score of 1,263. This marks a significant improvement over the previous Qwen-Image-2.0-Pro, which held the fifteenth spot. The model achieved substantial gains in categorical performance, notably rising to fourth in the Cartoon category and sixth in both Commercial Design and Text Rendering.

Read more
ArenaArenaAug 4

Arena Ranks DeepSeek V4 Flash High #21 on Agent Arena

Arena ranked DeepSeek-V4-Flash-20260731 (High) #21 overall and #3 among open-source models in its Agent Arena, based on 12,500 real-world sessions. The model shows a +1.98% net improvement, with strong Confirmed Success (+6.99%) but weaker Steerability (-2.31%). It ranks 6 spots higher than the V4-Pro and 13 spots higher than the previous V4-Flash model.

Read more
ArenaArenaAug 4

Arena Ranks MiniMax-H3 as #1 Open Model in Video Arena

Arena ranks MiniMax's MiniMax-H3 as the #1 open model on its Video Arena, topping both Text-to-Video and Image-to-Video. The model scored 1,455 in Text-to-Video — tied for #3 overall, three points behind Meta's Muse Video — and leads the next open model, Tencent's hunyuan-video-1.5, by 280 points. The model is now publicly available.

Read more
ArenaArenaAug 3

Arena Shares Research on Valid Inference Using Synthetic Data

Arena shared a research talk by Stanford PhD candidate Carrie Tan presenting task exchangeability, a framework for valid statistical inference using synthetic data. By calibrating synthetic data against historical tasks where real data is available, the method improves coverage from 3% to 97% in social-science surveys and provides steerability scores for the Agent Arena leaderboard.

Watch
ArenaArenaAug 3

Arena Ranks Alibaba Qwen3.8-Max Fourth on Frontend Code Leaderboard

Arena.ai ranked Alibaba’s Qwen3.8-Max fourth on its Frontend Code Arena leaderboard with a score of 1,668. The model also secured the fifth spot in the Text Arena and second in the Vision Arena. Qwen3.8-Max is available at $2 per million input tokens and $6 per million output tokens, with open weights expected next week.

Read more
ArenaArenaAug 1

Arena Ranks AI Models on Image-to-WebDev Generation Capabilities

Arena released Image-to-WebDev leaderboard scores for 40 models, measuring performance on website generation from images and screenshots. Anthropic Claude Opus 5 (Max) leads the rankings with 1,669 points. OpenAI GPT-5.6 Sol (xHigh), Grok-4.5, Kimi K3 (Max), Muse Spark 1.1, and GPT-5.6 Terra and Luna variants also feature in the top 25.

Read more
ArenaArenaAug 1

Arena Ranks DeepSeek-V4-Flash-High #7 on Frontend Code Leaderboard

Arena.ai ranked DeepSeek-V4-Flash-High #7 overall in the Frontend Code Arena with a score of 1,586. The model reshaped the Pareto frontier, delivering the best performance-per-dollar in its class at $0.14 per million input tokens and $0.28 per million output tokens. This ranking marks a 154-point improvement over the previous preview version.

Read more
ArenaArenaJul 31

Arena Ranks OpenAI GPT-5.6 Sol, Terra, and Luna on Fullstack Leaderboard

Arena published Fullstack Code Arena scores for OpenAI’s GPT-5.6 family. Sol (xHigh) ranks #3 in Fullstack with 1,638 points, followed by Terra (xHigh) at #10 with 1,579 points and Luna (xHigh) at #14 with 1,568 points. Arena also announced that Claude Opus 5 (Max) will join the Fullstack leaderboard in an upcoming update.

Frequently asked questions

Arena is Community-driven AI model evaluation platform with Arena leaderboards spanning code, text, vision, and search. HeadsUpAI tracks Arena across the AI ecosystem and curates every significant update — the latest being "Arena Study Finds Models Share 43% of Ideas on Average" (September 11, 2026) — so you get the whole story in a 30-second read.

The most recent Arena update is "Arena Study Finds Models Share 43% of Ideas on Average" (September 11, 2026). HeadsUpAI curates every significant Arena release as a 30-second read — what shipped and why it matters.

The latest Arena updates: "Arena Study Finds Models Share 43% of Ideas on Average", "Arena Ranks Meta Muse Spark 1.3 Thirteenth in Agent Arena", "Arena Ranks Tencent Hy4 Preview #2 Among Open Agent Models", "Arena Ranks DeepSeek-V4.1-Flash Fourteenth on WebDev Code Leaderboard", and "Arena Reports Proprietary Models Reopen Web Development Frontier Gap". HeadsUpAI has curated 96 Arena updates over the last 90 days, covering analysis, product updates, and research — listed newest first, presented straight, no hype, no bias.

Arena is Community-driven AI model evaluation platform with Arena leaderboards spanning code, text, vision, and search. On this page you'll find every significant Arena development HeadsUpAI has tracked recently — analysis, product updates, and research — so you can keep up with where Arena is heading without reading a dozen sources.

Continuously. HeadsUpAI adds new Arena updates as they're announced — usually within hours — and the 96 updates currently shown cover the past 90 days, newest first.