Artificial Analysis Ranks Nemotron 3 Ultra Fastest for Agentic Tasks

Artificial AnalysisArtificial Analysis

· Updated

Artificial Analysis evaluated NVIDIA's newly launched Nemotron 3 Ultra, finding it completes agentic tasks significantly faster than peers due to high inference speed. The model achieves competitive performance on Terminal-Bench v2.1, positioning it as a leading option for efficient autonomous AI workflows.

Independent evaluator Artificial Analysis reports that NVIDIA Nemotron 3 Ultra leads in speed for agentic tasks. Testing on Terminal-Bench v2.1 shows the model completes tasks faster than peers due to high inference speed (token generation rate) while scoring competitively. This evaluation positions NVIDIA Nemotron 3 Ultra on the Pareto frontier for performance versus time.
Artificial Analysis Intelligence Index Score
47.7
Output Tokens Per Second
>400
Total Parameters
~550 billion
Active Parameters
55 billion
AA-Omniscience Non-Hallucination Score
71%
GDPval-AA Elo
1378

This speed is critical for AI agents (autonomous systems using multi-step reasoning), as multi-turn interactions benefit from low latency. The model's ability to maintain competitive performance while being the fastest across all tested turn limits highlights its efficiency for complex workflows.

Nemotron 3 Ultra is also the most intelligent US open-weights model, scoring 47.7 on the Artificial Analysis Intelligence Index. With output speeds exceeding 400 tokens per second, it combines frontier intelligence with high throughput.

Artificial Analysis
Artificial Analysis
@ArtificialAnlys
X

Nemotron 3 Ultra was launched today, including a focus on low latency agentic performance. We tested it against peers under restricted turn-usage limits on Terminal-Bench v2.1 - @NVIDIA Nemotron 3 Ultra completes tasks at a much faster pace than peers due to its high inference speed while scoring competitively on the benchmark. In this analysis each model is given a ‘turn limit’ within which it can complete tasks, inside a customized version of the Terminus 2 harness which advises it of this limit. We apply 4 increasing turn limits and trace each result’s tradeoff of task latency and performance. Time per task, on the X axis, is calculated as decode time based on token usage and measured endpoint output speeds (for Nemotron 3 Ultra, speeds were measured on a pre-release deployment on @blackboxai), plus the actual time spent executing tools to complete the benchmark. Nemotron 3 Ultra is the fastest across all turn limits and sits on the Pareto frontier for performance versus time per task for this evaluation.

9retweets100likes
View on X

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update