Nemotron 3 Ultra was launched today, including a focus on low latency agentic performance. We tested it against peers under restricted turn-usage limits on Terminal-Bench v2.1 - @NVIDIA Nemotron 3 Ultra completes tasks at a much faster pace than peers due to its high inference speed while scoring competitively on the benchmark. In this analysis each model is given a ‘turn limit’ within which it can complete tasks, inside a customized version of the Terminus 2 harness which advises it of this limit. We apply 4 increasing turn limits and trace each result’s tradeoff of task latency and performance. Time per task, on the X axis, is calculated as decode time based on token usage and measured endpoint output speeds (for Nemotron 3 Ultra, speeds were measured on a pre-release deployment on @blackboxai), plus the actual time spent executing tools to complete the benchmark. Nemotron 3 Ultra is the fastest across all turn limits and sits on the Pareto frontier for performance versus time per task for this evaluation.
Artificial Analysis Ranks Nemotron 3 Ultra Fastest for Agentic Tasks
Artificial Analysis· Updated
Artificial Analysis evaluated NVIDIA's newly launched Nemotron 3 Ultra, finding it completes agentic tasks significantly faster than peers due to high inference speed. The model achieves competitive performance on Terminal-Bench v2.1, positioning it as a leading option for efficient autonomous AI workflows.
- Artificial Analysis Intelligence Index Score
- 47.7
- Output Tokens Per Second
- >400
- Total Parameters
- ~550 billion
- Active Parameters
- 55 billion
- AA-Omniscience Non-Hallucination Score
- 71%
- GDPval-AA Elo
- 1378
This speed is critical for AI agents (autonomous systems using multi-step reasoning), as multi-turn interactions benefit from low latency. The model's ability to maintain competitive performance while being the fastest across all tested turn limits highlights its efficiency for complex workflows.
Nemotron 3 Ultra is also the most intelligent US open-weights model, scoring 47.7 on the Artificial Analysis Intelligence Index. With output speeds exceeding 400 tokens per second, it combines frontier intelligence with high throughput.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →




