Vercel AI Gateway Adds NVIDIA Nemotron 3 Ultra for Cost-Efficient Agents

VercelVercel

Β· Updated

Vercel has integrated NVIDIA Nemotron 3 Ultra into its AI Gateway, making the large US open-weight model available for developers. This provides a cost-effective option for building and orchestrating complex, multi-step AI agent workflows, with up to 30% lower costs for agentic tasks.

Vercel has added NVIDIA Nemotron 3 Ultra to its AI Gateway. This Mixture-of-Experts (MoE) (an architecture using multiple specialized sub-networks) reasoning model is designed for orchestrating long-running agent workflows with a 1M token context window.
Model Identifier
`nvidia/nemotron-3-ultra-550b-a55b`
Context Window
1M tokens
Cost Reduction
Up to 30% lower for agentic tasks
Throughput
Up to 350 tokens per second
Architecture
Open Mixture-of-Experts reasoning model
Pricing
No markup, no platform fee on inference

The integration offers up to 30% lower cost for agentic tasks. As the largest US open-weight model, it delivers up to 350 tokens per second for high-throughput inference.

Developers can now use Nemotron 3 Ultra via the Vercel AI SDK by setting the model to nvidia/nemotron-3-ultra-550b-a55b. The AI Gateway (a unified API for calling models) offers usage tracking, dynamic provider sorting, and no markup on provider pricing.

Vercel Developers
Vercel Developers
@vercel_dev
X

Nemotron 3 Ultra is available on AI Gateway. The largest US open weight model with 30% lower cost for agentic tasks. πš–πš˜πšπšŽπš•: 'πš—πšŸπš’πšπš’πšŠ/πš—πšŽπš–πš˜πšπš›πš˜πš—-𝟹-πšžπš•πšπš›πšŠ-πŸ»πŸ»πŸΆπš‹-πšŠπŸ»πŸ»πš‹' https://t.co/wprOtqYVkY

1retweets25likes
View on X

Still wondering? A few quick answers below.

NVIDIA Nemotron 3 Ultra is an open Mixture-of-Experts (MoE) reasoning model. It is designed for orchestrating complex, multi-step AI agent workflows, including planning, tool use, sub-agent delegation, and error recovery. It features a 1M token context window.

Using Nemotron 3 Ultra on Vercel AI Gateway offers up to 30% lower cost for agentic tasks and high throughput of up to 350 tokens per second. The AI Gateway also provides a unified API for model calls, usage tracking, and performance optimizations without additional markup.

You can access NVIDIA Nemotron 3 Ultra by setting the model to `nvidia/nemotron-3-ultra-550b-a55b` within the Vercel AI SDK. This allows developers to integrate the model into their applications and leverage the AI Gateway's features.

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards β†’

Share this update