Nemotron 3 Ultra is available on AI Gateway. The largest US open weight model with 30% lower cost for agentic tasks. πππππ: 'ππππππ/ππππππππ-πΉ-πππππ-π»π»πΆπ-ππ»π»π' https://t.co/wprOtqYVkY
Vercel AI Gateway Adds NVIDIA Nemotron 3 Ultra for Cost-Efficient Agents
VercelΒ· Updated
Vercel has integrated NVIDIA Nemotron 3 Ultra into its AI Gateway, making the large US open-weight model available for developers. This provides a cost-effective option for building and orchestrating complex, multi-step AI agent workflows, with up to 30% lower costs for agentic tasks.
- Model Identifier
- `nvidia/nemotron-3-ultra-550b-a55b`
- Context Window
- 1M tokens
- Cost Reduction
- Up to 30% lower for agentic tasks
- Throughput
- Up to 350 tokens per second
- Architecture
- Open Mixture-of-Experts reasoning model
- Pricing
- No markup, no platform fee on inference
The integration offers up to 30% lower cost for agentic tasks. As the largest US open-weight model, it delivers up to 350 tokens per second for high-throughput inference.
Developers can now use Nemotron 3 Ultra via the Vercel AI SDK by setting the model to nvidia/nemotron-3-ultra-550b-a55b. The AI Gateway (a unified API for calling models) offers usage tracking, dynamic provider sorting, and no markup on provider pricing.
Still wondering? A few quick answers below.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards β





