Together AI Adds NVIDIA Nemotron Models for Agentic AI and Real-Time Voice

Together AITogether AI

· Updated

Together AI has made NVIDIA's Nemotron 3 Ultra and Nemotron 3.5 ASR models available on its AI Native Cloud. This integration provides developers with specialized capabilities for building high-throughput AI agents and low-latency multilingual voice systems. The move expands access to advanced models for autonomous workflows and real-time conversational AI.

Together AI now hosts NVIDIA Nemotron 3 Ultra and NVIDIA Nemotron 3.5 ASR on its AI Native Cloud. Nemotron 3 Ultra, a 550B-parameter model with 55B active parameters (Mixture of Experts), is designed for high-throughput agentic workloads. The Nemotron 3.5 ASR model, with 0.6B parameters, focuses on low-latency multilingual speech recognition (ASR) for real-time voice agents.
Nemotron 3 Ultra Parameters
550B (55B activated)
Nemotron 3.5 ASR Parameters
0.6B
Nemotron 3.5 ASR Languages
40 language-locales
Nemotron 3.5 ASR Latency
Sub-100ms

This integration addresses demand for specialized AI in autonomous systems and real-time interactions. Nemotron 3 Ultra is designed for high-throughput agentic workloads, while the Nemotron 3.5 ASR model supports 40 language-locales with sub-100ms latency, crucial for responsive global voice applications, building on Together AI's prior focus on Unified Voice Agent Cloud.

The Nemotron 3 Ultra model is available for building coding agents, deep research agents, and complex enterprise workflows, leveraging Together AI's inference stack on NVIDIA Blackwell GPUs. For real-time voice systems, Nemotron 3.5 ASR offers streaming multilingual speech recognition, powered by TensorRT engines and event-driven streaming I/O. Developers can start building on Together AI, the AI Native Cloud, with these NVIDIA Nemotron 3 Ultra and Nemotron 3.5 ASR models.

Together AI
Together AI
@togethercompute
X

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual speech recognition. AI natives can now build coding agents, deep research agents, and real-time voice systems on the AI Native Cloud.

2retweets23likes
View on X

Still wondering? A few quick answers below.

NVIDIA Nemotron 3 Ultra is a large language model specifically built for high-throughput agentic workloads. This includes tasks like autonomous coding, deep research, and orchestrating complex, multi-step enterprise workflows that require advanced reasoning.

NVIDIA Nemotron 3.5 ASR is a smaller model focused on streaming multilingual speech recognition. It supports 40 language-locale combinations, offers sub-100ms latency, and uses cache-aware FastConformer technology to maintain context across audio chunks.

Together AI provides the inference stack for both models on its AI Native Cloud. This includes high-throughput serving on NVIDIA Blackwell GPUs for Nemotron 3 Ultra's agentic workloads and TensorRT engines with event-driven streaming I/O for Nemotron 3.5 ASR's low-latency speech recognition.

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update