fal Makes NVIDIA Nemotron 3.5 ASR Available for Real-Time Multilingual Voice Agents

falfal

· Updated

fal has integrated NVIDIA's Nemotron 3.5 ASR, an open streaming speech recognition model, onto its platform. This provides ultra-low latency, multilingual transcription across 40 language-locale combinations, enabling high-performance real-time voice agents and applications like call centers.

fal has integrated NVIDIA Nemotron 3.5 ASR, an open streaming speech recognition (ASR) model, onto its platform. This model is engineered for high-quality, multilingual transcription, supporting 40 language-locale combinations with ultra-low latency, native punctuation, and capitalization.
Model
NVIDIA Nemotron 3.5 ASR (open streaming)
Languages
40 language-locale combinations
Features
Ultra-low latency, native punctuation & capitalization
Use Cases
Voice agents, call centers, meeting transcription, live captions, in-car assistants
Availability
fal API and Hugging Face

This high-performance, multilingual ASR model with ultra-low latency is designed for real-time AI agents that interact via voice. It enables developers to build systems requiring rapid and accurate speech understanding. This integration extends fal's platform, which also offers an MCP server for connecting agents to generative models.

Developers can now access NVIDIA Nemotron 3.5 ASR on fal for applications like voice agents, call centers, meeting transcription, live captions, and in-car assistants. The model is also available on Hugging Face, offering a production-ready API for integrating advanced speech recognition into real-time voice products. This complements fal's genmedia CLI for empowering agents with media generation.

fal
fal
@fal
X

Open speech models for real-time voice agents. @NVIDIAAI Nemotron 3.5 ASR is now available on fal. An open streaming speech recognition model supporting 40 language-locale combinations with ultra-low latency, native punctuation, and capitalization. Built for voice agents, call centers, meeting transcription, live captions, and in-car assistants.

2retweets48likes
View on X

Still wondering? A few quick answers below.

It's an open streaming Automatic Speech Recognition (ASR) model from NVIDIA, now available on fal. It's designed for high-quality, multilingual transcription in real-time applications.

The model supports 40 language-locale combinations, offers ultra-low latency, and includes native punctuation and capitalization for accurate transcription.

It's built for real-time voice agents, call centers, meeting transcription, live captions, and in-car assistants, enabling rapid and accurate speech understanding.

You can access NVIDIA Nemotron 3.5 ASR on the fal platform via its API, and it is also available on Hugging Face for broader use.

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update