Open speech models for real-time voice agents. @NVIDIAAI Nemotron 3.5 ASR is now available on fal. An open streaming speech recognition model supporting 40 language-locale combinations with ultra-low latency, native punctuation, and capitalization. Built for voice agents, call centers, meeting transcription, live captions, and in-car assistants.
fal Makes NVIDIA Nemotron 3.5 ASR Available for Real-Time Multilingual Voice Agents
fal· Updated
fal has integrated NVIDIA's Nemotron 3.5 ASR, an open streaming speech recognition model, onto its platform. This provides ultra-low latency, multilingual transcription across 40 language-locale combinations, enabling high-performance real-time voice agents and applications like call centers.
- Model
- NVIDIA Nemotron 3.5 ASR (open streaming)
- Languages
- 40 language-locale combinations
- Features
- Ultra-low latency, native punctuation & capitalization
- Use Cases
- Voice agents, call centers, meeting transcription, live captions, in-car assistants
- Availability
- fal API and Hugging Face
This high-performance, multilingual ASR model with ultra-low latency is designed for real-time AI agents that interact via voice. It enables developers to build systems requiring rapid and accurate speech understanding. This integration extends fal's platform, which also offers an MCP server for connecting agents to generative models.
Developers can now access NVIDIA Nemotron 3.5 ASR on fal for applications like voice agents, call centers, meeting transcription, live captions, and in-car assistants. The model is also available on Hugging Face, offering a production-ready API for integrating advanced speech recognition into real-time voice products. This complements fal's genmedia CLI for empowering agents with media generation.
Still wondering? A few quick answers below.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →




