NVIDIA / nemotron-3.5-asr-streaming-0.6b
Nemotron 3.5 ASR
Real-time streaming speech
Locales
40 locales
Parameters
600M
Latency
80ms – 1120ms
OpenVox fit
Live streaming STT
Model overview
Built into OpenVox for private local speech transcription and subtitles.
Nemotron 3.5 ASR Streaming 0.6B is NVIDIA’s next-generation cache-aware streaming speech recognition foundation model engineered for low-latency live audio capture and voice agent workflows.
Cache-aware streaming reuses encoder contextual states across consecutive audio chunks, preventing latency spikes and context fragmentation.
Supports adjustable chunk sizes ranging from 80ms to 1120ms to balance near-zero-latency live dictation with maximum batch accuracy.
Features locale and language prompting for accurate cross-lingual transcription across 40 supported regional language variants.
Powers real-time microphone dictation, voice agent listeners, and live transcription inside OpenVox without external servers.
All speech transcription runs 100% locally on your computer with offline model weights. No microphone audio, meeting recordings, or transcripts are ever sent to external cloud servers.
Open-source model: NVIDIA / nemotron-3.5-asr-streaming-0.6b
View on Hugging Face