Back to models
NM
Streaming STT

NVIDIA / nemotron-3.5-asr-streaming-0.6b

Nemotron 3.5 ASR

Real-time streaming speech

Locales

40 locales

Parameters

600M

Latency

80ms – 1120ms

OpenVox fit

Live streaming STT

Model overview

Built into OpenVox for private local speech transcription and subtitles.

Use in OpenVox

Nemotron 3.5 ASR Streaming 0.6B is NVIDIA’s next-generation cache-aware streaming speech recognition foundation model engineered for low-latency live audio capture and voice agent workflows.

Cache-aware streaming reuses encoder contextual states across consecutive audio chunks, preventing latency spikes and context fragmentation.

Supports adjustable chunk sizes ranging from 80ms to 1120ms to balance near-zero-latency live dictation with maximum batch accuracy.

Features locale and language prompting for accurate cross-lingual transcription across 40 supported regional language variants.

Powers real-time microphone dictation, voice agent listeners, and live transcription inside OpenVox without external servers.

All speech transcription runs 100% locally on your computer with offline model weights. No microphone audio, meeting recordings, or transcripts are ever sent to external cloud servers.

Open-source model: NVIDIA / nemotron-3.5-asr-streaming-0.6b

View on Hugging Face