Back to models
PK
Speech to Text

NVIDIA NeMo / parakeet-tdt-0.6b-v3

Parakeet TDT 0.6B v3

Ultra-fast batch transcription

Languages

25 European

Parameters

600M

License

CC BY 4.0

OpenVox fit

Fast batch STT

Model overview

Built into OpenVox for private local speech transcription and subtitles.

Use in OpenVox

Parakeet TDT 0.6B v3 is NVIDIA’s high-performance multilingual ASR model built on FastConformer and Token-and-Duration Transducer (TDT) decoding for ultra-fast transcription and precise subtitle timestamps.

Achieves an impressive 6.34% mean WER on the Open ASR Leaderboard, outperforming larger models on English meeting and clean speech benchmarks.

Uses Token-and-Duration Transducer (TDT) decoding to predict both tokens and durations simultaneously, achieving major inference speedups.

Officially supports 25 European languages with native word-level and segment-level timestamp extraction for pinpoint subtitle alignment.

In OpenVox, Parakeet TDT is the recommended engine for high-volume batch audio processing and rapid subtitle creation.

All speech transcription runs 100% locally on your computer with offline model weights. No microphone audio, meeting recordings, or transcripts are ever sent to external cloud servers.

Open-source model: NVIDIA NeMo / parakeet-tdt-0.6b-v3

View on Hugging Face