Model Gallery

50 models from 1 repositories

Filter by type:

Filter by tags:

nemo-parakeet-tdt-0.6b
NVIDIA NeMo Parakeet TDT 0.6B v3 is an automatic speech recognition (ASR) model from NVIDIA's NeMo toolkit. Parakeet models are state-of-the-art ASR models trained on large-scale English audio data.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt_ctc-110m
Hybrid TDT+CTC FastConformer, 110M. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-realtime_eou_120m-v1
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M. Use with streaming transcription. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-ctc-0.6b
CTC FastConformer, 0.6B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-rnnt-0.6b
RNNT FastConformer, 0.6B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt-0.6b-v2
TDT FastConformer, 0.6B (v2). F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt-0.6b-v3
TDT FastConformer, 0.6B (v3, multilingual). F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

orukeet
Orukeet is a 0.6B, 25-language fine-tune of Parakeet TDT v3 by Oruk. Q8 GGUF for the nemo-speech-cpp backend. Runs locally on CPU, with optional GPU acceleration through the backend's gpu option.

Repository: localaiLicense: cc-by-sa-4.0

parakeet-cpp-ctc-1.1b
CTC FastConformer, 1.1B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-rnnt-1.1b
RNNT FastConformer, 1.1B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt-1.1b
TDT FastConformer, 1.1B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt_ctc-1.1b
Hybrid TDT+CTC FastConformer, 1.1B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-nemotron-3.5-asr-streaming-0.6b
Multilingual (40+ locales), prompt-conditioned, cache-aware streaming FastConformer RNN-T, 0.6B. Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Byte-identical to NeMo at WER 0 offline and streaming, about 2.5x faster than NeMo on CPU with no GPU. Select a language with the request "language" field (for example en, de, es, ja-JP), or leave it empty for automatic detection. License OpenMDW-1.1.

Repository: localaiLicense: other

parakeet-cpp-moondream-ultra-f16
Moondream Ultra, F16: Moondream's post-trained derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). Runs on CPU and GPU. Full-precision file, about 1.4 GB. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-ultra-q8_0
Moondream Ultra, Q8_0: the same model as the F16 file, quantized to 8 bits (about 0.9 GB). Runs on CPU and GPU. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-packed
Moondream Redux, packed ternary weights (about 213 MB). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). CPU only and offline only: the library refuses to load it on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-f16
Moondream Redux, dequantized to F16 (about 1.4 GB). Same model as the packed file, but it runs on any backend, including GPU, and can stream. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-q8_0
Moondream Redux, dequantized and quantized to Q8_0 (about 0.9 GB). Runs on any backend, including GPU. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-silero-vad-f16
Silero VAD v6.2.3 as a GGUF (F16, about 1.3 MB) for the parakeet-cpp backend. It only detects speech: use it for the VAD endpoint, or as the vad_model: option of a parakeet-cpp ASR model that has no VAD head. Converted from the official Silero VAD model, MIT licensed, copyright Silero Team: https://github.com/snakers4/silero-vad

Repository: localaiLicense: mit

parakeet-cpp-vad-moondream-redux-packed
Voice activity detection with the VAD head of Moondream Redux, packed ternary weights (about 213 MB). CPU only. The file is the same as the ASR entry parakeet-cpp-moondream-redux-packed, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-ultra-q8_0
Voice activity detection with the VAD head of Moondream Ultra, Q8_0 (about 0.9 GB). Runs on CPU and GPU. The file is the same as the ASR entry parakeet-cpp-moondream-ultra-q8_0, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

Page 1