Model Gallery

27 models from 1 repositories

Filter by type:

Filter by tags:

silero-vad-sherpa
Silero VAD served through the sherpa-onnx backend. Uses the same ONNX weights as the dedicated silero-vad backend, loaded through sherpa-onnx's C VAD API. Pairs with the sherpa-onnx ASR entries for round-trip audio pipelines.

Repository: localaiLicense: mit

voice-hi_IN-priyamvada-medium
A fast, local neural text to speech system that sounds great and is optimized for the Raspberry Pi 4. Piper is used in a variety of [projects](https://github.com/rhasspy/piper#people-using-piper).

Repository: localaiLicense: mit

silero-vad
Silero VAD - pre-trained enterprise-grade Voice Activity Detector.

Repository: localai

silero-vad-ggml
Silero VAD - pre-trained enterprise-grade Voice Activity Detector.

Repository: localai

parakeet-cpp-moondream-ultra-f16
Moondream Ultra, F16: Moondream's post-trained derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). Runs on CPU and GPU. Full-precision file, about 1.4 GB. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-ultra-q8_0
Moondream Ultra, Q8_0: the same model as the F16 file, quantized to 8 bits (about 0.9 GB). Runs on CPU and GPU. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-packed
Moondream Redux, packed ternary weights (about 213 MB). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). CPU only and offline only: the library refuses to load it on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-f16
Moondream Redux, dequantized to F16 (about 1.4 GB). Same model as the packed file, but it runs on any backend, including GPU, and can stream. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-q8_0
Moondream Redux, dequantized and quantized to Q8_0 (about 0.9 GB). Runs on any backend, including GPU. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-silero-vad-f16
Silero VAD v6.2.3 as a GGUF (F16, about 1.3 MB) for the parakeet-cpp backend. It only detects speech: use it for the VAD endpoint, or as the vad_model: option of a parakeet-cpp ASR model that has no VAD head. Converted from the official Silero VAD model, MIT licensed, copyright Silero Team: https://github.com/snakers4/silero-vad

Repository: localaiLicense: mit

parakeet-cpp-vad-moondream-redux-packed
Voice activity detection with the VAD head of Moondream Redux, packed ternary weights (about 213 MB). CPU only. The file is the same as the ASR entry parakeet-cpp-moondream-redux-packed, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-ultra-q8_0
Voice activity detection with the VAD head of Moondream Ultra, Q8_0 (about 0.9 GB). Runs on CPU and GPU. The file is the same as the ASR entry parakeet-cpp-moondream-ultra-q8_0, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-redux
Voice activity detection with the VAD head of Moondream Redux cut out of the full model, as a small file (about 10 MB). It detects speech only and cannot transcribe: a transcription request fails with an error. The head is not retrained, so the output is the same as the full model's head, with a much smaller download and a faster load. Needs a parakeet.cpp build that can load VAD-only GGUF files. Older backend builds fail to load the file. For the full model, install parakeet-cpp-vad-moondream-redux-packed. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-ultra
Voice activity detection with the VAD head of Moondream Ultra, Q8_0 cut out of the full model, as a small file (about 6 MB). It detects speech only and cannot transcribe: a transcription request fails with an error. The head is not retrained, so the output is the same as the full model's head, with a much smaller download and a faster load. Needs a parakeet.cpp build that can load VAD-only GGUF files. Older backend builds fail to load the file. For the full model, install parakeet-cpp-vad-moondream-ultra-q8_0. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad
Voice activity detection served by the parakeet-cpp backend, with Silero VAD (about 1.3 MB). To use the VAD head of Moondream Redux or Ultra instead, install parakeet-cpp-vad-moondream-redux-packed or parakeet-cpp-vad-moondream-ultra-q8_0. The detectors are different, not builds of the same weights, so this entry does not pick between them.

Repository: localai

parakeet-cpp-tdt-0.6b-v3-silero-vad
TDT FastConformer, 0.6B (v3, multilingual) with Silero VAD. The vad_model option makes the backend cut long audio at pauses found by Silero before it transcribes, so recordings of any length work. The model has no VAD head of its own, so Silero does the cutting. Both files are GGUF for the parakeet-cpp backend. The Silero model is MIT licensed, copyright Silero Team.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-realtime-scene-speakers
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0, WeSpeaker ResNet34 CC-BY-4.0. Also loads WeSpeaker ResNet34 through the speaker_model option, so live speaker segments carry the name of a voice registered with /v1/voice/register (voice-detect-wespeaker-resnet34 model) once the speaker is identified.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-realtime-scene
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-realtime-scene-tdt
Parakeet TDT 0.6B v3 (multilingual, 25 European languages) paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options: one parakeet-cpp backend transcribes, labels speakers and tags sound events. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). TDT is not a streaming model, so in a realtime pipeline use it with server_vad: set it as both transcription and sound_detection and turn on pipeline.diarization, and each committed turn gets speaker segments (conversation.item.input_audio_transcription.segment, with text) and sound tags (conversation.item.sound_detection). Also labels speakers on /v1/audio/transcriptions. Speaker labels are per turn. License per model: transcription model CC-BY-4.0, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-realtime-scene-base
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Base (86M, the largest CED; more confident sound tags than CED-Tiny at a small extra cost: on CPU the live diarization + sound stream runs at 0.125 of real time against 0.103 with CED-Tiny) through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Base Apache-2.0.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-realtime-scene-tdt-base
Parakeet TDT 0.6B v3 (multilingual, 25 European languages) paired with Nemotron-3-Diarization and CED-Base (86M, the largest CED; about 0.03 s of CPU per second of audio per committed turn, against 0.005 for CED-Tiny) through the diarization_model and sound_model options: one parakeet-cpp backend transcribes, labels speakers and tags sound events. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). TDT is not a streaming model, so in a realtime pipeline use it with server_vad: set it as both transcription and sound_detection and turn on pipeline.diarization, and each committed turn gets speaker segments (conversation.item.input_audio_transcription.segment, with text) and sound tags (conversation.item.sound_detection). Also labels speakers on /v1/audio/transcriptions. Speaker labels are per turn. License per model: transcription model CC-BY-4.0, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: cc-by-4.0

Page 1