
Silero VAD served through the sherpa-onnx backend. Uses the same ONNX weights as the dedicated silero-vad backend, loaded through sherpa-onnx's C VAD API. Pairs with the sherpa-onnx ASR entries for round-trip audio pipelines.
Links
Tags

Silero VAD - pre-trained enterprise-grade Voice Activity Detector.
Links
Tags

Silero VAD - pre-trained enterprise-grade Voice Activity Detector.
Links
Tags
Repository: localaiLicense: mit

Silero VAD v6.2.3 as a GGUF (F16, about 1.3 MB) for the parakeet-cpp backend. It only detects speech: use it for the VAD endpoint, or as the vad_model: option of a parakeet-cpp ASR model that has no VAD head. Converted from the official Silero VAD model, MIT licensed, copyright Silero Team: https://github.com/snakers4/silero-vad
Links
Tags
Repository: localaiLicense: cc-by-4.0
TDT FastConformer, 0.6B (v3, multilingual) with Silero VAD. The vad_model option makes the backend cut long audio at pauses found by Silero before it transcribes, so recordings of any length work. The model has no VAD head of its own, so Silero does the cutting. Both files are GGUF for the parakeet-cpp backend. The Silero model is MIT licensed, copyright Silero Team.
Links
Tags
Repository: localaiLicense: other
Parakeet TDT+CTC 110M with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 338 MB) that holds five models for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT+CTC 110M (English, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true option cuts long audio at pauses found by Silero before it transcribes. The bundle options (diar_component, sound_component, speaker_component) load each model from the same file; see the audio-to-text docs for the option list. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet TDT+CTC 110M by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0 (as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained), WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The weights were converted to GGUF and quantised where stated; nothing was retrained.
Links
Tags
Repository: localaiLicense: other
Parakeet TDT 0.6B v3 (multilingual) with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 1101 MB) that holds five models for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT 0.6B v3 (multilingual, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true option cuts long audio at pauses found by Silero before it transcribes. The bundle options (diar_component, sound_component, speaker_component) load each model from the same file; see the audio-to-text docs for the option list. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet TDT 0.6B v3 by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0 (as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained), WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The weights were converted to GGUF and quantised where stated; nothing was retrained.
Links
Tags
Repository: localaiLicense: other
Moondream Redux with Silero VAD in one bundle GGUF file (about 215 MB) for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B), kept as published with the packed ternary weights. The bundle also holds Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions) and voice activity detection (/v1/vad). The vad:true option cuts long audio at pauses before it transcribes, with Silero as the VAD source. CPU only and offline only: the library refuses the packed Redux model on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet Redux by Moondream, derived from Parakeet TDT 0.6B v3 by NVIDIA, is CC-BY-4.0 (credit Moondream and NVIDIA); Silero VAD by the Silero Team is MIT. The weights were converted to GGUF; nothing was retrained.
Links
Tags
Silero VAD (audio.cpp) - voice activity detection over /v1/vad. The model ships inside the audio-cpp backend package, so installing this entry downloads nothing.
Links
Tags