parakeet-cpp-bundle-small
Parakeet TDT+CTC 110M with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 338 MB) that holds five models for the parakeet-cpp backend
(C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT+CTC 110M (English, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound
events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves
transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with
text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true
option cuts long audio at pauses found by Silero before it transcribes. The bundle options
(diar_component, sound_component, speaker_component) load each model from the same file; see the
audio-to-text docs for the option list.
A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file
next to the file. Parakeet TDT+CTC 110M by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0
(as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say
CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained),
WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The
weights were converted to GGUF and quantised where stated; nothing was retrained.