Model Gallery

9 models from 1 repositories

Filter by type:

Filter by tags:

parakeet-cpp-realtime-scene-speakers
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0, WeSpeaker ResNet34 CC-BY-4.0. Also loads WeSpeaker ResNet34 through the speaker_model option, so live speaker segments carry the name of a voice registered with /v1/voice/register (voice-detect-wespeaker-resnet34 model) once the speaker is identified.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-ced-tiny
CED-Tiny sound event tagger, Q8_0 GGUF for the parakeet-cpp backend (C++/ggml, loaded through third_party/ced.cpp). Served through /v1/audio/classification: 10 s windows are scored and averaged over the clip, then sorted by score with threshold and top_k applied. Smallest and fastest of the CED sizes; use ced-base for higher accuracy.

Repository: localaiLicense: apache-2.0

parakeet-cpp-ced-base
CED-Base sound event tagger, Q8_0 GGUF for the parakeet-cpp backend (C++/ggml, loaded through third_party/ced.cpp). Served through /v1/audio/classification: 10 s windows are scored and averaged over the clip, then sorted by score with threshold and top_k applied. Larger and more accurate than ced-tiny, still CPU-friendly.

Repository: localaiLicense: apache-2.0

parakeet-cpp-realtime-scene
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-realtime-scene-tdt
Parakeet TDT 0.6B v3 (multilingual, 25 European languages) paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options: one parakeet-cpp backend transcribes, labels speakers and tags sound events. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). TDT is not a streaming model, so in a realtime pipeline use it with server_vad: set it as both transcription and sound_detection and turn on pipeline.diarization, and each committed turn gets speaker segments (conversation.item.input_audio_transcription.segment, with text) and sound tags (conversation.item.sound_detection). Also labels speakers on /v1/audio/transcriptions. Speaker labels are per turn. License per model: transcription model CC-BY-4.0, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-realtime-scene-base
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Base (86M, the largest CED; more confident sound tags than CED-Tiny at a small extra cost: on CPU the live diarization + sound stream runs at 0.125 of real time against 0.103 with CED-Tiny) through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Base Apache-2.0.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-realtime-scene-tdt-base
Parakeet TDT 0.6B v3 (multilingual, 25 European languages) paired with Nemotron-3-Diarization and CED-Base (86M, the largest CED; about 0.03 s of CPU per second of audio per committed turn, against 0.005 for CED-Tiny) through the diarization_model and sound_model options: one parakeet-cpp backend transcribes, labels speakers and tags sound events. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). TDT is not a streaming model, so in a realtime pipeline use it with server_vad: set it as both transcription and sound_detection and turn on pipeline.diarization, and each committed turn gets speaker segments (conversation.item.input_audio_transcription.segment, with text) and sound tags (conversation.item.sound_detection). Also labels speakers on /v1/audio/transcriptions. Speaker labels are per turn. License per model: transcription model CC-BY-4.0, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-bundle-small
Parakeet TDT+CTC 110M with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 338 MB) that holds five models for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT+CTC 110M (English, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true option cuts long audio at pauses found by Silero before it transcribes. The bundle options (diar_component, sound_component, speaker_component) load each model from the same file; see the audio-to-text docs for the option list. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet TDT+CTC 110M by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0 (as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained), WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The weights were converted to GGUF and quantised where stated; nothing was retrained.

Repository: localaiLicense: other

parakeet-cpp-bundle-standard
Parakeet TDT 0.6B v3 (multilingual) with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 1101 MB) that holds five models for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT 0.6B v3 (multilingual, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true option cuts long audio at pauses found by Silero before it transcribes. The bundle options (diar_component, sound_component, speaker_component) load each model from the same file; see the audio-to-text docs for the option list. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet TDT 0.6B v3 by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0 (as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained), WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The weights were converted to GGUF and quantised where stated; nothing was retrained.

Repository: localaiLicense: other