Model Gallery

4 models from 1 repositories

Filter by type:

Filter by tags:

qwen3-omni-30b-a3b-instruct
Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation model. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech. This GGUF build runs on llama.cpp with the bundled mmproj for multimodal inputs.

Repository: localaiLicense: apache-2.0

qwen3-omni-30b-a3b-thinking
Qwen3-Omni-30B-A3B-Thinking is the reasoning-enhanced variant of Qwen3-Omni, a natively end-to-end multilingual omni-modal foundation model. It processes text, images, and audio and produces chain-of-thought reasoning before the final answer. This GGUF build runs on llama.cpp with the bundled mmproj.

Repository: localaiLicense: apache-2.0

parakeet-cpp-bundle-small
Parakeet TDT+CTC 110M with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 338 MB) that holds five models for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT+CTC 110M (English, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true option cuts long audio at pauses found by Silero before it transcribes. The bundle options (diar_component, sound_component, speaker_component) load each model from the same file; see the audio-to-text docs for the option list. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet TDT+CTC 110M by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0 (as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained), WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The weights were converted to GGUF and quantised where stated; nothing was retrained.

Repository: localaiLicense: other

parakeet-cpp-bundle-standard
Parakeet TDT 0.6B v3 (multilingual) with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 1101 MB) that holds five models for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT 0.6B v3 (multilingual, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true option cuts long audio at pauses found by Silero before it transcribes. The bundle options (diar_component, sound_component, speaker_component) load each model from the same file; see the audio-to-text docs for the option list. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet TDT 0.6B v3 by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0 (as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained), WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The weights were converted to GGUF and quantised where stated; nothing was retrained.

Repository: localaiLicense: other