Model Gallery

13 models from 1 repositories

Filter by type:

Filter by tags:

moondream2-20250414
Moondream is a small vision language model designed to run efficiently everywhere.

Repository: localaiLicense: apache-2.0

moondream2
a tiny vision language model that kicks ass and runs anywhere

Repository: localaiLicense: apache-2.0

parakeet-cpp-moondream-ultra-f16
Moondream Ultra, F16: Moondream's post-trained derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). Runs on CPU and GPU. Full-precision file, about 1.4 GB. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-ultra-q8_0
Moondream Ultra, Q8_0: the same model as the F16 file, quantized to 8 bits (about 0.9 GB). Runs on CPU and GPU. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-packed
Moondream Redux, packed ternary weights (about 213 MB). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). CPU only and offline only: the library refuses to load it on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-f16
Moondream Redux, dequantized to F16 (about 1.4 GB). Same model as the packed file, but it runs on any backend, including GPU, and can stream. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-q8_0
Moondream Redux, dequantized and quantized to Q8_0 (about 0.9 GB). Runs on any backend, including GPU. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-redux-packed
Voice activity detection with the VAD head of Moondream Redux, packed ternary weights (about 213 MB). CPU only. The file is the same as the ASR entry parakeet-cpp-moondream-redux-packed, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-ultra-q8_0
Voice activity detection with the VAD head of Moondream Ultra, Q8_0 (about 0.9 GB). Runs on CPU and GPU. The file is the same as the ASR entry parakeet-cpp-moondream-ultra-q8_0, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-redux
Voice activity detection with the VAD head of Moondream Redux cut out of the full model, as a small file (about 10 MB). It detects speech only and cannot transcribe: a transcription request fails with an error. The head is not retrained, so the output is the same as the full model's head, with a much smaller download and a faster load. Needs a parakeet.cpp build that can load VAD-only GGUF files. Older backend builds fail to load the file. For the full model, install parakeet-cpp-vad-moondream-redux-packed. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-ultra
Voice activity detection with the VAD head of Moondream Ultra, Q8_0 cut out of the full model, as a small file (about 6 MB). It detects speech only and cannot transcribe: a transcription request fails with an error. The head is not retrained, so the output is the same as the full model's head, with a much smaller download and a faster load. Needs a parakeet.cpp build that can load VAD-only GGUF files. Older backend builds fail to load the file. For the full model, install parakeet-cpp-vad-moondream-ultra-q8_0. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad
Voice activity detection served by the parakeet-cpp backend, with Silero VAD (about 1.3 MB). To use the VAD head of Moondream Redux or Ultra instead, install parakeet-cpp-vad-moondream-redux-packed or parakeet-cpp-vad-moondream-ultra-q8_0. The detectors are different, not builds of the same weights, so this entry does not pick between them.

Repository: localai

parakeet-cpp-bundle-moondream-redux
Moondream Redux with Silero VAD in one bundle GGUF file (about 215 MB) for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B), kept as published with the packed ternary weights. The bundle also holds Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions) and voice activity detection (/v1/vad). The vad:true option cuts long audio at pauses before it transcribes, with Silero as the VAD source. CPU only and offline only: the library refuses the packed Redux model on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet Redux by Moondream, derived from Parakeet TDT 0.6B v3 by NVIDIA, is CC-BY-4.0 (credit Moondream and NVIDIA); Silero VAD by the Silero Team is MIT. The weights were converted to GGUF; nothing was retrained.

Repository: localaiLicense: other