Model Gallery

17 models from 1 repositories

Filter by type:

Filter by tags:

dream-org_dream-v0-instruct-7b
This is the instruct model of Dream 7B, which is an open diffusion large language model with top-tier performance.

Repository: localaiLicense: apache-2.0

moondream2-20250414
Moondream is a small vision language model designed to run efficiently everywhere.

Repository: localaiLicense: apache-2.0

mn-12b-mag-mell-r1-iq-arm-imatrix
This is a merge of pre-trained language models created using mergekit. Mag Mell is a multi-stage merge, Inspired by hyper-merges like Tiefighter and Umbral Mind. Intended to be a general purpose "Best of Nemo" model for any fictional, creative use case. 6 models were chosen based on 3 categories; they were then paired up and merged via layer-weighted SLERP to create intermediate "specialists" which are then evaluated in their domain. The specialists were then merged into the base via DARE-TIES, with hyperparameters chosen to reduce interference caused by the overlap of the three domains. The idea with this approach is to extract the best qualities of each component part, and produce models whose task vectors represent more than the sum of their parts. The three specialists are as follows: Hero (RP, kink/trope coverage): Chronos Gold, Sunrose. Monk (Intelligence, groundedness): Bophades, Wissenschaft. Deity (Prose, flair): Gutenberg v4, Magnum 2.5 KTO. I've been dreaming about this merge since Nemo tunes started coming out in earnest. From our testing, Mag Mell demonstrates worldbuilding capabilities unlike any model in its class, comparable to old adventuring models like Tiefighter, and prose that exhibits minimal "slop" (not bad for no finetuning,) frequently devising electrifying metaphors that left us consistently astonished. I don't want to toot my own bugle though; I'm really proud of how this came out, but please leave your feedback, good or bad.Special thanks as usual to Toaster for his feedback and Fizz for helping fund compute, as well as the KoboldAI Discord for their resources. The following models were included in the merge: IntervitensInc/Mistral-Nemo-Base-2407-chatml nbeerbower/mistral-nemo-bophades-12B nbeerbower/mistral-nemo-wissenschaft-12B elinas/Chronos-Gold-12B-1.0 Fizzarolli/MN-12b-Sunrose nbeerbower/mistral-nemo-gutenberg-12B-v4 anthracite-org/magnum-12b-v2.5-kto

Repository: localaiLicense: unlicense

dreamgen_lucid-v1-nemo
Focused on role-play & story-writing. Suitable for all kinds of writers and role-play enjoyers: For world-builders who want to specify every detail in advance: plot, setting, writing style, characters, locations, items, lore, etc. For intuitive writers who start with a loose prompt and shape the narrative through instructions (OCC) as the story / role-play unfolds. Support for multi-character role-plays: Model can automatically pick between characters. Support for inline writing instructions (OOC): Controlling plot development (say what should happen, what the characters should do, etc.) Controlling pacing. etc. Support for inline writing assistance: Planning the next scene / the next chapter / story. Suggesting new characters. etc. Support for reasoning (opt-in).

Repository: localaiLicense: apache-2.0

moondream2
a tiny vision language model that kicks ass and runs anywhere

Repository: localaiLicense: apache-2.0

dreamshaper
A text-to-image model that uses Stable Diffusion 1.5 to generate images from text prompts. This model is DreamShaper model by Lykon.

Repository: localaiLicense: other

parakeet-cpp-moondream-ultra-f16
Moondream Ultra, F16: Moondream's post-trained derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). Runs on CPU and GPU. Full-precision file, about 1.4 GB. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-ultra-q8_0
Moondream Ultra, Q8_0: the same model as the F16 file, quantized to 8 bits (about 0.9 GB). Runs on CPU and GPU. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-packed
Moondream Redux, packed ternary weights (about 213 MB). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). CPU only and offline only: the library refuses to load it on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-f16
Moondream Redux, dequantized to F16 (about 1.4 GB). Same model as the packed file, but it runs on any backend, including GPU, and can stream. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-moondream-redux-q8_0
Moondream Redux, dequantized and quantized to Q8_0 (about 0.9 GB). Runs on any backend, including GPU. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-redux-packed
Voice activity detection with the VAD head of Moondream Redux, packed ternary weights (about 213 MB). CPU only. The file is the same as the ASR entry parakeet-cpp-moondream-redux-packed, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-ultra-q8_0
Voice activity detection with the VAD head of Moondream Ultra, Q8_0 (about 0.9 GB). Runs on CPU and GPU. The file is the same as the ASR entry parakeet-cpp-moondream-ultra-q8_0, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-redux
Voice activity detection with the VAD head of Moondream Redux cut out of the full model, as a small file (about 10 MB). It detects speech only and cannot transcribe: a transcription request fails with an error. The head is not retrained, so the output is the same as the full model's head, with a much smaller download and a faster load. Needs a parakeet.cpp build that can load VAD-only GGUF files. Older backend builds fail to load the file. For the full model, install parakeet-cpp-vad-moondream-redux-packed. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-ultra
Voice activity detection with the VAD head of Moondream Ultra, Q8_0 cut out of the full model, as a small file (about 6 MB). It detects speech only and cannot transcribe: a transcription request fails with an error. The head is not retrained, so the output is the same as the full model's head, with a much smaller download and a faster load. Needs a parakeet.cpp build that can load VAD-only GGUF files. Older backend builds fail to load the file. For the full model, install parakeet-cpp-vad-moondream-ultra-q8_0. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad
Voice activity detection served by the parakeet-cpp backend, with Silero VAD (about 1.3 MB). To use the VAD head of Moondream Redux or Ultra instead, install parakeet-cpp-vad-moondream-redux-packed or parakeet-cpp-vad-moondream-ultra-q8_0. The detectors are different, not builds of the same weights, so this entry does not pick between them.

Repository: localai

parakeet-cpp-bundle-moondream-redux
Moondream Redux with Silero VAD in one bundle GGUF file (about 215 MB) for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B), kept as published with the packed ternary weights. The bundle also holds Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions) and voice activity detection (/v1/vad). The vad:true option cuts long audio at pauses before it transcribes, with Silero as the VAD source. CPU only and offline only: the library refuses the packed Redux model on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet Redux by Moondream, derived from Parakeet TDT 0.6B v3 by NVIDIA, is CC-BY-4.0 (credit Moondream and NVIDIA); Silero VAD by the Silero Team is MIT. The weights were converted to GGUF; nothing was retrained.

Repository: localaiLicense: other