Model Gallery

191 models from 1 repositories

Filter by type:

Filter by tags:

ornith-1.0-9b-q4
Ornith-1.0-9B is an MIT-licensed Qwen3.5 model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and F16 vision projector. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ornith-1.0-9b-q8
Ornith-1.0-9B in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-9b-q4
Ornith-1.5-9B is an MIT-licensed Qwen3.5 model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ornith-1.5-9b-q8
Ornith-1.5-9B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

qwen3.8-27b-q4
Qwen3.8-27B is Qwen's dense 27B vision-language model for reasoning, coding, tool use, and long-running agent tasks. It accepts text, images, and video, and it supports a native context window of 262K tokens. This default entry uses the official Q4_K_M GGUF and Q8_0 vision projector. The linked variants add MTP speculative decoding or use the higher-quality Q8_0 model.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-q4-mtp
Qwen3.8-27B with the official Q4_K_M model and Q4_0 MTP draft model. MTP speculative decoding can increase generation speed by proposing multiple tokens for the target model to verify.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-q8
Qwen3.8-27B in the official Q8_0 GGUF format. This variant provides higher model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-ridge
Qwen3.8-27B Ridge is a 3.69-bit mixed quantization that keeps the Gated-DeltaNet state path at Q8_0 and preserves the embedded MTP head. It reduces the model weights to 12.59 GB while retaining multimodal, reasoning, coding, tool-use, and long-context capabilities.

Repository: localaiLicense: apache-2.0

muse-glimmer-30b
Muse Glimmer is Meta Superintelligence Labs' Apache-2.0 dense 30B model for autonomous agentic work, coding, tool use, long-horizon reasoning, and multimodal understanding. It supports more than 100 languages, interleaved text and image input through its 1.8B-parameter perception encoder, and a 131K-token context window. This entry uses the publisher's higher-quality dynamic K-quant GGUF and official quantized vision projector. Automatic variant selection can use the smaller 17 GB quantization or a DFlash-accelerated build when it fits.

Repository: localaiLicense: apache-2.0

muse-glimmer-30b-dflash
Muse Glimmer's higher-quality dynamic K-quant GGUF with the official quantized perception encoder and DFlash drafter. DFlash proposes blocks of up to 16 tokens for the target to verify in parallel, accelerating output without changing model quality. Flash attention is enabled for this path.

Repository: localaiLicense: apache-2.0

muse-glimmer-30b-17gb
Muse Glimmer's smaller 17 GB K-quant GGUF with the official quantized perception encoder. It preserves the model's agentic, coding, tool-use, multilingual, and image-understanding capabilities for hosts with less memory than the dynamic quantization requires.

Repository: localaiLicense: apache-2.0

muse-glimmer-30b-17gb-dflash
Muse Glimmer's smaller 17 GB K-quant GGUF with the official quantized perception encoder and DFlash drafter. This is the lowest-memory published build that retains image understanding and block-speculative decoding. Flash attention is enabled for the DFlash path.

Repository: localaiLicense: apache-2.0

qwen3.5-9b-defiant-fable-mtp
Qwen3.5 9B Defiant Fable is an Apache-2.0 multimodal fine-tune for reasoning, coding, creative writing, and roleplay. It retains the 256K context window and vision support of Qwen3.5 while reducing refusals. This default entry uses the NEO-imatrix Q4_K_M build with multi-token prediction enabled for faster generation.

Repository: localaiLicense: apache-2.0

qwen3.5-9b-defiant-fable
Qwen3.5 9B Defiant Fable in the plain NEO-imatrix Q4_K_M GGUF format. This fallback offers the same multimodal reasoning, coding, and creative capabilities without enabling multi-token prediction.

Repository: localaiLicense: apache-2.0

grug-27b
Grug 27B is a multimodal Qwen3.5-derived model for chat, reasoning, vision, and tool use. This entry uses the QAT Q4_K_M GGUF build.

Repository: localaiLicense: apache-2.0

grug-27b-q8
Grug 27B Q8 is the higher-precision Q8_0 GGUF build for multimodal chat, reasoning, vision, and tool use.

Repository: localaiLicense: apache-2.0

grug-27b-mtp
Grug 27B MTP is the Q4_K_M GGUF build with multi-token prediction enabled for speculative decoding, plus the shared vision projector.

Repository: localaiLicense: apache-2.0

kimi-k3
📰  Tech Blog |     📄  Full Report ## 1. Model Introduction Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning. ...

Repository: localaiLicense: other

qwen3.6-35b-a3b-uncensored-genesis-hermes-v6
Qwen3.6-35B-A3B Uncensored Genesis Hermes V6 is LuffyTheFox's multimodal, agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It combines Genesis tensor calibration with Hermes function-calling data while retaining the 35B mixture-of-experts architecture, roughly 3B active parameters per token, and the native 262K-token context window. This entry installs the Q8_0 GGUF together with its F16 multimodal projector for llama.cpp. The model card recommends Jinja chat templates and at least a 128K context for its thinking behavior. License: Apache-2.0.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7
Qwen3.6-35B-A3B Genesis Hermes V7 is LuffyTheFox's Apache-2.0 multimodal, agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It combines Genesis tensor calibration with Hermes function-calling data while retaining the 35B mixture-of-experts architecture, roughly 3B active parameters per token, and the native 262K-token context window. This entry's own payload uses the model card's recommended APEX GGUF and the shared F16 multimodal projector. Automatic variant selection may instead choose Compact APEX, an MTP-enabled APEX build, or Q8_K_P based on serving features and available memory. The model card recommends Jinja chat templates and at least a 128K context for its thinking behavior.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7-apex-compact
Qwen3.6-35B-A3B Genesis Hermes V7 in the smaller APEX Compact GGUF format, with the shared F16 multimodal projector. This build preserves the model's multimodal, reasoning, coding, and agentic capabilities for hosts with less memory than the recommended full APEX build.

Repository: localaiLicense: apache-2.0

Page 1