Model Gallery

92 models from 1 repositories

Filter by type:

Filter by tags:

hy-mt2-1.8b-q8
Hy-MT2-1.8B in the higher-quality 1.9 GB Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

carbon-3b-q8
Carbon-3B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.

Repository: localaiLicense: apache-2.0

carbon-8b-q8
Carbon-8B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.

Repository: localaiLicense: apache-2.0

ornith-1.0-9b-q8
Ornith-1.0-9B in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-9b-q8
Ornith-1.5-9B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

qwen3.8-27b-q8
Qwen3.8-27B in the official Q8_0 GGUF format. This variant provides higher model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-9b-q8
Qwen3.8-9B in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-4b-q8
Qwen3.8-4B in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

qwen3.8-2b-q8
Qwen3.8-2B in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity while remaining suitable for compact hosts.

Repository: localaiLicense: apache-2.0

twil-lm3-q8
TwIL-LM3 in the publisher's near-lossless Q8_0 GGUF format for quality-sensitive use on hosts with enough memory.

Repository: localaiLicense: webai-non-commercial-license-ver.-1.0

nemotron-3.5-lightning-30b-a3b-q8
NVIDIA Nemotron 3.5 Lightning 30B-A3B in the official high-quality Q8_0 GGUF format for hosts with enough memory.

Repository: localaiLicense: openmdw-1.1

grug-27b-q8
Grug 27B Q8 is the higher-precision Q8_0 GGUF build for multimodal chat, reasoning, vision, and tool use.

Repository: localaiLicense: apache-2.0

instella-moe-16b-a3b-think-q8
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the near-lossless Q8_0 GGUF quantization.

Repository: localaiLicense: other

north-mini-code-1.0-q8
North Mini Code 1.0 is Cohere Labs' Apache-2.0 sparse mixture-of-experts coding model with 30B total parameters and 3B active parameters. It targets code generation, agentic software engineering, terminal tasks, tool use, and interleaved reasoning with a 256K-token context window. This entry uses the Q8_0 GGUF quantization.

Repository: localaiLicense: apache-2.0

mellum2-12b-a2.5b-instruct-q8
Mellum2-12B-A2.5B-Instruct is an Apache-2.0 mixture-of-experts model from JetBrains with 12 billion total parameters, 2.5 billion activated per token, and a 131,072-token context window. This entry uses the higher-quality Q8_0 GGUF quantization.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7-q8-k-p
Qwen3.6-35B-A3B Genesis Hermes V7 in the high-quality Q8_K_P GGUF format, with the shared F16 multimodal projector. This is the largest non-MTP build in the published V7 set.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev-q8
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry uses the higher-quality Q8_0 GGUF quantization.

Repository: localaiLicense: apache-2.0

qwopus3.6-27b-fusion-q8
Qwopus3.6-27B Fusion is an experimental Qwen3.6-27B reasoning and coding merge. This entry uses the near-lossless Q8_0 GGUF quantization and the shared Q8_0 vision projector.

Repository: localaiLicense: qwen

tess-4-27b-q8
Tess-4-27B is an Apache-2.0 agentic and reasoning model built on Qwen3.6-27B. This entry uses the near-lossless Q8_0 GGUF quantization and the shared F16 vision projector.

Repository: localaiLicense: apache-2.0

qwen3.6-14b-a3b-fablevibes-q8
Qwen3.6-14B-A3B-FableVibes is an Apache-2.0 mixture-of-experts reasoning model distilled from Fable 5 and Claude Opus traces, with additional tool calling and coding data. This entry uses the near-lossless Q8_0 GGUF quantization and its matching Q8_0 multimodal projector.

Repository: localaiLicense: apache-2.0

agents-a1-4b-q8
Agents-A1-4B is InternScience's Apache-2.0 dense 4B agentic model, based on Qwen3.5. It is trained for long-horizon search, engineering and scientific research, instruction following, tool use, and multimodal tasks. This entry uses the official Q8_0 GGUF quantization and vision projector.

Repository: localaiLicense: apache-2.0

Page 1