Model Gallery

14 models from 1 repositories

Filter by type:

Filter by tags:

qwen3.8-27b-ridge
Qwen3.8-27B Ridge is a 3.69-bit mixed quantization that keeps the Gated-DeltaNet state path at Q8_0 and preserves the embedded MTP head. It reduces the model weights to 12.59 GB while retaining multimodal, reasoning, coding, tool-use, and long-context capabilities.

Repository: localaiLicense: apache-2.0

nemotron-3.5-lightning-30b-a3b-nvfp4
NVIDIA Nemotron 3.5 Lightning 30B-A3B in the official NVFP4 GGUF format. This is the smallest linked build and retains the model's reasoning, coding, tool-use, multilingual, and long-context capabilities.

Repository: localaiLicense: openmdw-1.1

muse-glimmer-30b-17gb
Muse Glimmer's smaller 17 GB K-quant GGUF with the official quantized perception encoder. It preserves the model's agentic, coding, tool-use, multilingual, and image-understanding capabilities for hosts with less memory than the dynamic quantization requires.

Repository: localaiLicense: apache-2.0

btl-4-compact
BTL-4 Compact is Bad Theory Labs' text-only 35B mixture-of-experts model compressed into a single 9.96 GB IQ2_XXS GGUF. Around 2.1B parameters are active per token, and the model is tuned for agentic work, tool use, coding, and reasoning. The compact build omits the vision tower and disables the source model's MTP layer for compatibility with stock llama.cpp.

Repository: localaiLicense: apache-2.0

instella-moe-16b-a3b-think
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the Q4_K_M GGUF quantization.

Repository: localaiLicense: other

instella-moe-16b-a3b-think-q8
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the near-lossless Q8_0 GGUF quantization.

Repository: localaiLicense: other

parable-granite-4.1-3b-claude-fable-5
# Parable-Granite-4.1-3B-Claude-Fable-5 Granite 4.1 3B fine-tuned on genuine Claude Fable 5 and GPT-5.5 agent traces (planning, tool use, reasoning from real agent sessions). Agent-flavored small model: terminal workflows, idiomatic code fixes, explanations. v2 recipe: completion-masked SFT, replay mix, seed-averaged weights. Published corpus and eval harness.

Repository: localaiLicense: apache-2.0

parable-granite-4.1-8b-claude-fable-5
# Parable-Granite-4.1-8B-Claude-Fable-5 Granite 4.1 8B fine-tuned on genuine Claude Fable 5 and GPT-5.5 agent traces. Strongest Parable model: multi-step scripts, configs, terminal workflows, reasoning.

Repository: localaiLicense: apache-2.0

inkling
# Inkling BF16 | NVFP4 | Playground | Tinker Cookbook | Acceptable Use ## 1. General Information Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers. **Languages:** English, with general multilingual capabilities across other languages. ## 2. Getting Started Try Inkling on the Tinker Playground or access via API using the Tinker Cookbook. Inkling supports local deployment using the following open-source libraries: * SGLang (recipe, PR) * vLLM (recipe, PR) * TokenSpeed (recipe, PR) * Unsloth (recipe, PR) * Huggingface (recipe, PR) ...

Repository: localaiLicense: apache-2.0

arex-turbo
AREX-Turbo is BAAI's compact 4B deep-research agent, fine-tuned from Qwen3.5-4B for long-horizon search, evidence aggregation, constraint verification, and tool-assisted reasoning. It supports text and image input with a 262K-token context window. This entry uses the recommended Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

arex-turbo-q8
AREX-Turbo is BAAI's compact 4B deep-research agent, fine-tuned from Qwen3.5-4B for long-horizon search, evidence aggregation, constraint verification, and tool-assisted reasoning. It supports text and image input with a 262K-token context window. This entry uses the higher-quality Q8_0 GGUF quantization.

Repository: localaiLicense: apache-2.0

qwen3-coder-30b-a3b-vllm-cpp
Qwen3-Coder-30B-A3B on vllm.cpp: a coding and agentic-tool-use model, 30B total parameters with about 3B active per token, gated token-exact against vLLM on this engine. The tool-call parser is named explicitly rather than auto-detected, and that matters here. Qwen3-Coder's tool dialect is byte-identical on the wire to another family's, so template sniffing cannot separate the two and would fall back to the wrong parser. With qwen3_coder named, tool calls arrive as real tool_calls on the OpenAI response. This is the bf16 checkpoint, roughly 57 GB of weights, which is what the engine was gated on. Being bf16 rather than NVFP4 it does not need Blackwell on its own account, but LocalAI's CUDA images for this backend are currently built for Blackwell-family GPUs only, so on an older card use the CPU build.

Repository: localaiLicense: apache-2.0

lfm2.5-1.2b-nova-function-calling
The **LFM2.5-1.2B-Nova-Function-Calling-GGUF** is a quantized version of the original model, optimized for efficiency with **Unsloth**. It supports text and multimodal tasks, using different quantization levels (e.g., Q2_K, Q3_K, Q4_K, etc.) to balance performance and memory usage. The model is designed for function calling and is faster than the original version, making it suitable for tasks like code generation, reasoning, and multi-modal input processing.

Repository: localaiLicense: apache-2.0

steelskull_l3.3-shakudo-70b
L3.3-Shakudo-70b is the result of a multi-stage merging process by Steelskull, designed to create a powerful and creative roleplaying model with a unique flavor. The creation process involved several advanced merging techniques, including weight twisting, to achieve its distinct characteristics. Stage 1: The Cognitive Foundation & Weight Twisting The process began by creating a cognitive and tool-use focused base model, L3.3-Cogmoblated-70B. This was achieved through a `model_stock` merge of several models known for their reasoning and instruction-following capabilities. This base was built upon `nbeerbower/Llama-3.1-Nemotron-lorablated-70B`, a model intentionally "ablated" to skew refusal behaviors. This technique, known as weight twisting, helps the final model adopt more desirable response patterns by building upon a foundation that is already aligned against common refusal patterns. Stage 2: The Twin Hydrargyrum - Flavor and Depth Two distinct models were then created from the Cogmoblated base: L3.3-M1-Hydrargyrum-70B: This model was merged using `SCE`, a technique that enhances creative writing and prose style, giving the model its unique "flavor." The Top_K for this merge were set at 0.22 . L3.3-M2-Hydrargyrum-70B: This model was created using a `Della_Linear` merge, which focuses on integrating the "depth" of various roleplaying and narrative models. The settings for this merge were set at: (lambda: 1.1) (weight: 0.2) (density: 0.7) (epsilon: 0.2) Final Stage: Shakudo The final model, L3.3-Shakudo-70b, was created by merging the two Hydrargyrum variants using a 50/50 `nuslerp`. This final step combines the rich, creative prose (flavor) from the SCE merge with the strong roleplaying capabilities (depth) from the Della_Linear merge, resulting in a model with a distinct and refined narrative voice. A special thank you to Nectar.ai for their generous support of the open-source community and my projects. Additionally, a heartfelt thanks to all the Ko-fi supporters who have contributed—your generosity is deeply appreciated and helps keep this work going and the Pods spinning.

Repository: localaiLicense: llama3.3