Model Gallery

257 models from 1 repositories

Filter by type:

Filter by tags:

deepseek-v4-flash-0731
# DeepSeek-V4-Flash-0731 Technical Report๐Ÿ‘๏ธ ## Introduction **DeepSeek-V4-Flash-0731** is the official release of **DeepSeek-V4-Flash**, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available. Notes: 1. For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the `max` reasoning effort level with `temperature = 1.0, top_p = 0.95`. 2. โ€  DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems. ## Chat Template ...

Repository: localaiLicense: mit

pocket-35b-q3
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the balanced Q3_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry offers standard Q4_K_M and Q8_0 GGUF quantizations alongside APEX mixed-precision variants with quality, balanced, compact, and mini profiles. The APEX I-profiles use importance-matrix calibration.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev-apex-i-quality
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry offers standard Q4_K_M and Q8_0 GGUF quantizations alongside APEX mixed-precision variants with quality, balanced, compact, and mini profiles. The APEX I-profiles use importance-matrix calibration.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev-apex-i-balanced
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry offers standard Q4_K_M and Q8_0 GGUF quantizations alongside APEX mixed-precision variants with quality, balanced, compact, and mini profiles. The APEX I-profiles use importance-matrix calibration.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev-apex-i-compact
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry offers standard Q4_K_M and Q8_0 GGUF quantizations alongside APEX mixed-precision variants with quality, balanced, compact, and mini profiles. The APEX I-profiles use importance-matrix calibration.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev-apex-i-mini
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry offers standard Q4_K_M and Q8_0 GGUF quantizations alongside APEX mixed-precision variants with quality, balanced, compact, and mini profiles. The APEX I-profiles use importance-matrix calibration.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev-apex-quality
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry offers standard Q4_K_M and Q8_0 GGUF quantizations alongside APEX mixed-precision variants with quality, balanced, compact, and mini profiles. The APEX I-profiles use importance-matrix calibration.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev-apex-balanced
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry offers standard Q4_K_M and Q8_0 GGUF quantizations alongside APEX mixed-precision variants with quality, balanced, compact, and mini profiles. The APEX I-profiles use importance-matrix calibration.

Repository: localaiLicense: apache-2.0

kat-coder-v2.5-dev-apex-compact
KAT-Coder-V2.5-Dev is an Apache-2.0 agentic coding model from Kwaipilot, post-trained from Qwen3.6-35B-A3B. It has 35 billion total parameters with 3 billion activated per token, a 262K-token context window, and text-only weights tuned for repository-level coding and tool use. This entry offers standard Q4_K_M and Q8_0 GGUF quantizations alongside APEX mixed-precision variants with quality, balanced, compact, and mini profiles. The APEX I-profiles use importance-matrix calibration.

Repository: localaiLicense: apache-2.0

minicpm5-1b-claude-opus-fable5-v2-thinking
# MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking GGUF quantizations for local deployment: **MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF** ไธญๆ–‡่ฏดๆ˜Ž **MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking** is a compact 1B **Thinking** language model built on openbmb/MiniCPM5-1B. Compared with V1, this V2 release is further fine-tuned on **Fable 5** data with a stronger focus on **tool calling / function calling**, while also improving **coding** and **instruction-following**. It keeps MiniCPM5's native Thinking chat template and XML tool-call format. Previous version: **MiniCPM5-1B-Claude-Opus-Fable5-Thinking** (V1) For llama.cpp / Ollama / LM Studio deployment, see the **GGUF repository**. ## Overview ## Capabilities - **Tool calling (enhanced in V2)** โ€” more reliable XML / function-calling style tool use on top of MiniCPM5's native format - **Coding** โ€” code generation, debugging, and software-engineering-style tasks - **Instruction following** โ€” more reliable adherence to user prompts and structured constraints - **Thinking mode** โ€” chain-of-thought reasoning via the MiniCPM5 chat template - **Long context** โ€” up to **128K tokens** (131,072 tokens per `config.json`) ...

Repository: localaiLicense: apache-2.0

qwopus3.6-35b-a3b-coder-mtp
# ๐ŸŒŸ Qwopus3.6-35B-A3B-v1 ## ๐Ÿ’ก Base Model Overview **Qwen3.6-35B-A3B** is an advanced hybrid sparse MoE (Mixture-of-Experts) model developed by Alibaba Cloud. It features 35B total parameters with only 3B active parameters per token, ensuring high inference efficiency. Architecturally, it combines Gated DeltaNet linear attention with standard gated attention layers, routing tokens across **256 experts**. It natively supports a massive **262k context window** and is specifically designed for high-performance agentic coding, deep reasoning, and multimodal tasks. ## ๐Ÿš€ Model Refinement & Logic Tuning ๏ผˆQwopus3.6-35B-A3B-v1๏ผ‰ ๐Ÿช**Qwopus3.6-35B-A3B-v1** is a reasoning-enhanced MoE (Mixture of Experts) model fine-tuned on top of **Qwen3.6-35B-A3B**. ### ๐Ÿ›  Training Strategy The fine-tuning process for this model is structured into **three distinct stages of distributed SFT (Supervised Fine-Tuning)**, progressively scaling reasoning complexity and data diversity. This systematic approach ensures the model inherits the base MoE capabilities while sharpening its logic-handling depth. ...

Repository: localaiLicense: apache-2.0

laguna-xs-2.1-apex-i-balanced
Laguna XS 2.1 in the 24.3 GB APEX-I Balanced format, an importance-matrix build balancing fidelity and memory use for llama.cpp. License: OpenMDW 1.1.

Repository: localaiLicense: other

laguna-xs-2.1-apex-balanced
Laguna XS 2.1 in the 24.3 GB APEX Balanced format, balancing fidelity and memory use for llama.cpp. License: OpenMDW 1.1.

Repository: localaiLicense: other

qwopus3.6-27b-coder-compat-mtp
๐Ÿช Qwopus-3.6-27B-Coder Coder SFT Release Agentic Coding & Tool-Use Reasoning Model Fine-Tuned on Qwopus3.6-27B-v2 ๐Ÿงฌ Trace Inversion & Negentropy ๐Ÿง  27B Dense Model โšก Agentic Coding ๐Ÿ› ๏ธ Tool Calling & Agent ๐Ÿ† SWE-bench Verified: 67.0% (off-thinking) ๐Ÿ’ก What is Qwopus-3.6-27B-Coder? ๐Ÿช Qwopus-3.6-27B-Coder is a reasoning-enhanced agentic coding model built on top of Qwopus3.6-27B-v2. It inherits the powerful reasoning foundation of the v2 base โ€” which achieved 87.43% MMLU-Pro and 75.25% SWE-bench Verified โ€” and further specializes it for agentic code generation, structured tool calling, debugging, and instruction-following in developer workflows. The model is designed to excel at repository-level coding tasks, multi-turn tool orchestration, and complex logical reasoning under realistic agent environments. ๐Ÿงฉ Agentic Coding Optimized for repository-level coding, debugging, patch generation, and structured multi-step development workflows. ๐Ÿ› ๏ธ Tool Calling Learns from real agent trajectories with tool definitions, tool calls, and environment feedback for robust multi-turn execution. ...

Repository: localaiLicense: apache-2.0

qwen3.5-9b-hauhaucs-aggressive
Qwen3.5 9B Aggressive is HauhauCS's refusal-removed fine-tune of the multimodal Qwen3.5 9B model. It retains the base model's reasoning, tool use, image and video understanding, and 262K-token native context window. This entry uses the balanced Q4_K_M GGUF quantization and includes the matching BF16 multimodal projector. The Q8_0 variant offers higher fidelity.

Repository: localaiLicense: apache-2.0

glm-5.2
# GLM-5.2 ๐Ÿ‘‹ Join our WeChat or Discord community. ๐Ÿ“– Check out the GLM-5.2 blog and GLM-5 Technical report. ๐Ÿ“ Use GLM-5.2 API services on Z.ai API Platform. ๐Ÿ”œ Try GLM-5.2 here. [Paper] [GitHub] ## Introduction We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**. GLM-5.2's new capabilities include: - **Solid 1M Context:** A solid 1M-token context that stably sustains long-horizon work - **Advanced Coding with Flexible Effort**: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency - **Improved Architecture**: We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9ร— at a 1M context length. We also improve GLM-5.2โ€™s MTP layer for speculative decoding, increasing the acceptance length by up to 20% - **Pure Open**: An MIT open-source license โ€” no regional limits, technical access without borders ## Benchmark ## Serve GLM-5.2 Locally ...

Repository: localaiLicense: mit

qwen3.6-35b-a3b-nvfp4-mtp
# Qwen3.6-35B-A3B [](https://chat.qwen.ai) > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. ## Qwen3.6 Highlights This release delivers substantial upgrades, particularly in - **Agentic Coding:** the model now handles frontend workflows and repository-level reasoning with greater fluency and precision. - **Thinking Preservation:** we've introduced a new option to retain reasoning context from historical messages, streamlining iterative development and reducing overhead. For more details, please refer to our blog post Qwen3.6-35B-A3B. ## Model Overview ...

Repository: localai

qwopus3.6-27b-v2-mtp-nvfp4
๐Ÿช Qwopus3.6-27B-v2-MTP MTP Release Multi-Token Prediction reasoning model fine-tuned from Qwen3.6-27B ๐Ÿงฌ Trace Inversion & Negentropy ๐Ÿง  27B Parameters โšก Speculative Decoding ๐Ÿ› ๏ธ Coding / DevOps / Math ๐Ÿ’ก What is Qwopus3.6-27B-v2-MTP? ๐Ÿช Qwopus3.6-27B-v2-MTP is a speed-oriented reasoning release built on top of Qwen3.6-27B. It keeps the Qwopus line's focus on reconstructed reasoning traces, coding discipline, DevOps procedures, and mathematical derivations, while adding Multi-Token Prediction for faster generation. The goal is simple: preserve the depth and structure of a 27B reasoning model while making real interactive use noticeably faster. โšก MTP DecodingAuxiliary future-token prediction improves throughput on long reasoning, code, math, and strict-format prompts. ๐Ÿงฉ Structured ReasoningInherits the Qwopus training recipe built around reconstructed step-by-step reasoning trajectories. ๐Ÿงช GB10 TestedValidated on a 30-question local benchmark across Logic, Coding, DevOps, Math, and Edge tasks. ๐Ÿš€ Practical SpeedDesigned for workflows where strong answers matter, but waiting several extra minutes per task does not. ...

Repository: localai

qwopus3.6-27b-coder-mtp-nvfp4
๐Ÿช Qwopus-3.6-27B-Coder Coder SFT Release Agentic Coding & Tool-Use Reasoning Model Fine-Tuned on Qwopus3.6-27B-v2 ๐Ÿงฌ Trace Inversion & Negentropy ๐Ÿง  27B Dense Model โšก Agentic Coding ๐Ÿ› ๏ธ Tool Calling & Agent ๐Ÿ† SWE-bench Verified: 67.0% (off-thinking) ๐Ÿ’ก What is Qwopus-3.6-27B-Coder? ๐Ÿช Qwopus-3.6-27B-Coder is a reasoning-enhanced agentic coding model built on top of Qwopus3.6-27B-v2. It inherits the powerful reasoning foundation of the v2 base โ€” which achieved 87.43% MMLU-Pro (300ex) and 75.25% SWE-bench Verified โ€” and further specializes it for agentic code generation, structured tool calling, debugging, and instruction-following in developer workflows. The model is designed to excel at repository-level coding tasks, multi-turn tool orchestration, and complex logical reasoning under realistic agent environments. ๐Ÿงฉ Agentic Coding Optimized for repository-level coding, debugging, patch generation, and structured multi-step development workflows. ๐Ÿ› ๏ธ Tool Calling Learns from real agent trajectories with tool definitions, tool calls, and environment feedback for robust multi-turn execution. ...

Repository: localai

qwen3.6-27b-nvfp4-mtp
# Qwen3.6-27B [](https://chat.qwen.ai) > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. ## Qwen3.6 Highlights This release delivers substantial upgrades, particularly in - **Agentic Coding:** the model now handles frontend workflows and repository-level reasoning with greater fluency and precision. - **Thinking Preservation:** we've introduced a new option to retain reasoning context from historical messages, streamlining iterative development and reducing overhead. For more details, please refer to our blog post Qwen3.6-27B. ## Model Overview ...

Repository: localai

Page 1