Model Gallery

25 models from 1 repositories

Filter by type:

Filter by tags:

ternary-bonsai-2-27b
Ternary Bonsai 2 27B (PrismML) is a 27B-class reasoning model with ternary transformer weights. This PTQ1_0 build packs the trits densely at 1.75 bits per weight (5.95 GB) and includes the Q8_0 vision projector. PTQ1_0 is a Prism-private GGUF type, so the entry uses the bonsai backend (PrismML's llama.cpp fork) instead of stock llama.cpp.

Repository: localaiLicense: apache-2.0

ternary-bonsai-8b
Ternary Bonsai 8B (PrismML) is a 1.58-bit ternary language model on the Qwen3-8B dense architecture. Each weight takes a value from {-1, 0, +1} with one shared FP16 scale per group of 128 weights (GGUF Q2_0, ~2.18 GB deployed, 7.5x smaller than FP16). The extra zero state recovers more of the full-precision model than the 1-bit build: it ranks 2nd among compared 6-9B models at 75.5 average despite being ~1/8th their size. Q2_0 is the recommended, ternary-lossless variant. The Q2_0 kernels are only in the PrismML llama.cpp fork, so this runs on LocalAI's `bonsai` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

ternary-bonsai-8b-q2-g64
Ternary Bonsai 8B (PrismML), GGUF Q2_0 with group-64 packing (each FP16 scale shared across 64 weights instead of 128). Slightly larger (~2.31 GB) but matches llama.cpp's native 64-value Q2_0 block layout. Runs on LocalAI's `bonsai` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

ternary-bonsai-8b-pq2
Ternary Bonsai 8B (PrismML), GGUF PQ2_0 (packed Q2_0) ternary variant (~2.18 GB). Same {-1, 0, +1} weight alphabet as Q2_0. Runs on LocalAI's `bonsai` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

ternary-bonsai-27b
Ternary Bonsai 27B (PrismML) is the quality-oriented operating point of the Bonsai 27B family: full 27B-class reasoning in ternary {-1, 0, +1} weights on the Qwen3.6-27B hybrid-attention backbone (262K context). At a true 1.71 bits/weight it deploys in ~7.2 GB (GGUF Q2_0_g128) and retains 95% of FP16 intelligence (80.49 average across 15 thinking-mode benchmarks) - a higher score than a conventional IQ2_XXS build at less than two-thirds its footprint. Ships an optional 4-bit vision tower (mmproj), included. The Q2_0 weights and hybrid-attention kernels are only in the PrismML llama.cpp fork, so this runs on LocalAI's `bonsai` backend. A GPU is recommended. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

ternary-bonsai-27b-pq2
Ternary Bonsai 27B (PrismML), GGUF PQ2_0 (packed Q2_0) ternary variant (~7.17 GB) with the 4-bit vision tower (mmproj) included. Runs on LocalAI's `bonsai` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

ternary-bonsai-27b-q2-g64
Ternary Bonsai 27B (PrismML), GGUF Q2_0 with group-64 packing (~7.59 GB), matching llama.cpp's native 64-value Q2_0 block layout, with the 4-bit vision tower (mmproj) included. Runs on LocalAI's `bonsai` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

maple-preview-tq1-0-head-q4-k
Maple-Preview is DeepGrove's 20B mixture-of-experts reasoning model with about 1B active parameters. This build uses TQ1_0 ternary GGUF weights with a Q4_K output head. It uses the publisher's CPU configuration and embedded Jinja chat template with an 8K default context. The native context is 128K tokens.

Repository: localaiLicense: mit

maple-preview-tq1-0-head-f16
Maple-Preview is DeepGrove's 20B mixture-of-experts reasoning model with about 1B active parameters. This build uses TQ1_0 ternary GGUF weights with a F16 output head. It uses the publisher's CPU configuration and embedded Jinja chat template with an 8K default context. The native context is 128K tokens.

Repository: localaiLicense: mit

maple-preview-tq2-0-head-q4-k
Maple-Preview is DeepGrove's 20B mixture-of-experts reasoning model with about 1B active parameters. This build uses TQ2_0 ternary GGUF weights with a Q4_K output head. It uses the publisher's CPU configuration and embedded Jinja chat template with an 8K default context. The native context is 128K tokens.

Repository: localaiLicense: mit

maple-preview-tq2-0-head-f16
Maple-Preview is DeepGrove's 20B mixture-of-experts reasoning model with about 1B active parameters. This build uses TQ2_0 ternary GGUF weights with a F16 output head. It uses the publisher's CPU configuration and embedded Jinja chat template with an 8K default context. The native context is 128K tokens.

Repository: localaiLicense: mit

parakeet-cpp-moondream-redux-packed
Moondream Redux, packed ternary weights (about 213 MB). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B). CPU only and offline only: the library refuses to load it on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). The model has a voice-activity-detection head, and the vad:true option cuts long audio at pauses before transcription. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-vad-moondream-redux-packed
Voice activity detection with the VAD head of Moondream Redux, packed ternary weights (about 213 MB). CPU only. The file is the same as the ASR entry parakeet-cpp-moondream-redux-packed, so the two share it on disk. Use it for the VAD endpoint. License CC-BY-4.0: credit Moondream and NVIDIA.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-bundle-moondream-redux
Moondream Redux with Silero VAD in one bundle GGUF file (about 215 MB) for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet). Redux is Moondream's ternary-encoder derivative of NVIDIA parakeet-tdt-0.6b-v3 (TDT, 0.6B), kept as published with the packed ternary weights. The bundle also holds Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions) and voice activity detection (/v1/vad). The vad:true option cuts long audio at pauses before it transcribes, with Silero as the VAD source. CPU only and offline only: the library refuses the packed Redux model on a GPU backend and does not stream it. For a GPU or for streaming, use the redux-f16 or redux-q8_0 entry instead. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet Redux by Moondream, derived from Parakeet TDT 0.6B v3 by NVIDIA, is CC-BY-4.0 (credit Moondream and NVIDIA); Silero VAD by the Silero Team is MIT. The weights were converted to GGUF; nothing was retrained.

Repository: localaiLicense: other

vibevoice-asr-bitnet-crispasr
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model for English, Chinese, French, Italian, Korean, Portuguese, and Vietnamese. Its BitNet-trained language-model projections use ternary TQ2_0 weights. This recommended CrispASR build keeps the VAE encoder in Q8_0 and the embedding in F16, and is approximately 1.55 GB.

Repository: localaiLicense: mit

vibevoice-asr-bitnet-embed-q8-crispasr
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model. This approximately 1.34 GB CrispASR build uses ternary TQ2_0 language-model projections, a Q8_0 VAE encoder, and a Q8_0 embedding.

Repository: localaiLicense: mit

vibevoice-asr-bitnet-vae-q5-crispasr
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model. This approximately 1.33 GB CrispASR build uses ternary TQ2_0 language-model projections, a Q5_0 VAE encoder, and an F16 embedding.

Repository: localaiLicense: mit

vibevoice-asr-bitnet-both-q5-crispasr
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model. This approximately 1.12 GB CrispASR build uses ternary TQ2_0 language-model projections, a Q5_0 VAE encoder, and a Q8_0 embedding.

Repository: localaiLicense: mit

vibevoice-asr-bitnet-vae-q4-crispasr
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model. This approximately 1.26 GB CrispASR build uses ternary TQ2_0 language-model projections, a Q4_0 VAE encoder, and an F16 embedding.

Repository: localaiLicense: mit

vibevoice-asr-bitnet-aggro-crispasr
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model for English, Chinese, French, Italian, Korean, Portuguese, and Vietnamese. This compact CrispASR build combines ternary TQ2_0 language-model projections with a Q4_0 VAE encoder and Q8_0 embedding, reducing the model to approximately 1.05 GB while producing the same JFK benchmark transcription as the larger builds published alongside it.

Repository: localaiLicense: mit

swift-bonsai-2-pq2
Swift Bonsai 2 is ukisai's experimental 27B reasoning model derived from Ternary Bonsai 2. This PQ2_0 GGUF contains the updated merged Swift weights and runs text chat through LocalAI's Bonsai backend.

Repository: localaiLicense: apache-2.0

Page 1