Model Gallery

17 models from 1 repositories

Filter by type:

Filter by tags:

insightface-antelopev2
Largest insightface pack (SCRFD-10GF + ResNet100@Glint360K recognizer + genderage, ~407MB). Higher recognition accuracy than `buffalo_l` on harder benchmarks; pays for it in GPU memory. NON-COMMERCIAL RESEARCH USE ONLY.

Repository: localaiLicense: insightface-non-commercial

face-detect-antelopev2
Face recognition with insightface's `antelopev2` pack (SCRFD-10G detector + ArcFace glint360k R100, 512-d embedder), converted to a C++/ggml GGUF for the `face-detect` backend. The higher-accuracy insightface pack: heavier, but the best fit when recognition quality matters more than speed. The architecture (`facedetect.arch`) is read from the GGUF metadata, so this entry alone selects the antelopev2 engine. If this GGUF embeds the MiniFASNet anti-spoof ensemble, it is available via the FaceVerify `anti_spoof` request flag. NON-COMMERCIAL RESEARCH USE ONLY.

Repository: localaiLicense: insightface-non-commercial

wespeaker-resnet34
Speaker recognition with WeSpeaker's ResNet34 trained on VoxCeleb, exported to ONNX. 256-d embeddings, CPU-friendly — avoids the PyTorch runtime entirely (onnxruntime only). APACHE 2.0. Pair with the `speaker-recognition` backend's OnnxDirectEngine. Use when ECAPA-TDNN's torch dependency is undesirable (small images, edge deployments).

Repository: localaiLicense: cc-by-4.0

voice-detect-wespeaker-resnet34
Speaker recognition with WeSpeaker's ResNet34 trained on VoxCeleb, converted to a C++/ggml GGUF for the `voice-detect` backend. 256-d embeddings, CPU-friendly and runtime-free (no onnxruntime or torch). CC-BY-4.0. Use when you want WeSpeaker's ResNet34 topology instead of ECAPA-TDNN. The embedding architecture (`voicedetect.arch`) is read from the GGUF metadata, so this entry alone selects the engine.

Repository: localaiLicense: cc-by-4.0

qwen3-8b-shiningvaliant3
Shining Valiant 3 is a science, AI design, and general reasoning specialist built on Qwen 3. Finetuned on our newest science reasoning data generated with Deepseek R1 0528! AI to build AI: our high-difficulty AI reasoning data makes Shining Valiant 3 your friend for building with current AI tech and discovering new innovations and improvements! Improved general and creative reasoning to supplement problem-solving and general chat performance. Small model sizes allow running on local desktop and mobile, plus super-fast server inference!

Repository: localaiLicense: apache-2.0

qwen3-stargate-sg1-uncensored-abliterated-8b-i1
This repo contains the full precision source code, in "safe tensors" format to generate GGUFs, GPTQ, EXL2, AWQ, HQQ and other formats. The source code can also be used directly. This model is specifically for SG1 (Stargate Series), science fiction, story generation (all genres) but also does coding and general tasks too. This model can also be used for Role play. This model will produce uncensored content (see notes below). Fine tune (6 epochs, using Unsloth for Win 11) on an inhouse generated dataset to simulate / explore the Stargate SG1 Universe. This version has the "canon" of all 10 seasons of SG1. Model also contains, but not trained, on content from Stargate Atlantis, and Universe. Fine tune process adds knowledge to the model, and alter all aspects of its operations. Float32 (32 bit precision) was used to further increase the model's quality. This model is based on "Goekdeniz-Guelmez/Josiefied-Qwen3-8B-abliterated-v1". Example generations at the bottom of this page. This is a Stargate (SG1) fine tune (1,331,953,664 of 9,522,689,024 (13.99% trained)), SIX epochs on this model. As this is an instruct model, it will also benefit from a detailed system prompt too.

Repository: localaiLicense: apache-2.0

steelskull_l3.3-electra-r1-70b
L3.3-Electra-R1-70b is the newest release of the Unnamed series, this is the 6th iteration based of user feedback. Built on a custom DeepSeek R1 Distill base (TheSkullery/L3.1x3.3-Hydroblated-R1-70B-v4.4), Electra-R1 integrates specialized components through the SCE merge method. The model uses float32 dtype during processing with a bfloat16 output dtype for optimized performance. Electra-R1 serves newest gold standard and baseline. User feedback consistently highlights its superior intelligence, coherence, and unique ability to provide deep character insights. Through proper prompting, the model demonstrates advanced reasoning capabilities and unprompted exploration of character inner thoughts and motivations. The model utilizes the custom Hydroblated-R1 base, created for stability and enhanced reasoning. The SCE merge method's settings are precisely tuned based on extensive community feedback (of over 10 diffrent models from Nevoria to Cu-Mai), ensuring optimal component integration while maintaining model coherence and reliability. This foundation establishes Electra-R1 as the benchmark upon which its variant models build and expand.

Repository: localaiLicense: eva-llama3.3

llama-3.1-8b-stheno-v3.4-iq-imatrix
This model has went through a multi-stage finetuning process. - 1st, over a multi-turn Conversational-Instruct - 2nd, over a Creative Writing / Roleplay along with some Creative-based Instruct Datasets. - - Dataset consists of a mixture of Human and Claude Data. Prompting Format: - Use the L3 Instruct Formatting - Euryale 2.1 Preset Works Well - Temperature + min_p as per usual, I recommend 1.4 Temp + 0.2 min_p. - Has a different vibe to previous versions. Tinker around. Changes since previous Stheno Datasets: - Included Multi-turn Conversation-based Instruct Datasets to boost multi-turn coherency. # This is a separate set, not the ones made by Kalomaze and Nopm, that are used in Magnum. They're completely different data. - Replaced Single-Turn Instruct with Better Prompts and Answers by Claude 3.5 Sonnet and Claude 3 Opus. - Removed c2 Samples -> Underway of re-filtering and masking to use with custom prefills. TBD - Included 55% more Roleplaying Examples based of [Gryphe's](https://huggingface.co/datasets/Gryphe/Sonnet3.5-Charcard-Roleplay) Charcard RP Sets. Further filtered and cleaned on. - Included 40% More Creative Writing Examples. - Included Datasets Targeting System Prompt Adherence. - Included Datasets targeting Reasoning / Spatial Awareness. - Filtered for the usual errors, slop and stuff at the end. Some may have slipped through, but I removed nearly all of it. Personal Opinions: - Llama3.1 was more disappointing, in the Instruct Tune? It felt overbaked, atleast. Likely due to the DPO being done after their SFT Stage. - Tuning on L3.1 base did not give good results, unlike when I tested with Nemo base. unfortunate. - Still though, I think I did an okay job. It does feel a bit more distinctive. - It took a lot of tinkering, like a LOT to wrangle this.

Repository: localaiLicense: cc-by-nc-4.0

qwen3-6b-almost-human-xmen-x4-x2-x1-dare-e32
**Model Name:** Qwen3-6B-Almost-Human-XMEN-X4-X2-X1-Dare-e32 **Author:** DavidAU (based on original Qwen3-6B architecture) **Repository:** [DavidAU/Qwen3-Almost-Human-XMEN-X4-X2-X1-Dare-e32](https://huggingface.co/DavidAU/Qwen3-Almost-Human-XMEN-X4-X2-X1-Dare-e32) **Base Model:** Qwen3-6B (original Qwen3 6B from Alibaba) **License:** Apache 2.0 **Quantization Status:** Full-precision (float32) source model available; GGUF quantizations also provided by third parties (e.g., mradermacher) --- ### 🌟 Model Description **Qwen3-6B-Almost-Human-XMEN-X4-X2-X1-Dare-e32** is a creatively enhanced, instruction-tuned variant of the Qwen3-6B model, meticulously fine-tuned to emulate the literary voice and psychological depth of **Philip K. Dick**. Developed by DavidAU using **Unsloth** and trained on multiple proprietary datasets—including works of PK Dick, personal notes, letters, and creative writing—this model excels in **narrative richness, emotional nuance, and complex reasoning**. It is the result of a **"DARE-TIES" merge** combining four distinct training variants: X4, X2, and two X1 models, with the final fusion mastered in **32-bit precision (float32)** for maximum fidelity. The model incorporates **Brainstorm 20x**, a novel reasoning enhancement technique that expands and recalibrates the model’s internal reasoning centers 20 times to improve coherence, detail, and creative depth—without compromising instruction-following. --- ### ✨ Key Features - **Enhanced Prose & Storytelling:** Generates vivid, immersive, and deeply human-like narratives with foreshadowing, similes, metaphors, and emotional engagement. - **Strong Reasoning & Creativity:** Ideal for brainstorming, roleplay, long-form writing, and complex problem-solving. - **High Context (256K):** Supports extensive conversations and long-form content. - **Optimized for Creative & Coding Tasks:** Performs exceptionally well with detailed prompts and step-by-step refinement. - **Full-Precision Source Available:** Original float32 model is provided—ideal for advanced users and model developers. --- ### 🛠️ Recommended Use Cases - Creative writing & fiction generation - Roleplaying and character-driven dialogue - Complex brainstorming and ideation - Code generation with narrative context - Literary and philosophical exploration > 🔍 **Note:** The GGUF quantized version (e.g., by mradermacher) is **not the original**—it’s a derivative. For the **true base model**, use the **DavidAU/Qwen3-Almost-Human-X1-6B-e32** repository, which hosts the original, full-precision model. --- ### 📌 Tips for Best Results - Use **CHATML or Jinja templates** - Set `temperature: 0.3–0.7`, `top_p: 0.8`, `repetition_penalty: 1.05–1.1` - Enable **smoothing factor (1.5)** in tools like KoboldCpp or Text-Gen-WebUI for smoother output - Use **Q6 or Q8 GGUF quants** for best performance on complex tasks --- ✨ **In short:** A poetic, introspective, and deeply human-like AI—crafted to feel like a real mind, not just a machine. Perfect for those who want **intelligence with soul**.

Repository: localaiLicense: apache-2.0

almost-human-x3-32bit-1839-6b-i1
**Model Name:** Almost-Human-X3-32bit-1839-6B **Base Model:** Qwen3-Jan-v1-256k-ctx-6B-Brainstorm20x **Author:** DavidAU **Repository:** [DavidAU/Almost-Human-X3-32bit-1839-6B](https://huggingface.co/DavidAU/Almost-Human-X3-32bit-1839-6B) **License:** Apache 2.0 --- ### 🔍 **Overview** A high-precision, full-precision (float32) fine-tuned variant of the Qwen3-Jan model, specifically trained to emulate the literary and philosophical depth of Philip K. Dick. This model is the third in the "Almost-Human" series, built with advanced **"Brainstorm 20x"** methodology to enhance reasoning, coherence, and narrative quality—without sacrificing instruction-following ability. ### 🎯 **Key Features** - **Full Precision (32-bit):** Trained at 16-bit for 3 epochs, then finalized at float32 for maximum fidelity and performance. - **Extended Context (256k tokens):** Ideal for long-form writing, complex reasoning, and detailed code generation. - **Advanced Reasoning via Brainstorm 20x:** The model’s reasoning centers are expanded, calibrated, and interconnected 20 times, resulting in: - Richer, more nuanced prose - Stronger emotional engagement - Deeper narrative focus and foreshadowing - Fewer clichés, more originality - Enhanced coherence and detail - **Optimized for Creativity & Code:** Excels at brainstorming, roleplay, storytelling, and multi-step coding tasks. ### 🛠️ **Usage Tips** - Use **CHATML or Jinja templates** for best results. - Recommended settings: Temperature 0.3–0.7 (higher for creativity), Top-p 0.8, Repetition penalty 1.05–1.1. - Best used with **"smoothing" (1.5)** in GUIs like KoboldCpp or oobabooga. - For complex tasks, use **Q6 or Q8 GGUF quantizations**. ### 📦 **Model Formats** - **Full precision (safe tensors)** – for training or high-fidelity inference - **GGUF, GPTQ, EXL2, AWQ, HQQ** – available via quantization (see [mradermacher/Almost-Human-X3-32bit-1839-6B-i1-GGUF](https://huggingface.co/mradermacher/Almost-Human-X3-32bit-1839-6B-i1-GGUF) for quantized versions) --- ### 💬 **Ideal For** - Creative writing, speculative fiction, and philosophical storytelling - Complex code generation with deep reasoning - Roleplay, character-driven dialogue, and immersive narratives - Researchers and developers seeking a highly expressive, human-like model > 📌 **Note:** This is the original source model. The GGUF versions by mradermacher are quantized derivatives — not the base model. --- **Explore the source:** [DavidAU/Almost-Human-X3-32bit-1839-6B](https://huggingface.co/DavidAU/Almost-Human-X3-32bit-1839-6B) **Quantization guide:** [mradermacher/Almost-Human-X3-32bit-1839-6B-i1-GGUF](https://huggingface.co/mradermacher/Almost-Human-X3-32bit-1839-6B-i1-GGUF)

Repository: localaiLicense: apache-2.0

parakeet-cpp-nemotron-3-diarization-speakers
Nemotron-3-Diarization (Sortformer) with WeSpeaker ResNet34 speaker identification, for the parakeet-cpp backend. Speakers you register with /v1/voice/register (using the voice-detect-wespeaker-resnet34 model) come back by name in /v1/audio/diarization, next to the SPEAKER_NN label. Speakers that are not registered keep only their SPEAKER_NN label. The diarization model is OpenMDW-1.1, the speaker model is CC-BY-4.0. Naming was measured on one two-voice fixture only; check the threshold on your own audio.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-nemotron-3-diarization-asr-speakers
Nemotron-3-Diarization (Sortformer) paired with the Parakeet TDT+CTC 110M ASR model through the asr_model option, both Q8_0/F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Served through /v1/audio/diarization with include_text: each speaker segment comes back with its transcribed text in one call. Diarization model is OpenMDW-1.1, ASR model is CC-BY-4.0. Also loads WeSpeaker ResNet34 (CC-BY-4.0) through the speaker_model option: speakers registered with /v1/voice/register (voice-detect-wespeaker-resnet34 model) come back by name, next to the SPEAKER_NN label.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-multilingual-diarization-speakers
Parakeet TDT 0.6B v3 (multilingual, 25 European languages) paired with Nemotron-3-Diarization (Sortformer) through the diarization_model option and WeSpeaker ResNet34 through the speaker_model option, all GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). One call to /v1/audio/diarization with include_text and include_speaker_profiles returns the speaker turns, the text of each turn in the spoken language, and one voice-print embedding per speaker, so a client does not need a separate diarization call and transcription call. Use it where the English-only 110M ASR of parakeet-cpp-nemotron-3-diarization-asr-speakers is not enough. Also serves /v1/audio/transcriptions. License per model: transcription model CC-BY-4.0, diarization model OpenMDW-1.1, WeSpeaker encoder CC-BY-4.0.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-realtime-scene-speakers
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0, WeSpeaker ResNet34 CC-BY-4.0. Also loads WeSpeaker ResNet34 through the speaker_model option, so live speaker segments carry the name of a voice registered with /v1/voice/register (voice-detect-wespeaker-resnet34 model) once the speaker is identified.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-bundle-small
Parakeet TDT+CTC 110M with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 338 MB) that holds five models for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT+CTC 110M (English, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true option cuts long audio at pauses found by Silero before it transcribes. The bundle options (diar_component, sound_component, speaker_component) load each model from the same file; see the audio-to-text docs for the option list. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet TDT+CTC 110M by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0 (as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained), WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The weights were converted to GGUF and quantised where stated; nothing was retrained.

Repository: localaiLicense: other

parakeet-cpp-bundle-standard
Parakeet TDT 0.6B v3 (multilingual) with diarization, sound events, speaker naming and VAD. One bundle GGUF file (about 1101 MB) that holds five models for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet): Parakeet TDT 0.6B v3 (multilingual, Q8_0), Nemotron-3-Diarization (Q8_0), CED-Small sound events (Q8_0), WeSpeaker ResNet34-LM speaker encoder (F32) and Silero VAD (F16). One install serves transcription (/v1/audio/transcriptions), voice activity detection (/v1/vad), speaker diarization with text and named speakers (/v1/audio/diarization) and sound events (/v1/audio/classification). The vad:true option cuts long audio at pauses found by Silero before it transcribes. The bundle options (diar_component, sound_component, speaker_component) load each model from the same file; see the audio-to-text docs for the option list. A bundle has no single license: each model keeps its own, listed in the file header and in the NOTICE file next to the file. Parakeet TDT 0.6B v3 by NVIDIA is CC-BY-4.0, Nemotron-3-Diarization by NVIDIA is OpenMDW-1.1, CED-Small is Apache-2.0 (as stated on the model card; the upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream; the model here is converted, not trained), WeSpeaker ResNet34-LM by the WeSpeaker project is CC-BY-4.0, Silero VAD by the Silero Team is MIT. The weights were converted to GGUF and quantised where stated; nothing was retrained.

Repository: localaiLicense: other

cua-s1-forms-vllm-cpp
cua-s1-forms is a small, jev-like ("System One") one-pass option scorer for GUI form filling. Given a UI element and a list of typed options, it returns one probability per option in a single forward pass. Every actionable element on a form is scored independently and in parallel. In LocalAI, serve via POST /api/score with this model. The vllm.cpp engine runs the scoring pipeline through the vllm_decide C ABI. 2.8 MB, float32, CPU and GPU.

Repository: localaiLicense: mit