qwen3-tts-llamacpp
Qwen3-TTS 1.7B Base served by the llama.cpp backend, using upstream's own
GGUF conversion. Runs on the full llama-cpp accelerator matrix (CUDA, ROCm,
SYCL, Vulkan, Metal). Streaming output and zero-shot voice cloning: set
`voice` to a reference clip or a saved Voice Library profile, which is
required since the Base checkpoint has no built-in speaker. 24kHz mono,
10 languages. Q8_0 backbone (~1.8 GB) plus a Q8_0 projector.