Llama.cpp & GGUF Quantization Internals: K-Quants, IQ4_XS, and Metal GPU Offloading (2026)9/12/20265 min read