🔍 Read the full analysis: Using Transformers To Run Quantized Llama.cpp Models on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Hugging Face has added support for loading GGUF quantized checkpoints in Transformers through the from_pretrained API. The initial rollout targets Apple Silicon and Qwen3.5, uses llama.cpp’s ggml kernels when compatible, and is available on the Transformers main branch ahead of a stable release.
Hugging Face has added GGUF model loading to its Transformers library, allowing users to run compatible quantized checkpoints through the familiar from_pretrained API, as described in the original analysis. The initial support targets Apple Silicon Macs and Qwen3.5, and is currently on the library’s main branch ahead of a stable release.
Users can select a GGUF checkpoint from the Hugging Face Hub and pass it to from_pretrained with the gguf_file argument to run local models on device. Hugging Face says the integration reuses llama.cpp’s ggml kernels to keep inference performance close to llama.cpp. When weights remain packed on Apple’s Metal backend, Transformers can load compatible ggml/Metal layer kernels and use ggml-org/ggml-attn for attention.
The setup has specific requirements: an Apple Silicon Mac, a supported PyTorch version, the latest Transformers code and a compatible kernels library. If the compatible quantization kernel is unavailable, the loader can dequantize the model instead, which uses more memory. If the ggml attention kernel cannot be fetched, the system falls back to standard sdpa attention with a warning; users can also select sdpa directly.
The same checkpoints can be served through transformers serve, which provides an OpenAI-compatible API on localhost. Clients such as Jan or Pi can connect through a custom provider. Hugging Face says its performance comparison uses llama.cpp as the reference and covers three GGUF checkpoints: a small dense model, a larger dense model and a mixture-of-experts model.
GGUF Models Reach Transformers
The change makes GGUF checkpoints available within a widely used model-development library, reducing the need for developers to switch to a separate llama.cpp-based application when working with those files. GGUF is used by local inference tools including Ollama, LM Studio and Jan, and the format packages weights and model metadata in a single file.
Quantization can reduce the memory needed to run a model locally. Hugging Face’s example for Unsloth’s Qwen3.5-4B lists a BF16 file at 8.42 GB and a Q4_K_M version at 2.74 GB. Those sizes illustrate the storage difference; they do not establish a particular quality or speed outcome. Hugging Face says the quality effect of lower precision depends on the model and task, so users should assess it against their own workloads.
For developers already using Transformers, loading Hub-hosted GGUF files through the existing API may simplify local experiments and serving workflows. The practical reach of the feature is currently limited by its Apple Silicon focus, architecture coverage and dependency on compatible kernels.
Top picks for "transformer quantiz llama"
As an affiliate, we earn on qualifying purchases.
A New Route for GGUF
GGUF is a format developed by the llama.cpp project for storing model weights and related metadata, including tokenizer information and, optionally, a chat template. It supports multiple quantization levels, which trade memory use against numerical precision. A Q4_K_M checkpoint uses mostly four-bit weights while retaining higher precision for some tensors.
Before this addition, developers generally ran GGUF checkpoints through llama.cpp or tools built around it. Hugging Face’s announcement describes Transformers support as a way to use those checkpoints through its own APIs. The company’s cited example sizes for Qwen3.5-4B are 3.53 GB for Q6_K, 3.14 GB for Q5_K_M and 2.74 GB for Q4_K_M, compared with 8.42 GB for BF16. It recommends starting with Q4_K_M and considering higher-precision variants when memory allows, while advising evaluation on the intended task.
The announcement follows Hugging Face co-founder Julien Chaumond’s public demonstration of Qwen3.6 27B running in the Pi coding agent through llama.cpp on a MacBook Pro. Chaumond said it felt “very, very close” to Claude Opus for non-trivial tasks on Hugging Face codebases. That was his assessment of a demonstration, not an independent benchmark of the new Transformers integration.
“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”
— Hugging Face announcement
Support Still Has Limits
The announcement identifies Apple Silicon and Qwen3.5 as the initial targets. It does not give a timeline for support on CUDA, Linux or Windows, or specify when additional model architectures will be covered. The feature is on the Transformers main branch, and a stable-release date has not been announced.
Hugging Face refers to benchmarks against llama.cpp across three checkpoints, but the available source material does not provide the full figures or enough hardware and test detail to assess how results vary across setups. It is also unclear how quickly compatible kernels will be available for different PyTorch versions. Quantization’s effect on output quality remains dependent on the model and task, according to Hugging Face.
Stable Release and Wider Coverage
The next practical milestone is inclusion in a stable Transformers release, which would let users install the feature without using the main branch. Hugging Face has not announced when that release will arrive. Until then, users need the current main-branch code and compatible kernel builds to use the packed-weight path.
Further expansion could add hardware backends and model architectures, but the announcement provides no schedule or confirmed roadmap for those additions. Users can follow Hugging Face’s GGUF documentation and kernels library for updates on supported models, quantization types and dependencies.
Key Questions
What did Hugging Face add?
It added support for loading GGUF quantized checkpoints in Transformers using the from_pretrained API, with a GGUF file passed through the gguf_file argument.
Which devices and models are supported initially?
The announced initial support targets Apple Silicon Macs and the Qwen3.5 architecture. The source does not give a schedule for other hardware or architectures.
Does Transformers use llama.cpp kernels?
Hugging Face says the integration reuses llama.cpp’s ggml kernels for performance. If compatible quantization kernels are unavailable, the loader can dequantize the model, using more memory.
Is the feature in a stable Transformers release?
No. It is currently available on the Transformers main branch. Hugging Face has not announced a date for a stable release containing the feature.
How much can quantization reduce model file size?
In Hugging Face’s Qwen3.5-4B example, the listed size falls from 8.42 GB in BF16 to 2.74 GB in Q4_K_M. The resulting quality depends on the model and task, so users should evaluate their own workload.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
