AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Using Transformers To Run Quantized Llama.cpp Models on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face has added support for loading GGUF quantized checkpoints in Transformers through the from_pretrained API. The initial rollout targets Apple Silicon and Qwen3.5, uses llama.cpp’s ggml kernels when compatible, and is available on the Transformers main branch ahead of a stable release.

Hugging Face has added GGUF model loading to its Transformers library, allowing users to run compatible quantized checkpoints through the familiar from_pretrained API, as described in the original analysis. The initial support targets Apple Silicon Macs and Qwen3.5, and is currently on the library’s main branch ahead of a stable release.

Users can select a GGUF checkpoint from the Hugging Face Hub and pass it to from_pretrained with the gguf_file argument to run local models on device. Hugging Face says the integration reuses llama.cpp’s ggml kernels to keep inference performance close to llama.cpp. When weights remain packed on Apple’s Metal backend, Transformers can load compatible ggml/Metal layer kernels and use ggml-org/ggml-attn for attention.

The setup has specific requirements: an Apple Silicon Mac, a supported PyTorch version, the latest Transformers code and a compatible kernels library. If the compatible quantization kernel is unavailable, the loader can dequantize the model instead, which uses more memory. If the ggml attention kernel cannot be fetched, the system falls back to standard sdpa attention with a warning; users can also select sdpa directly.

The same checkpoints can be served through transformers serve, which provides an OpenAI-compatible API on localhost. Clients such as Jan or Pi can connect through a custom provider. Hugging Face says its performance comparison uses llama.cpp as the reference and covers three GGUF checkpoints: a small dense model, a larger dense model and a mixture-of-experts model.

At a glance
announcementWhen: Available on the Transformers main bran…
The developmentHugging Face added a Transformers feature that loads GGUF quantized models directly through from_pretrained.
At a glance
announcementWhen: announced April 2026; available via tra…
The developmentHugging Face announced that the transformers library can now run llama.cpp-style GGUF quantized models natively, using ggml kernels for near-llama.cpp performance on Apple Silicon.

GGUF Models Reach Transformers

The change makes GGUF checkpoints available within a widely used model-development library, reducing the need for developers to switch to a separate llama.cpp-based application when working with those files. GGUF is used by local inference tools including Ollama, LM Studio and Jan, and the format packages weights and model metadata in a single file.

Quantization can reduce the memory needed to run a model locally. Hugging Face’s example for Unsloth’s Qwen3.5-4B lists a BF16 file at 8.42 GB and a Q4_K_M version at 2.74 GB. Those sizes illustrate the storage difference; they do not establish a particular quality or speed outcome. Hugging Face says the quality effect of lower precision depends on the model and task, so users should assess it against their own workloads.

For developers already using Transformers, loading Hub-hosted GGUF files through the existing API may simplify local experiments and serving workflows. The practical reach of the feature is currently limited by its Apple Silicon focus, architecture coverage and dependency on compatible kernels.

Amazon

Top picks for "transformer quantiz llama"

As an affiliate, we earn on qualifying purchases.

A New Route for GGUF

GGUF is a format developed by the llama.cpp project for storing model weights and related metadata, including tokenizer information and, optionally, a chat template. It supports multiple quantization levels, which trade memory use against numerical precision. A Q4_K_M checkpoint uses mostly four-bit weights while retaining higher precision for some tensors.

Before this addition, developers generally ran GGUF checkpoints through llama.cpp or tools built around it. Hugging Face’s announcement describes Transformers support as a way to use those checkpoints through its own APIs. The company’s cited example sizes for Qwen3.5-4B are 3.53 GB for Q6_K, 3.14 GB for Q5_K_M and 2.74 GB for Q4_K_M, compared with 8.42 GB for BF16. It recommends starting with Q4_K_M and considering higher-precision variants when memory allows, while advising evaluation on the intended task.

The announcement follows Hugging Face co-founder Julien Chaumond’s public demonstration of Qwen3.6 27B running in the Pi coding agent through llama.cpp on a MacBook Pro. Chaumond said it felt “very, very close” to Claude Opus for non-trivial tasks on Hugging Face codebases. That was his assessment of a demonstration, not an independent benchmark of the new Transformers integration.

“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”

— Hugging Face announcement

Support Still Has Limits

The announcement identifies Apple Silicon and Qwen3.5 as the initial targets. It does not give a timeline for support on CUDA, Linux or Windows, or specify when additional model architectures will be covered. The feature is on the Transformers main branch, and a stable-release date has not been announced.

Hugging Face refers to benchmarks against llama.cpp across three checkpoints, but the available source material does not provide the full figures or enough hardware and test detail to assess how results vary across setups. It is also unclear how quickly compatible kernels will be available for different PyTorch versions. Quantization’s effect on output quality remains dependent on the model and task, according to Hugging Face.

Stable Release and Wider Coverage

The next practical milestone is inclusion in a stable Transformers release, which would let users install the feature without using the main branch. Hugging Face has not announced when that release will arrive. Until then, users need the current main-branch code and compatible kernel builds to use the packed-weight path.

Further expansion could add hardware backends and model architectures, but the announcement provides no schedule or confirmed roadmap for those additions. Users can follow Hugging Face’s GGUF documentation and kernels library for updates on supported models, quantization types and dependencies.

Key Questions

What did Hugging Face add?

It added support for loading GGUF quantized checkpoints in Transformers using the from_pretrained API, with a GGUF file passed through the gguf_file argument.

Which devices and models are supported initially?

The announced initial support targets Apple Silicon Macs and the Qwen3.5 architecture. The source does not give a schedule for other hardware or architectures.

Does Transformers use llama.cpp kernels?

Hugging Face says the integration reuses llama.cpp’s ggml kernels for performance. If compatible quantization kernels are unavailable, the loader can dequantize the model, using more memory.

Is the feature in a stable Transformers release?

No. It is currently available on the Transformers main branch. Hugging Face has not announced a date for a stable release containing the feature.

How much can quantization reduce model file size?

In Hugging Face’s Qwen3.5-4B example, the listed size falls from 8.42 GB in BF16 to 2.74 GB in Q4_K_M. The resulting quality depends on the model and task, so users should evaluate their own workload.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Hardware Innovation: Designing The Foundation Before The Function

New hardware designs focus on inference, emphasizing thermal efficiency, memory interconnects, and specialization to meet growing AI demand.

Community volunteer action tracker for local boards

A new volunteer action tracker is being tested to improve follow-up on community board decisions, aiming to streamline volunteer coordination and accountability.

The runway.How enterprise-revenuelock becomes the load-bearing valuation argument.

OpenAI and Anthropic are preparing record-breaking IPOs, emphasizing enterprise revenue as the key to justified valuations amid uncertain margins and profitability.

Behind The Scenes Of AI II: The Engine Room In Twelve Machines

An in-depth look at the core processes powering AI chatbots, revealing how twelve key machines in the engine room operate and why they matter.