🔍 Read the full analysis: Transformers And Llama.cpp Quants: What’s New? on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has added early support for loading GGUF-quantized checkpoints in its Transformers library through the familiar from_pretrained API. The feature is on the main branch, with initial support focused on Apple Silicon and Qwen3.5; broader hardware and architecture support has no announced timeline.
Hugging Face has added support for GGUF checkpoints to its Transformers library, letting users load supported quantized models through the from_pretrained API and run them locally. The early rollout targets Apple Silicon Macs and Qwen3.5, bringing a widely used llama.cpp model format into the PyTorch-based library while the feature remains on the main branch ahead of a stable release.
Users can select a GGUF checkpoint hosted on the Hugging Face Hub and pass its file through the gguf_file argument to from_pretrained. Hugging Face says the model can then generate text without extra configuration. The implementation reuses ggml kernels associated with llama.cpp, with the stated goal of keeping inference performance close to llama.cpp. The announcement describes performance comparisons across three checkpoints: a small dense model, a larger dense model and a mixture-of-experts model. The supplied material does not include the full benchmark figures or enough hardware details to generalize their results.
On compatible Apple Silicon systems, Transformers can keep weights packed on Metal and load compatible ggml/Metal layer kernels. The implementation uses ggml-org/ggml-attn for attention when available. If that kernel cannot be fetched, it falls back to PyTorch’s standard scaled dot product attention, or SDPA, with a warning. Users can also select SDPA directly. The setup requires an Apple Silicon Mac, a supported PyTorch version and compatible versions of Transformers and the quantization kernels library. Without a compatible quantization kernel, the loader dequantizes the model, which uses more memory.
The same checkpoints can also be served through transformers serve, which exposes an OpenAI-compatible API on localhost. Clients such as Jan or Pi can connect through a custom provider configured for that endpoint. Hugging Face’s example sizes for Unsloth’s Qwen3.5-4B show the memory tradeoff: the BF16 file is 8.42 GB, compared with 3.53 GB for Q6_K, 3.14 GB for Q5_K_M and 2.74 GB for Q4_K_M. Those are file sizes, not a guarantee of total runtime memory use or model quality.
GGUF Enters the Transformers Workflow
The change gives developers who already use Transformers a way to access compatible GGUF checkpoints without switching to a separate llama.cpp-based application for loading. GGUF bundles weights and model metadata, and its quantization options can reduce checkpoint size by storing many weights at lower precision. That can make some models easier to run on machines with limited memory, although actual use also depends on the model, software setup and available memory.
For users, the practical effect is a familiar loading interface and an option to serve models to local clients through an API. This may simplify workflows that combine Hub-hosted models with tools built around Transformers. It does not establish that every GGUF checkpoint will work, that output quality will match a full-precision model, or that performance will match llama.cpp on every machine. Hugging Face advises users to evaluate quantized models on their own tasks because the effect on quality varies by model and workload.
Top picks for "transformer llama quant"
As an affiliate, we earn on qualifying purchases.
Quantized Models Meet PyTorch
GGUF was developed for the llama.cpp ecosystem and is used by local inference tools including Ollama, LM Studio and Jan. Before this addition, users generally ran GGUF checkpoints through llama.cpp-derived software rather than loading them within the Transformers workflow. The new support links those ecosystems for the architectures and hardware covered by this rollout; it does not replace the tools or extend support to every model family.
GGUF files can include tokenizer information and an optional chat template alongside model weights. Quantization variants such as Q4_K_M use mostly four-bit weights while retaining higher precision for some tensors. Hugging Face recommends starting with Q4_K_M and considering Q5_K_M or Q6_K when more memory is available, while stressing that users should check results on their own workloads. The announcement follows a period of public interest in local model use, including a demonstration by Hugging Face co-founder Julien Chaumond of Qwen3.6 27B running in the Pi coding agent through llama.cpp on a MacBook Pro. Chaumond described its performance on some coding tasks as feeling “very, very close” to Claude Opus; that was his assessment, not a benchmark result reported here.
“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”
— Hugging Face announcement
Support Limits Remain
The current support described by Hugging Face is limited to Apple Silicon and Qwen3.5. The announcement gives no schedule for CUDA, Linux or Windows support, and it does not say when additional model architectures will be added. The feature is on the Transformers main branch, and no stable release date has been announced.
Performance will depend on the specific hardware, model and kernels available. Although Hugging Face says it compared Transformers with llama.cpp on three checkpoint types, the source material here does not provide the full figures or test conditions. It is also unclear how quickly support will expand to other model families or how much memory particular setups will require beyond the checkpoint file sizes. Quantization’s effect on output quality remains model- and task-dependent.
Stable Release and Wider Support
The next stated milestone is inclusion in a stable Transformers release, which would let users install the feature without using the main branch. Hugging Face has not given a release date. Users considering the current version need to check the latest Transformers instructions and compatibility information for PyTorch and the quantization kernels library before loading a checkpoint.
Broader device and architecture coverage may follow, but the announcement does not set dates or promise specific additions. Users can watch the Transformers and GGUF documentation and the kernels library for support updates. Until the scope changes, people using other hardware or architectures should check compatibility rather than assume the new loader will work for their setup.
Key Questions
What changed in Transformers?
Users can load supported GGUF checkpoints through from_pretrained by supplying a gguf_file argument. The feature is currently on the main branch.
Which systems and models are supported so far?
The initial rollout described by Hugging Face focuses on Apple Silicon Macs and the Qwen3.5 architecture. The announcement gives no timeline for other platforms or architectures.
Does GGUF quantization reduce memory use?
It can reduce checkpoint size. For Unsloth’s Qwen3.5-4B, the cited file sizes are 8.42 GB for BF16 and 2.74 GB for Q4_K_M. Runtime memory use and quality depend on the setup, model and task.
When will the feature reach a stable release?
Hugging Face has not announced a date. The feature is available on the Transformers main branch ahead of a future stable release.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
