AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Transformers And Llama.cpp Quants: What’s New? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face has added early support for loading GGUF-quantized checkpoints in its Transformers library through the familiar from_pretrained API. The feature is on the main branch, with initial support focused on Apple Silicon and Qwen3.5; broader hardware and architecture support has no announced timeline.

Hugging Face has added support for GGUF checkpoints to its Transformers library, letting users load supported quantized models through the from_pretrained API and run them locally. The early rollout targets Apple Silicon Macs and Qwen3.5, bringing a widely used llama.cpp model format into the PyTorch-based library while the feature remains on the main branch ahead of a stable release.

Users can select a GGUF checkpoint hosted on the Hugging Face Hub and pass its file through the gguf_file argument to from_pretrained. Hugging Face says the model can then generate text without extra configuration. The implementation reuses ggml kernels associated with llama.cpp, with the stated goal of keeping inference performance close to llama.cpp. The announcement describes performance comparisons across three checkpoints: a small dense model, a larger dense model and a mixture-of-experts model. The supplied material does not include the full benchmark figures or enough hardware details to generalize their results.

On compatible Apple Silicon systems, Transformers can keep weights packed on Metal and load compatible ggml/Metal layer kernels. The implementation uses ggml-org/ggml-attn for attention when available. If that kernel cannot be fetched, it falls back to PyTorch’s standard scaled dot product attention, or SDPA, with a warning. Users can also select SDPA directly. The setup requires an Apple Silicon Mac, a supported PyTorch version and compatible versions of Transformers and the quantization kernels library. Without a compatible quantization kernel, the loader dequantizes the model, which uses more memory.

The same checkpoints can also be served through transformers serve, which exposes an OpenAI-compatible API on localhost. Clients such as Jan or Pi can connect through a custom provider configured for that endpoint. Hugging Face’s example sizes for Unsloth’s Qwen3.5-4B show the memory tradeoff: the BF16 file is 8.42 GB, compared with 3.53 GB for Q6_K, 3.14 GB for Q5_K_M and 2.74 GB for Q4_K_M. Those are file sizes, not a guarantee of total runtime memory use or model quality.

At a glance
updateWhen: Available on the Transformers main bran…
The developmentHugging Face added GGUF model loading to Transformers, allowing users to run supported quantized checkpoints through from_pretrained.
At a glance
announcementWhen: announced April 2026; available via tra…
The developmentHugging Face announced that the transformers library can now run llama.cpp-style GGUF quantized models natively, using ggml kernels for near-llama.cpp performance on Apple Silicon.

GGUF Enters the Transformers Workflow

The change gives developers who already use Transformers a way to access compatible GGUF checkpoints without switching to a separate llama.cpp-based application for loading. GGUF bundles weights and model metadata, and its quantization options can reduce checkpoint size by storing many weights at lower precision. That can make some models easier to run on machines with limited memory, although actual use also depends on the model, software setup and available memory.

For users, the practical effect is a familiar loading interface and an option to serve models to local clients through an API. This may simplify workflows that combine Hub-hosted models with tools built around Transformers. It does not establish that every GGUF checkpoint will work, that output quality will match a full-precision model, or that performance will match llama.cpp on every machine. Hugging Face advises users to evaluate quantized models on their own tasks because the effect on quality varies by model and workload.

Amazon

Top picks for "transformer llama quant"

As an affiliate, we earn on qualifying purchases.

Quantized Models Meet PyTorch

GGUF was developed for the llama.cpp ecosystem and is used by local inference tools including Ollama, LM Studio and Jan. Before this addition, users generally ran GGUF checkpoints through llama.cpp-derived software rather than loading them within the Transformers workflow. The new support links those ecosystems for the architectures and hardware covered by this rollout; it does not replace the tools or extend support to every model family.

GGUF files can include tokenizer information and an optional chat template alongside model weights. Quantization variants such as Q4_K_M use mostly four-bit weights while retaining higher precision for some tensors. Hugging Face recommends starting with Q4_K_M and considering Q5_K_M or Q6_K when more memory is available, while stressing that users should check results on their own workloads. The announcement follows a period of public interest in local model use, including a demonstration by Hugging Face co-founder Julien Chaumond of Qwen3.6 27B running in the Pi coding agent through llama.cpp on a MacBook Pro. Chaumond described its performance on some coding tasks as feeling “very, very close” to Claude Opus; that was his assessment, not a benchmark result reported here.

“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”

— Hugging Face announcement

Support Limits Remain

The current support described by Hugging Face is limited to Apple Silicon and Qwen3.5. The announcement gives no schedule for CUDA, Linux or Windows support, and it does not say when additional model architectures will be added. The feature is on the Transformers main branch, and no stable release date has been announced.

Performance will depend on the specific hardware, model and kernels available. Although Hugging Face says it compared Transformers with llama.cpp on three checkpoint types, the source material here does not provide the full figures or test conditions. It is also unclear how quickly support will expand to other model families or how much memory particular setups will require beyond the checkpoint file sizes. Quantization’s effect on output quality remains model- and task-dependent.

Stable Release and Wider Support

The next stated milestone is inclusion in a stable Transformers release, which would let users install the feature without using the main branch. Hugging Face has not given a release date. Users considering the current version need to check the latest Transformers instructions and compatibility information for PyTorch and the quantization kernels library before loading a checkpoint.

Broader device and architecture coverage may follow, but the announcement does not set dates or promise specific additions. Users can watch the Transformers and GGUF documentation and the kernels library for support updates. Until the scope changes, people using other hardware or architectures should check compatibility rather than assume the new loader will work for their setup.

Key Questions

What changed in Transformers?

Users can load supported GGUF checkpoints through from_pretrained by supplying a gguf_file argument. The feature is currently on the main branch.

Which systems and models are supported so far?

The initial rollout described by Hugging Face focuses on Apple Silicon Macs and the Qwen3.5 architecture. The announcement gives no timeline for other platforms or architectures.

Does GGUF quantization reduce memory use?

It can reduce checkpoint size. For Unsloth’s Qwen3.5-4B, the cited file sizes are 8.42 GB for BF16 and 2.74 GB for Q4_K_M. Runtime memory use and quality depend on the setup, model and task.

When will the feature reach a stable release?

Hugging Face has not announced a date. The feature is available on the Transformers main branch ahead of a future stable release.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Musk

Search interest in Elon Musk has surged, driven by rising media coverage and public attention, though specific developments remain unconfirmed.

The Fundamental Reason AI Labs Are Committing To Recursive Self-Enhancement

AI research labs are increasingly pursuing recursive self-improvement, aiming for autonomous model upgrades. This shift could reshape AI development and innovation.

Choose The Best Mesh WiFi System For 2026

Discover the best mesh WiFi systems for 2026, including WiFi 7 and WiFi 6 options, to enhance coverage, speed, and future-proofing for your home network.

From Losses To Profit: How SenseTime’s AI Growth Led To Its First IFRS Net Profit

SenseTime reports its first IFRS net profit alongside 23.4% revenue increase and higher gross margins, signaling a shift toward profitability.