Gemma 4 GGUF Resources: Quants, VRAM, and How to Run

Looking for Gemma 4 GGUF or gemma4 gguf files? Official and community weights are free on Hugging Face, Kaggle, Ollama, and ModelScope. This page covers every variant — E2B, E4B, 26B MoE, and 31B Dense — in SafeTensors, quantized GGUF (Q4 / Q5 / Q8), GPTQ, and MLX, with sizes, VRAM notes, and download links.

All Gemma 4 models are released under the Apache 2.0 license, which means you can download, use, modify, and redistribute them freely for any purpose — including commercial applications.

Gemma 4 GGUF Download Sizes on Hugging Face

Google's official QAT model cards list deployable GGUF checkpoints for the Gemma 4 family. The size rows below are community quantizations and are labelled by publisher; inspect each linked repository before downloading.

ModelTotal ParamsQ4_K_MQ5_K_MQ8_0BF16Hugging Face Repo
Gemma 4 E2B-it5B3.11 GB3.36 GB5.05 GB9.31 GBunsloth/gemma-4-E2B-it-GGUF
Gemma 4 E4B-it8B4.98 GB5.48 GB8.19 GB15.1 GBunsloth/gemma-4-E4B-it-GGUF
Gemma 4 26B-A4B-it27B (MoE, 4B active)16.9 GB21.2 GB26.9 GB—unsloth/gemma-4-26B-A4B-it-GGUF
Gemma 4 31B-it33B (Dense)18.3 GB21.7 GB32.6 GB—unsloth/gemma-4-31B-it-GGUF

The rows above are community quantization listings; model formats and supported variants are cross-checked against Google's official Gemma 4 model cards. File sizes can change as repositories are revised; check the linked repository before downloading.

Model Format Guide

Understanding the different model file formats available for Gemma 4:

SafeTensors (.safetensors)

The default format on Hugging Face. Safe, fast-loading tensors designed to prevent code execution vulnerabilities. Used with Hugging Face Transformers, vLLM, and other Python-based frameworks.

Research, fine-tuning, Python frameworks, vLLM serving

GGUF (.gguf)

The standard format for llama.cpp and Ollama. Supports various quantization levels (Q4, Q5, Q8, etc.) to reduce model size and memory requirements. Optimized for CPU and mixed CPU/GPU inference.

Local inference, Ollama, llama.cpp, KoboldCpp, LM Studio

GPTQ

GPU-optimized quantization format that maintains high accuracy while significantly reducing VRAM requirements. Available through community contributions on Hugging Face.

GPU inference with reduced VRAM, production serving

MLX Format

Apple's native ML format optimized for Apple Silicon (M1/M2/M3/M4). Leverages unified memory architecture for efficient inference on Mac hardware.

Mac with Apple Silicon, MLX framework

Quantization Guide

Quantization reduces model size and memory usage at the cost of some accuracy. Here's how different levels compare for Gemma 4:

FormatBitsQualityNotes
BF16 / FP16 (Full Precision)16-bit100%Full model quality with no accuracy loss. Requires the most VRAM and disk space.
INT8 / Q88-bit~98-99%Minimal quality loss. Halves VRAM requirements compared to FP16. Recommended for most GPU deployments.
Q5_K_M5-bit~95-97%Good balance of quality and size. Popular choice for local inference with GGUF format.
INT4 / Q4_K_M4-bit~93-95%Significant size reduction with acceptable quality for most use cases. Enables running larger models on consumer hardware.

Download via Command Line

Hugging Face CLI

Install the Hugging Face CLI and download models directly:

pip install huggingface_hub

# Full-precision SafeTensors (official Google repo)
huggingface-cli download google/gemma-4-31B-it

# GGUF quantized (community, unsloth — most downloaded)
huggingface-cli download unsloth/gemma-4-31B-it-GGUF \
  --include "gemma-4-31B-it-Q4_K_M.gguf"

Git LFS

Clone model repositories with Git Large File Storage:

git lfs install
git clone https://huggingface.co/google/gemma-4-31B-it

Ollama CLI

Pull models directly into Ollama:

# Pull any variant
ollama pull gemma4:e2b
ollama pull gemma4:e4b
ollama pull gemma4:26b
ollama pull gemma4:31b

Download FAQ

Where is the best place to download Gemma 4?

Hugging Face is the most comprehensive source with all formats and variants. For one-command local setup, use Ollama. For users in China, ModelScope offers faster download speeds.

What format should I download?

For Ollama or llama.cpp: download GGUF files. For Python/vLLM: use SafeTensors format. For Mac with Apple Silicon: use MLX format. If unsure, start with Ollama which handles format selection automatically.

How large are Gemma 4 model files?

Full precision sizes: E2B (~4GB), E4B (~8GB), 26B MoE (~52GB), 31B Dense (~62GB). Q4 quantized versions are roughly 4x smaller. Ollama's default downloads use optimized quantization.

Do I need a Hugging Face account to download?

No. Gemma 4 models are publicly accessible under Apache 2.0 license. You can download without an account, though having one enables faster downloads and access to the Hugging Face CLI.

What is a GGUF file?

GGUF (GPT-Generated Unified Format) is a binary format designed for efficient local inference with llama.cpp and Ollama. It supports various quantization levels, allowing you to trade accuracy for smaller file sizes and lower memory usage.

Can I download Gemma 4 in China?

Yes. ModelScope (魔搭社区) mirrors Gemma 4 models with fast download speeds within China. Alternatively, use a mirror or proxy for Hugging Face downloads.

Download and Deploy

Get Gemma 4 model weights and start deploying. Check our deployment guide for step-by-step setup instructions.

Gemma 4 GGUF Downloads: E4B, Quants & Local Setup | Gemma 4