Fix Hugging Face CUDA Out of Memory at Model Load
- Load in 16-bit:
dtype=torch.float16(on transformers 4.x the argument istorch_dtype). Qwen2.5-3B went from an OOM in float32 to 5,898 MiB peak on an 8 GB card, with the same greedy output. - Check your transformers version. 4.57 loads in float32 unless you pass a dtype; 5.x loads in the checkpoint's dtype (bfloat16 here), so the same call that crashed on 4.57 fit on 5.17.
- If 16-bit still does not fit, use
device_map="auto"withmax_memory. It loads anything that fits in GPU plus RAM, but generation fell from 17.9 to about 1 token/s. - If it has to live on the GPU, quantize with bitsandbytes. It works on this Pascal card: 3,323 MiB in 8-bit, 2,010 MiB in 4-bit, at 4.2 and 3.8 tokens/s.
You call from_pretrained() and the process dies before a single token is generated. This is the real output from loading Qwen2.5-3B in float32 onto the GTX 1070 (transformers 5.17, last lines of the traceback):
File ".../transformers/core_model_loading.py", line 1240, in _materialize_copy
tensor = tensor.to(device=device, dtype=dtype)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 86.00 MiB. GPU 0 has
a total capacity of 7.91 GiB of which 24.44 MiB is free. Including non-PyTorch memory,
this process has 7.73 GiB memory in use. Of the allocated memory 7.64 GiB is allocated
by PyTorch, and 6.43 MiB is reserved by PyTorch but unallocated. If reserved but
unallocated memory is large try setting PYTORCH_ALLOC_CONF=expandable_segments:True
to avoid fragmentation. See documentation for Memory Management
(https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
The "Tried to allocate 86.00 MiB" is misleading on its own. The failing allocation is one weight matrix; what matters is that 7.64 GiB was already allocated by PyTorch and the reserved-but-unallocated part was only 6.43 MiB. That rules out fragmentation, so expandable_segments will not help. The weights simply do not fit. On transformers 4.57.6 the same load fails in _load_state_dict_into_meta_model with the same message, and the older pattern of loading to CPU first and then calling model.to("cuda") fails inside torch/nn/modules/module.py in convert.
This is different from the training OOM covered in the torch.cuda.OutOfMemoryError training guide. At load time there are no gradients, optimizer states or activations yet. The weights alone are too big for the card.
Reproduced on Real Hardware
Test: load Qwen/Qwen2.5-3B (3,085,938,688 parameters, ungated, stored as bfloat16 in 6.17 GB of safetensors), then greedy-generate 32 tokens from "The capital of France is". Each row ran in a fresh process on an otherwise idle GPU (160 MiB used by the desktop). "Peak allocated" is torch.cuda.max_memory_allocated(); "nvidia-smi" is the highest reading polled every 0.2 s, which includes the CUDA context. I first tried PyTorch 2.14, but its generate() tried to compile a Triton kernel and failed because the machine has no Python headers (Python.h), so the numbers below use 2.10.
| What I ran | Result | Peak allocated | nvidia-smi | Generation |
|---|---|---|---|---|
dtype=torch.float32, device_map="cuda" | OOM (error above) | 7,833 MiB | 7,704 MiB | none |
transformers 4.57.6, no dtype, device_map="cuda" | OOM, loads float32 by default | 7,829 MiB | 8,076 MiB | none |
float32 on CPU, then model.to("cuda") | OOM in module.py convert | 7,779 MiB | 8,026 MiB | none |
| transformers 5.17, no dtype | Loads as bfloat16 | 5,898 MiB | 6,230 MiB | 17.3 tok/s |
dtype=torch.float16 | Fits | 5,898 MiB | 6,230 MiB | 17.9 tok/s |
dtype=torch.bfloat16 (Pascal) | Fits, same text | 5,898 MiB | 6,230 MiB | 16.8 tok/s |
float32, device_map="auto" | 21 modules on GPU, 19 on CPU | 6,873 MiB | 7,152 MiB | 1.0 tok/s |
float32, device_map="auto", max_memory={0: "6GiB", "cpu": "20GiB"} | 17 on GPU, 23 on CPU | 5,697 MiB | 5,978 MiB | 0.8 tok/s |
float16, device_map="auto", max_memory={0: "2GiB", ...} | 10 on GPU, 30 on CPU | 1,825 MiB | 2,102 MiB | 1.1 tok/s |
| bitsandbytes 8-bit | Fits | 3,323 MiB | 3,672 MiB | 4.2 tok/s |
| bitsandbytes 4-bit NF4, double quant | Fits | 2,010 MiB | 2,298 MiB | 3.8 tok/s |
What the numbers show:
- Every configuration that loaded produced the same first sentence, "The capital of France is Paris." The 16-bit and offloaded runs printed identical text (first 120 characters compared); 8-bit and 4-bit diverged in the second sentence.
- Offloading has a price. Moving layers to CPU cut generation speed by 16 to 22 times on this machine, because the offloaded weights cross PCIe on every forward pass.
- On this Pascal card, bitsandbytes quantization saved memory but was slower than plain float16 (4.2 and 3.8 tok/s against 17.9).
Why Loading a Model Runs Out of Memory
The weights need roughly num_parameters × bytes_per_parameter: 4 bytes for float32, 2 for float16/bfloat16, about 1 for int8 and about 0.5 for int4. For Qwen2.5-3B, model.get_memory_footprint() reported 11,772 MiB in float32 and 5,886 MiB in float16. The float32 figure is more than the 7.91 GiB (about 8,100 MiB) the card has, so the load fails partway through, once about 7.6 GiB of weights are on the GPU.
A generation call adds little on top of the weights for short prompts: 5,886 MiB after loading, 5,898 MiB peak after 32 tokens. Long prompts and large batches grow the KV cache, so leave headroom if you generate thousands of tokens.
You can get an estimate before downloading anything by summing the checkpoint file sizes on the Hub. The checkpoint stores weights in its own dtype (Qwen2.5 ships bfloat16), so the file size is roughly the 16-bit load size:
from huggingface_hub import HfApi
def checkpoint_gb(repo_id):
info = HfApi().model_info(repo_id, files_metadata=True)
return sum(f.size for f in info.siblings
if f.rfilename.endswith(".safetensors")) / 1e9
for repo in ["Qwen/Qwen2.5-3B", "Qwen/Qwen2.5-1.5B"]:
print(repo, round(checkpoint_gb(repo), 2), "GB")
# Qwen/Qwen2.5-3B 6.17 GB
# Qwen/Qwen2.5-1.5B 3.09 GB
Double it for float32, halve it for 8-bit. Avoid estimating the parameter count from hidden_size and num_hidden_layers with the textbook 4·h² + 2·h·ffn per layer: modern models use grouped-query attention and three MLP matrices, and for Qwen2.5-3B that formula gives 2.54B parameters against the real 3.09B. AutoConfig also has no num_parameters attribute, so code that reads config.num_parameters silently falls back to its guess. Once a model is loaded, model.num_parameters() is exact.
Start with dtype=torch.float16
The most effective change is loading in half precision. It halves the weight memory and, in this test, did not change the output:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen2.5-3B"
# Crashes on an 8 GB card: 11,772 MiB of float32 weights
# model = AutoModelForCausalLM.from_pretrained(model_name, dtype=torch.float32, device_map="cuda")
model = AutoModelForCausalLM.from_pretrained(
model_name,
dtype=torch.float16, # transformers 4.x: torch_dtype=torch.float16
device_map="cuda",
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
print(model.dtype) # torch.float16
print(round(model.get_memory_footprint() / 2**20)) # 5886
inputs = tokenizer("The capital of France is", return_tensors="pt").to("cuda")
with torch.no_grad():
out = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))
# The capital of France is Paris. The capital of Germany is Berlin. ...
Two version details matter here:
- The argument was renamed. On transformers 5.17,
torch_dtype=still works but prints`torch_dtype` is deprecated! Use `dtype` instead!. On 4.x, usetorch_dtype; the 4.57.6 run above used it and loaded in 5,898 MiB peak, the same asdtypeon 5.17. - The default changed. With no dtype argument, transformers 4.57.6 loaded Qwen2.5-3B in float32 and ran out of memory. Transformers 5.17 loaded it as
torch.bfloat16, the dtype stored in the checkpoint's config, and it fit. If an old tutorial "works" for a colleague and crashes for you, compare transformers versions first.
bfloat16 is often described as unsupported on pre-Ampere cards. On the GTX 1070 with PyTorch 2.10, torch.cuda.is_bf16_supported() returned True, bfloat16 loaded in the same 5,898 MiB and generated the same text at 16.8 tok/s against 17.9 for float16. It works, but Pascal has no native bfloat16 math, so on older cards float16 is the safer default and bfloat16 is the one to prefer on Ampere and newer.
Still not enough? device_map="auto" with accelerate
When the model does not fit even in 16-bit, device_map="auto" (requires pip install accelerate) fills the GPU first and places the remaining layers in CPU RAM. Set max_memory to leave headroom on the GPU for the KV cache and anything else on the card:
import torch
from collections import Counter
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-3B",
dtype=torch.float32,
device_map="auto",
max_memory={0: "6GiB", "cpu": "20GiB"},
)
print(Counter(str(d) for d in model.hf_device_map.values()))
# Counter({'cpu': 23, '0': 17})
Transformers prints Some parameters are on the meta device because they were offloaded to the cpu. when this happens. That line is expected. Generation still ran on the GPU (inputs go to "cuda"), but at 0.8 tok/s instead of 17.9, because each forward pass copies the CPU-resident layers over. With float16 and a 2 GiB cap, 30 of the 40 modules went to CPU and the GPU peak was 1,825 MiB, still at about 1 tok/s.
Offloading moves the problem to CPU RAM. Loading this model in float32 on the CPU peaked at 17.3 GiB of process memory (the checkpoint is bfloat16, so it is upcast while loading). The float32 device_map="auto" runs peaked at 11.4 to 12.6 GiB, and float16 with a 2 GiB GPU cap at 10.6 GiB. If RAM runs out too, accelerate can also offload to disk through offload_folder; I did not test that here, and it is slower again.
low_cpu_mem_usage=True made no difference on transformers 5.17: CPU loading peaked at 17.3 GiB with and without it.
Need everything on GPU: quantize with bitsandbytes
bitsandbytes stores linear-layer weights in 8 or 4 bits and dequantizes them during the matrix multiply. It worked on the GTX 1070 (compute capability 6.1) with bitsandbytes 0.50.2, with no extra setup:
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
# pip install bitsandbytes accelerate
model_name = "Qwen/Qwen2.5-3B"
model_8bit = AutoModelForCausalLM.from_pretrained(
model_name,
quantization_config=BitsAndBytesConfig(load_in_8bit=True),
device_map="cuda",
)
print(round(model_8bit.get_memory_footprint() / 2**20)) # 3240
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
bnb_config_4bit = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
bnb_4bit_use_double_quant=True,
)
model_4bit = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-3B",
quantization_config=bnb_config_4bit,
device_map="cuda",
)
print(round(model_4bit.get_memory_footprint() / 2**20)) # 1917
print(round(torch.cuda.max_memory_allocated() / 2**20)) # 1969 right after loading
4-bit came out at 1,917 MiB, not a quarter of 5,886. Only the linear layers are quantized; the embedding matrix (151,936 × 2,048, shared with the output layer in this model) stays 16-bit and alone is about 593 MiB. Small models with large vocabularies gain less from 4-bit than the bit width suggests.
On this card quantized generation was 4 to 5 times slower than float16, so quantize because you need the model to fit, not for speed. The load also printed four `torch.jit.script_method` is deprecated warnings from a dependency; they are harmless.
Sometimes the fix is just a smaller model
If the task allows it, a smaller model from the same family is the most reliable fix. Qwen2.5-1.5B is a 3.09 GB checkpoint, half of the 3B, and fits on an 8 GB card in 16-bit with room for long contexts. A model that loads in float16 on the GPU ran 16 to 22 times faster here than the same model offloaded to CPU, so a smaller model on the GPU often beats a bigger one that barely fits.
Quick Reference: Weight Memory by Size
The generic rows are arithmetic (parameters × bytes per parameter, in GiB), weights only. The Qwen2.5-3B row is measured with get_memory_footprint(). Leave 1 to 2 GiB of headroom on top for the CUDA context and KV cache.
| Parameters | float32 | float16 / bfloat16 | int8 | int4 |
|---|---|---|---|---|
| 1B | 3.7 GiB | 1.9 GiB | 0.9 GiB | 0.5 GiB |
| 3B | 11.2 GiB | 5.6 GiB | 2.8 GiB | 1.4 GiB |
| 7B | 26.1 GiB | 13.0 GiB | 6.5 GiB | 3.3 GiB |
| 13B | 48.4 GiB | 24.2 GiB | 12.1 GiB | 6.1 GiB |
| 70B | 260.8 GiB | 130.4 GiB | 65.2 GiB | 32.6 GiB |
| Qwen2.5-3B, measured (3.09B) | 11.5 GiB | 5.7 GiB | 3.2 GiB | 1.9 GiB |
Where I'd actually start
Try 16-bit first. If the 16-bit size from the Hub file listing is still bigger than your card minus a gigabyte or two, decide between speed and fit: device_map="auto" keeps full precision but ran at about 1 token/s here, while 8-bit or 4-bit keeps everything on the GPU at 4 tokens/s. If you find yourself stacking offload on top of quantization just to get a model to load, drop to a smaller model in the same family.