Skip to content

Fix Hugging Face CUDA Out of Memory at Model Load

Tested with: transformers 5.17.0 (and 4.57.6 for the old default), accelerate 1.15.0, bitsandbytes 0.50.2, PyTorch 2.10.0 (CUDA 12.6), Python 3.12, NVIDIA GeForce GTX 1070 8 GB, driver 580, 23 GB RAM. Model: Qwen/Qwen2.5-3B. Last run 2026-09-27.

TL;DR, in the order to try them:
  1. Load in 16-bit: dtype=torch.float16 (on transformers 4.x the argument is torch_dtype). Qwen2.5-3B went from an OOM in float32 to 5,898 MiB peak on an 8 GB card, with the same greedy output.
  2. Check your transformers version. 4.57 loads in float32 unless you pass a dtype; 5.x loads in the checkpoint's dtype (bfloat16 here), so the same call that crashed on 4.57 fit on 5.17.
  3. If 16-bit still does not fit, use device_map="auto" with max_memory. It loads anything that fits in GPU plus RAM, but generation fell from 17.9 to about 1 token/s.
  4. If it has to live on the GPU, quantize with bitsandbytes. It works on this Pascal card: 3,323 MiB in 8-bit, 2,010 MiB in 4-bit, at 4.2 and 3.8 tokens/s.

You call from_pretrained() and the process dies before a single token is generated. This is the real output from loading Qwen2.5-3B in float32 onto the GTX 1070 (transformers 5.17, last lines of the traceback):

  File ".../transformers/core_model_loading.py", line 1240, in _materialize_copy
    tensor = tensor.to(device=device, dtype=dtype)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 86.00 MiB. GPU 0 has
a total capacity of 7.91 GiB of which 24.44 MiB is free. Including non-PyTorch memory,
this process has 7.73 GiB memory in use. Of the allocated memory 7.64 GiB is allocated
by PyTorch, and 6.43 MiB is reserved by PyTorch but unallocated. If reserved but
unallocated memory is large try setting PYTORCH_ALLOC_CONF=expandable_segments:True
to avoid fragmentation.  See documentation for Memory Management
(https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)

The "Tried to allocate 86.00 MiB" is misleading on its own. The failing allocation is one weight matrix; what matters is that 7.64 GiB was already allocated by PyTorch and the reserved-but-unallocated part was only 6.43 MiB. That rules out fragmentation, so expandable_segments will not help. The weights simply do not fit. On transformers 4.57.6 the same load fails in _load_state_dict_into_meta_model with the same message, and the older pattern of loading to CPU first and then calling model.to("cuda") fails inside torch/nn/modules/module.py in convert.

This is different from the training OOM covered in the torch.cuda.OutOfMemoryError training guide. At load time there are no gradients, optimizer states or activations yet. The weights alone are too big for the card.


Reproduced on Real Hardware

Test: load Qwen/Qwen2.5-3B (3,085,938,688 parameters, ungated, stored as bfloat16 in 6.17 GB of safetensors), then greedy-generate 32 tokens from "The capital of France is". Each row ran in a fresh process on an otherwise idle GPU (160 MiB used by the desktop). "Peak allocated" is torch.cuda.max_memory_allocated(); "nvidia-smi" is the highest reading polled every 0.2 s, which includes the CUDA context. I first tried PyTorch 2.14, but its generate() tried to compile a Triton kernel and failed because the machine has no Python headers (Python.h), so the numbers below use 2.10.

What I ran Result Peak allocated nvidia-smi Generation
dtype=torch.float32, device_map="cuda"OOM (error above)7,833 MiB7,704 MiBnone
transformers 4.57.6, no dtype, device_map="cuda"OOM, loads float32 by default7,829 MiB8,076 MiBnone
float32 on CPU, then model.to("cuda")OOM in module.py convert7,779 MiB8,026 MiBnone
transformers 5.17, no dtypeLoads as bfloat165,898 MiB6,230 MiB17.3 tok/s
dtype=torch.float16Fits5,898 MiB6,230 MiB17.9 tok/s
dtype=torch.bfloat16 (Pascal)Fits, same text5,898 MiB6,230 MiB16.8 tok/s
float32, device_map="auto"21 modules on GPU, 19 on CPU6,873 MiB7,152 MiB1.0 tok/s
float32, device_map="auto", max_memory={0: "6GiB", "cpu": "20GiB"}17 on GPU, 23 on CPU5,697 MiB5,978 MiB0.8 tok/s
float16, device_map="auto", max_memory={0: "2GiB", ...}10 on GPU, 30 on CPU1,825 MiB2,102 MiB1.1 tok/s
bitsandbytes 8-bitFits3,323 MiB3,672 MiB4.2 tok/s
bitsandbytes 4-bit NF4, double quantFits2,010 MiB2,298 MiB3.8 tok/s

What the numbers show:

  • Every configuration that loaded produced the same first sentence, "The capital of France is Paris." The 16-bit and offloaded runs printed identical text (first 120 characters compared); 8-bit and 4-bit diverged in the second sentence.
  • Offloading has a price. Moving layers to CPU cut generation speed by 16 to 22 times on this machine, because the offloaded weights cross PCIe on every forward pass.
  • On this Pascal card, bitsandbytes quantization saved memory but was slower than plain float16 (4.2 and 3.8 tok/s against 17.9).

Why Loading a Model Runs Out of Memory

The weights need roughly num_parameters × bytes_per_parameter: 4 bytes for float32, 2 for float16/bfloat16, about 1 for int8 and about 0.5 for int4. For Qwen2.5-3B, model.get_memory_footprint() reported 11,772 MiB in float32 and 5,886 MiB in float16. The float32 figure is more than the 7.91 GiB (about 8,100 MiB) the card has, so the load fails partway through, once about 7.6 GiB of weights are on the GPU.

A generation call adds little on top of the weights for short prompts: 5,886 MiB after loading, 5,898 MiB peak after 32 tokens. Long prompts and large batches grow the KV cache, so leave headroom if you generate thousands of tokens.

You can get an estimate before downloading anything by summing the checkpoint file sizes on the Hub. The checkpoint stores weights in its own dtype (Qwen2.5 ships bfloat16), so the file size is roughly the 16-bit load size:

from huggingface_hub import HfApi

def checkpoint_gb(repo_id):
    info = HfApi().model_info(repo_id, files_metadata=True)
    return sum(f.size for f in info.siblings
               if f.rfilename.endswith(".safetensors")) / 1e9

for repo in ["Qwen/Qwen2.5-3B", "Qwen/Qwen2.5-1.5B"]:
    print(repo, round(checkpoint_gb(repo), 2), "GB")
# Qwen/Qwen2.5-3B 6.17 GB
# Qwen/Qwen2.5-1.5B 3.09 GB

Double it for float32, halve it for 8-bit. Avoid estimating the parameter count from hidden_size and num_hidden_layers with the textbook 4·h² + 2·h·ffn per layer: modern models use grouped-query attention and three MLP matrices, and for Qwen2.5-3B that formula gives 2.54B parameters against the real 3.09B. AutoConfig also has no num_parameters attribute, so code that reads config.num_parameters silently falls back to its guess. Once a model is loaded, model.num_parameters() is exact.


Start with dtype=torch.float16

The most effective change is loading in half precision. It halves the weight memory and, in this test, did not change the output:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen2.5-3B"

# Crashes on an 8 GB card: 11,772 MiB of float32 weights
# model = AutoModelForCausalLM.from_pretrained(model_name, dtype=torch.float32, device_map="cuda")

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    dtype=torch.float16,     # transformers 4.x: torch_dtype=torch.float16
    device_map="cuda",
)
tokenizer = AutoTokenizer.from_pretrained(model_name)

print(model.dtype)                                    # torch.float16
print(round(model.get_memory_footprint() / 2**20))    # 5886

inputs = tokenizer("The capital of France is", return_tensors="pt").to("cuda")
with torch.no_grad():
    out = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))
# The capital of France is Paris. The capital of Germany is Berlin. ...

Two version details matter here:

  • The argument was renamed. On transformers 5.17, torch_dtype= still works but prints `torch_dtype` is deprecated! Use `dtype` instead!. On 4.x, use torch_dtype; the 4.57.6 run above used it and loaded in 5,898 MiB peak, the same as dtype on 5.17.
  • The default changed. With no dtype argument, transformers 4.57.6 loaded Qwen2.5-3B in float32 and ran out of memory. Transformers 5.17 loaded it as torch.bfloat16, the dtype stored in the checkpoint's config, and it fit. If an old tutorial "works" for a colleague and crashes for you, compare transformers versions first.

bfloat16 is often described as unsupported on pre-Ampere cards. On the GTX 1070 with PyTorch 2.10, torch.cuda.is_bf16_supported() returned True, bfloat16 loaded in the same 5,898 MiB and generated the same text at 16.8 tok/s against 17.9 for float16. It works, but Pascal has no native bfloat16 math, so on older cards float16 is the safer default and bfloat16 is the one to prefer on Ampere and newer.


Still not enough? device_map="auto" with accelerate

When the model does not fit even in 16-bit, device_map="auto" (requires pip install accelerate) fills the GPU first and places the remaining layers in CPU RAM. Set max_memory to leave headroom on the GPU for the KV cache and anything else on the card:

import torch
from collections import Counter
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2.5-3B",
    dtype=torch.float32,
    device_map="auto",
    max_memory={0: "6GiB", "cpu": "20GiB"},
)
print(Counter(str(d) for d in model.hf_device_map.values()))
# Counter({'cpu': 23, '0': 17})

Transformers prints Some parameters are on the meta device because they were offloaded to the cpu. when this happens. That line is expected. Generation still ran on the GPU (inputs go to "cuda"), but at 0.8 tok/s instead of 17.9, because each forward pass copies the CPU-resident layers over. With float16 and a 2 GiB cap, 30 of the 40 modules went to CPU and the GPU peak was 1,825 MiB, still at about 1 tok/s.

Offloading moves the problem to CPU RAM. Loading this model in float32 on the CPU peaked at 17.3 GiB of process memory (the checkpoint is bfloat16, so it is upcast while loading). The float32 device_map="auto" runs peaked at 11.4 to 12.6 GiB, and float16 with a 2 GiB GPU cap at 10.6 GiB. If RAM runs out too, accelerate can also offload to disk through offload_folder; I did not test that here, and it is slower again.

low_cpu_mem_usage=True made no difference on transformers 5.17: CPU loading peaked at 17.3 GiB with and without it.


Need everything on GPU: quantize with bitsandbytes

bitsandbytes stores linear-layer weights in 8 or 4 bits and dequantizes them during the matrix multiply. It worked on the GTX 1070 (compute capability 6.1) with bitsandbytes 0.50.2, with no extra setup:

import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig

# pip install bitsandbytes accelerate
model_name = "Qwen/Qwen2.5-3B"

model_8bit = AutoModelForCausalLM.from_pretrained(
    model_name,
    quantization_config=BitsAndBytesConfig(load_in_8bit=True),
    device_map="cuda",
)
print(round(model_8bit.get_memory_footprint() / 2**20))   # 3240
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig

bnb_config_4bit = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_use_double_quant=True,
)
model_4bit = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2.5-3B",
    quantization_config=bnb_config_4bit,
    device_map="cuda",
)
print(round(model_4bit.get_memory_footprint() / 2**20))   # 1917
print(round(torch.cuda.max_memory_allocated() / 2**20))   # 1969 right after loading

4-bit came out at 1,917 MiB, not a quarter of 5,886. Only the linear layers are quantized; the embedding matrix (151,936 × 2,048, shared with the output layer in this model) stays 16-bit and alone is about 593 MiB. Small models with large vocabularies gain less from 4-bit than the bit width suggests.

On this card quantized generation was 4 to 5 times slower than float16, so quantize because you need the model to fit, not for speed. The load also printed four `torch.jit.script_method` is deprecated warnings from a dependency; they are harmless.


Sometimes the fix is just a smaller model

If the task allows it, a smaller model from the same family is the most reliable fix. Qwen2.5-1.5B is a 3.09 GB checkpoint, half of the 3B, and fits on an 8 GB card in 16-bit with room for long contexts. A model that loads in float16 on the GPU ran 16 to 22 times faster here than the same model offloaded to CPU, so a smaller model on the GPU often beats a bigger one that barely fits.


Quick Reference: Weight Memory by Size

The generic rows are arithmetic (parameters × bytes per parameter, in GiB), weights only. The Qwen2.5-3B row is measured with get_memory_footprint(). Leave 1 to 2 GiB of headroom on top for the CUDA context and KV cache.

Parameters float32 float16 / bfloat16 int8 int4
1B 3.7 GiB 1.9 GiB 0.9 GiB 0.5 GiB
3B 11.2 GiB 5.6 GiB 2.8 GiB 1.4 GiB
7B 26.1 GiB 13.0 GiB 6.5 GiB 3.3 GiB
13B 48.4 GiB 24.2 GiB 12.1 GiB 6.1 GiB
70B 260.8 GiB 130.4 GiB 65.2 GiB 32.6 GiB
Qwen2.5-3B, measured (3.09B) 11.5 GiB 5.7 GiB 3.2 GiB 1.9 GiB

Where I'd actually start

Try 16-bit first. If the 16-bit size from the Hub file listing is still bigger than your card minus a gigabyte or two, decide between speed and fit: device_map="auto" keeps full precision but ran at about 1 token/s here, while 8-bit or 4-bit keeps everything on the GPU at 4 tokens/s. If you find yourself stacking offload on top of quantization just to get a model to load, drop to a smaller model in the same family.