Skip to content

Hugging Face No Space Left on Device (Errno 28) Fix

Tested with: huggingface_hub 1.33.0 (hf-xet 1.6.0), transformers 5.17.0, plus huggingface_hub 0.36.2 with transformers 4.57.6 for the legacy checks; Python 3.12.14 in the python:3.12-slim Docker image on Ubuntu 24.04. Last run 2026-09-27.

TL;DR:
  1. The disk holding the Hugging Face cache (~/.cache/huggingface by default) is full. Point the cache at a bigger disk: export HF_HOME=/data/huggingface.
  2. Free space with hf cache ls, hf cache rm model/<repo_id> and hf cache prune. huggingface-cli no longer works in huggingface_hub 1.x.
  3. Do not trust rm -rf models--* on huggingface_hub 1.33.0: Xet downloads keep the weights in a shared hub/blobs/ store, and in my run deleting the model folder freed 0 bytes until hf cache prune ran.
  4. TRANSFORMERS_CACHE is silently ignored by transformers 5.17.0. Use HF_HOME or HF_HUB_CACHE.

You call from_pretrained() or hf_hub_download() and the download dies partway. With huggingface_hub 1.33.0 the error you see depends on which download backend handled the file. Most large model files on the Hub are now served through Xet, and then the error looks like this (full traceback trimmed to the last frames):

UserWarning: Not enough free disk space to download the file. The expected file size is: 548.11 MB. The target location /hf/hub/models--openai-community--gpt2/blobs only has 41.93 MB free disk space.
...
  File "/usr/local/lib/python3.12/site-packages/huggingface_hub/file_download.py", line 589, in xet_get
    with session.new_file_download_group(
         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
RuntimeError: Task error: File reconstruction error: IO Error: No space left on device (os error 28)

With Xet disabled (HF_HUB_DISABLE_XET=1) or for files that go over plain HTTP, you get the classic Python error instead:

  File "/usr/local/lib/python3.12/site-packages/huggingface_hub/file_download.py", line 464, in http_get
    temp_file.write(chunk)
OSError: [Errno 28] No space left on device

Note the UserWarning printed before the download starts. huggingface_hub checks free space first and tells you the expected size and what is available, but it only warns and then downloads anyway. If you see that warning scroll past, the failure is already decided.

This error has nothing to do with your GPU or model code, which is what makes it annoying the first time: you go looking for a CUDA or driver problem before checking df -h.

Reproduced

Every row ran in a fresh python:3.12-slim container with the cache on a size-limited tmpfs (docker run --tmpfs /hf:size=40m -e HF_HOME=/hf ...), downloading openai-community/gpt2, whose model.safetensors is 548,105,171 bytes.

What I ranExact resultFix that worked
hf_hub_download("openai-community/gpt2", "model.safetensors") into a 40 MB cache, default settings (Xet)RuntimeError: Task error: File reconstruction error: IO Error: No space left on device (os error 28)HF_HOME on a 1 GB tmpfs: download succeeded
Same with HF_HUB_DISABLE_XET=1OSError: [Errno 28] No space left on device (from temp_file.write(chunk))Same
snapshot_download("openai-community/gpt2", allow_patterns=["*.json", "*.safetensors"]) into 40 MB11 of 12 files fetched, then the same RuntimeError on the safetensors fileBigger disk, plus tighter patterns (below)
Download killed with SIGKILL after 4 s, no XetLeft blobs/248dfc39...a707.e22da17a.incomplete, 20,971,520 bytes, plus a 0-byte .lockhf cache prune -y: "Deleted 1 incomplete download(s); freed 21.0M."
rm -rf hub/models--openai-community--gpt2 after a Xet downloaddf still showed 523M used, hf cache ls said "No results found."hf cache prune -y: "Deleted 1 unreferenced shared blob(s); freed 548.1M."
huggingface-cli delete-cache on 1.33.0Warning: `huggingface-cli` is deprecated and no longer works. Use `hf` instead.hf cache rm model/openai-community/gpt2 -y: "freed 551.0M"
hf cache scan on 1.33.0Error: No such command 'scan'.hf cache ls
TRANSFORMERS_CACHE=/big/tc with transformers 4.57.6FutureWarning: Using `TRANSFORMERS_CACHE` is deprecated and will be removed in v5 of Transformers. Use `HF_HOME` instead. (still honored)HF_HOME
TRANSFORMERS_CACHE=/big/tc with transformers 5.17.0No warning, /big/tc never created, files went to the default cacheHF_HOME or HF_HUB_CACHE

One thing I expected and did not see: a failed download that raises the error cleans up after itself. After both 40 MB failures the tmpfs was back to 72K used, with only a 0-byte .lock file left. Leftover .incomplete files come from downloads that were killed (Ctrl+C twice, OOM killer, a pod eviction), not from the out-of-space error itself.


Why the Default Cache Fills Up

Every from_pretrained(), hf_hub_download() or snapshot_download() call stores the files on disk so the next load does not download again. By default the cache lives at:

~/.cache/huggingface/hub/

On many cloud VMs and WSL2 setups the root or home partition is small, and model repos are big. A repo is often much bigger than the one file you need. openai-community/gpt2 has 26 files totalling 5,632,417,295 bytes (about 5.6 GB) because it ships the same weights as safetensors, PyTorch, TensorFlow, Flax, Rust and three ONNX exports. from_pretrained() only fetches what it needs, but a bare snapshot_download() takes everything.

Check usage first:

df -h ~/.cache/huggingface   # which partition, and how full
hf cache ls                  # what is cached, per repo

Real output of hf cache ls after one download:

ID                            SIZE LAST_ACCESSED  LAST_MODIFIED     REFS
--------------------------- ------ -------------- ----------------- ----
model/openai-community/gpt2 551.0M 45 seconds ago a few seconds ago main

Found 1 repo(s) for a total of 1 revision(s) and 551.0M on disk.

Prefer hf cache ls over du -sh ~/.cache/huggingface/hub/*/. With huggingface_hub 1.33.0, a Xet download put the weights in a shared store at hub/blobs/63/63bed808... and hard-linked them into the model folder. du counts a hard-linked file once, so in my run it reported 8.0K for models--openai-community--gpt2 and 523M for hub/blobs. The layout looked like this:

hub/
ā”œā”€ā”€ CACHEDIR.TAG
ā”œā”€ā”€ blobs/                       # shared Xet blob store (new)
│   └── 63/63bed80836ee...8758   # the 548 MB safetensors file
└── models--openai-community--gpt2/
    ā”œā”€ā”€ blobs/
    │   ā”œā”€ā”€ 10c66461e4c1...      # config.json
    │   └── 248dfc391186...      # model.safetensors (hard link to the shared blob)
    ā”œā”€ā”€ refs/main
    └── snapshots/607a30d783df.../
        ā”œā”€ā”€ config.json -> ../../blobs/10c66461e4c1...
        └── model.safetensors -> ../../blobs/248dfc391186...

With HF_HUB_DISABLE_XET=1 there was no top-level hub/blobs/ and the model folder held the full 523M, the older layout.


Fix 1: Move the Cache with HF_HOME

HF_HOME sets the root for everything the Hugging Face libraries store; the model cache goes to $HF_HOME/hub. Find a partition with room (df -h), then:

mkdir -p /data/huggingface
export HF_HOME=/data/huggingface

# confirm it is picked up
python -c "from huggingface_hub import constants as c; print(c.HF_HOME, c.HF_HUB_CACHE)"
# /data/huggingface /data/huggingface/hub

Add the export line to ~/.bashrc or ~/.zshrc to make it permanent. To keep models you already have, move them instead of re-downloading:

mkdir -p /data/huggingface/hub
mv ~/.cache/huggingface/hub/* /data/huggingface/hub/

If you only want to move the model cache and leave tokens and other files where they are, set HF_HUB_CACHE instead. With HF_HOME=/hf and HF_HUB_CACHE=/big/hub, the constants resolved to /hf /big/hub.

Do not use TRANSFORMERS_CACHE. transformers 4.57.6 still honors it with a FutureWarning that says it "will be removed in v5 of Transformers". In 5.17.0 it is gone: no warning, the directory is never created, and downloads land in the default cache, so the disk keeps filling up while you think you moved it.


Fix 2: Delete Old Models with hf cache rm

If you cannot move the cache (shared server, no second disk), delete what you no longer need. In huggingface_hub 1.x the CLI is hf. The old huggingface-cli entry point is still installed but only prints "huggingface-cli is deprecated and no longer works. Use hf instead." The hf cache subcommands in 1.33.0 are list (alias ls), rm, prune and verify; there is no scan.

pip install -U huggingface_hub
hf cache ls
hf cache rm model/openai-community/gpt2          # asks for confirmation
hf cache rm model/openai-community/gpt2 -y       # no prompt, for scripts
hf cache rm model/openai-community/gpt2 --dry-run

Real output of the -y run:

About to delete 1 repo(s) totalling 551.0M.
  - model/openai-community/gpt2 (entire repo)
Cache deletion done. Saved 551.0M.
āœ“ Deleted 1 repo(s) and 1 revision(s); freed 551.0M.

df confirmed it: the 2 GB tmpfs went back to 120K used. hf cache rm also accepts revision hashes and hf:// file URIs if you only want to drop one revision or file. Both rm and prune take --cache-dir when the cache is not in the default place.

On huggingface_hub 0.x the old commands still run: huggingface-cli scan-cache on 0.36.2 printed "'huggingface-cli scan-cache' is deprecated. Use 'hf cache scan' instead." and then worked, and huggingface-cli delete-cache still offered its --disable-tui and --sort options.


Fix 3: Clean Up Leftovers with hf cache prune

hf cache prune removes detached revisions, incomplete downloads and, on 1.33.0, shared blobs no model folder points to any more. Run it with --dry-run first:

hf cache prune --dry-run
hf cache prune -y

After a download killed partway (plain HTTP), the dry run found the partial file:

About to delete 1 incomplete download(s) (21.0M total).
āœ“ Dry run: no files were deleted.

and the real run removed it: "Deleted 1 incomplete download(s); freed 21.0M." A killed Xet download left a 0-byte .incomplete file instead, and prune removed that too.

The bigger reason to know prune is manual deletion. Deleting a model folder by hand used to be enough. On 1.33.0 with a Xet download, it was not:

rm -rf ~/.cache/huggingface/hub/models--openai-community--gpt2
df -h    # before and after in my run: 523M used both times

The 548 MB file survived in hub/blobs/ because that hard link was still there, and hf cache ls said "No results found." so nothing looked wrong. hf cache prune found it:

About to delete 1 unreferenced shared blob(s) (548.1M total).
Cache deletion done. Saved 0.0.
āœ“ Deleted 1 unreferenced shared blob(s); freed 548.1M.

(The "Saved 0.0." line is what 1.33.0 printed; the last line and df, back to 76K used, show the space really came back.) So: use hf cache rm to delete models, and if you already deleted folders by hand, run hf cache prune afterwards.


When you cannot set environment variables (a fixed entrypoint, a tool that launches its own processes) but do have a bigger disk, move the directory and leave a symlink:

mkdir -p /data
mv ~/.cache/huggingface /data/huggingface
ln -s /data/huggingface ~/.cache/huggingface

I tested this with no HF_HOME set: huggingface_hub still reported /root/.cache/huggingface/hub as its cache, and the downloaded config.json ended up under the symlink target on the other filesystem.

In Docker, mount the big disk and point HF_HOME at it, which is how every test in this post was run:

docker run --gpus all \
  -v /data/huggingface:/hf -e HF_HOME=/hf \
  my-ml-image python train.py

Fix 5: Per-Call Cache Override with cache_dir

For one script, pass cache_dir. It overrides HF_HOME for that call only:

from transformers import AutoConfig

AutoConfig.from_pretrained("openai-community/gpt2", cache_dir="/big/cd")
# /big/cd now contains: .locks, CACHEDIR.TAG, models--openai-community--gpt2

AutoModel*.from_pretrained, AutoTokenizer.from_pretrained, hf_hub_download and snapshot_download take the same argument. If you use it in several places, read the path from one config value or an environment variable so moving the cache is a one-line change.


Fix 6: Download Only the Files You Need

If you use snapshot_download(), filter the files. One trap: in allow_patterns, * also matches /. My first try with allow_patterns=["*.json", "*.safetensors"] on gpt2 also pulled the onnx/ subfolder's JSON files. Naming top-level files explicitly gave exactly what a safetensors load needs:

from huggingface_hub import snapshot_download
import os

path = snapshot_download(
    "openai-community/gpt2",
    allow_patterns=["*.safetensors", "config.json", "generation_config.json",
                    "tokenizer*", "vocab.json", "merges.txt"],
)
print(sorted(os.listdir(path)))
# ['config.json', 'generation_config.json', 'merges.txt', 'model.safetensors',
#  'tokenizer.json', 'tokenizer_config.json', 'vocab.json']

That came to 551.0M in hf cache ls, against about 5.6 GB for the whole repo. ignore_patterns works the other way round, for example ignore_patterns=["*.bin", "*.h5", "*.msgpack", "*.ot", "onnx/*"].


Quick-Reference: Which Fix to Use

Situation Best fix Effort
You have a second disk or data partition export HF_HOME=/data/huggingface in .bashrc One-time, 2 commands
Old models you no longer need hf cache rm model/<repo_id> Minutes
Deleted folders by hand but space did not come back, or a killed download hf cache prune One command
Cannot change env vars Symlink ~/.cache/huggingface to a larger partition Three commands
One script, no system changes from_pretrained(..., cache_dir="/data/hf_cache") One line of code
snapshot_download pulling every weight format allow_patterns / ignore_patterns One argument

Which one should you actually do

If you have any access to a bigger disk, set HF_HOME once and stop thinking about it. It fixes the problem at the source instead of every few weeks when the cache fills again. For cleanup, stick to hf cache rm and hf cache prune rather than rm -rf: on current huggingface_hub the files on disk are no longer laid out one folder per model, and prune is what knows where the space went.