Hugging Face No Space Left on Device (Errno 28) Fix
- The disk holding the Hugging Face cache (
~/.cache/huggingfaceby default) is full. Point the cache at a bigger disk:export HF_HOME=/data/huggingface. - Free space with
hf cache ls,hf cache rm model/<repo_id>andhf cache prune.huggingface-clino longer works in huggingface_hub 1.x. - Do not trust
rm -rf models--*on huggingface_hub 1.33.0: Xet downloads keep the weights in a sharedhub/blobs/store, and in my run deleting the model folder freed 0 bytes untilhf cache pruneran. TRANSFORMERS_CACHEis silently ignored by transformers 5.17.0. UseHF_HOMEorHF_HUB_CACHE.
You call from_pretrained() or hf_hub_download() and the download dies partway. With huggingface_hub 1.33.0 the error you see depends on which download backend handled the file. Most large model files on the Hub are now served through Xet, and then the error looks like this (full traceback trimmed to the last frames):
UserWarning: Not enough free disk space to download the file. The expected file size is: 548.11 MB. The target location /hf/hub/models--openai-community--gpt2/blobs only has 41.93 MB free disk space.
...
File "/usr/local/lib/python3.12/site-packages/huggingface_hub/file_download.py", line 589, in xet_get
with session.new_file_download_group(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
RuntimeError: Task error: File reconstruction error: IO Error: No space left on device (os error 28)
With Xet disabled (HF_HUB_DISABLE_XET=1) or for files that go over plain HTTP, you get the classic Python error instead:
File "/usr/local/lib/python3.12/site-packages/huggingface_hub/file_download.py", line 464, in http_get
temp_file.write(chunk)
OSError: [Errno 28] No space left on device
Note the UserWarning printed before the download starts. huggingface_hub checks free space first and tells you the expected size and what is available, but it only warns and then downloads anyway. If you see that warning scroll past, the failure is already decided.
This error has nothing to do with your GPU or model code, which is what makes it annoying the first time: you go looking for a CUDA or driver problem before checking df -h.
Reproduced
Every row ran in a fresh python:3.12-slim container with the cache on a size-limited tmpfs (docker run --tmpfs /hf:size=40m -e HF_HOME=/hf ...), downloading openai-community/gpt2, whose model.safetensors is 548,105,171 bytes.
| What I ran | Exact result | Fix that worked |
|---|---|---|
hf_hub_download("openai-community/gpt2", "model.safetensors") into a 40 MB cache, default settings (Xet) | RuntimeError: Task error: File reconstruction error: IO Error: No space left on device (os error 28) | HF_HOME on a 1 GB tmpfs: download succeeded |
Same with HF_HUB_DISABLE_XET=1 | OSError: [Errno 28] No space left on device (from temp_file.write(chunk)) | Same |
snapshot_download("openai-community/gpt2", allow_patterns=["*.json", "*.safetensors"]) into 40 MB | 11 of 12 files fetched, then the same RuntimeError on the safetensors file | Bigger disk, plus tighter patterns (below) |
Download killed with SIGKILL after 4 s, no Xet | Left blobs/248dfc39...a707.e22da17a.incomplete, 20,971,520 bytes, plus a 0-byte .lock | hf cache prune -y: "Deleted 1 incomplete download(s); freed 21.0M." |
rm -rf hub/models--openai-community--gpt2 after a Xet download | df still showed 523M used, hf cache ls said "No results found." | hf cache prune -y: "Deleted 1 unreferenced shared blob(s); freed 548.1M." |
huggingface-cli delete-cache on 1.33.0 | Warning: `huggingface-cli` is deprecated and no longer works. Use `hf` instead. | hf cache rm model/openai-community/gpt2 -y: "freed 551.0M" |
hf cache scan on 1.33.0 | Error: No such command 'scan'. | hf cache ls |
TRANSFORMERS_CACHE=/big/tc with transformers 4.57.6 | FutureWarning: Using `TRANSFORMERS_CACHE` is deprecated and will be removed in v5 of Transformers. Use `HF_HOME` instead. (still honored) | HF_HOME |
TRANSFORMERS_CACHE=/big/tc with transformers 5.17.0 | No warning, /big/tc never created, files went to the default cache | HF_HOME or HF_HUB_CACHE |
One thing I expected and did not see: a failed download that raises the error cleans up after itself. After both 40 MB failures the tmpfs was back to 72K used, with only a 0-byte .lock file left. Leftover .incomplete files come from downloads that were killed (Ctrl+C twice, OOM killer, a pod eviction), not from the out-of-space error itself.
Why the Default Cache Fills Up
Every from_pretrained(), hf_hub_download() or snapshot_download() call stores the files on disk so the next load does not download again. By default the cache lives at:
~/.cache/huggingface/hub/
On many cloud VMs and WSL2 setups the root or home partition is small, and model repos are big. A repo is often much bigger than the one file you need. openai-community/gpt2 has 26 files totalling 5,632,417,295 bytes (about 5.6 GB) because it ships the same weights as safetensors, PyTorch, TensorFlow, Flax, Rust and three ONNX exports. from_pretrained() only fetches what it needs, but a bare snapshot_download() takes everything.
Check usage first:
df -h ~/.cache/huggingface # which partition, and how full
hf cache ls # what is cached, per repo
Real output of hf cache ls after one download:
ID SIZE LAST_ACCESSED LAST_MODIFIED REFS
--------------------------- ------ -------------- ----------------- ----
model/openai-community/gpt2 551.0M 45 seconds ago a few seconds ago main
Found 1 repo(s) for a total of 1 revision(s) and 551.0M on disk.
Prefer hf cache ls over du -sh ~/.cache/huggingface/hub/*/. With huggingface_hub 1.33.0, a Xet download put the weights in a shared store at hub/blobs/63/63bed808... and hard-linked them into the model folder. du counts a hard-linked file once, so in my run it reported 8.0K for models--openai-community--gpt2 and 523M for hub/blobs. The layout looked like this:
hub/
āāā CACHEDIR.TAG
āāā blobs/ # shared Xet blob store (new)
ā āāā 63/63bed80836ee...8758 # the 548 MB safetensors file
āāā models--openai-community--gpt2/
āāā blobs/
ā āāā 10c66461e4c1... # config.json
ā āāā 248dfc391186... # model.safetensors (hard link to the shared blob)
āāā refs/main
āāā snapshots/607a30d783df.../
āāā config.json -> ../../blobs/10c66461e4c1...
āāā model.safetensors -> ../../blobs/248dfc391186...
With HF_HUB_DISABLE_XET=1 there was no top-level hub/blobs/ and the model folder held the full 523M, the older layout.
Fix 1: Move the Cache with HF_HOME
HF_HOME sets the root for everything the Hugging Face libraries store; the model cache goes to $HF_HOME/hub. Find a partition with room (df -h), then:
mkdir -p /data/huggingface
export HF_HOME=/data/huggingface
# confirm it is picked up
python -c "from huggingface_hub import constants as c; print(c.HF_HOME, c.HF_HUB_CACHE)"
# /data/huggingface /data/huggingface/hub
Add the export line to ~/.bashrc or ~/.zshrc to make it permanent. To keep models you already have, move them instead of re-downloading:
mkdir -p /data/huggingface/hub
mv ~/.cache/huggingface/hub/* /data/huggingface/hub/
If you only want to move the model cache and leave tokens and other files where they are, set HF_HUB_CACHE instead. With HF_HOME=/hf and HF_HUB_CACHE=/big/hub, the constants resolved to /hf /big/hub.
Do not use TRANSFORMERS_CACHE. transformers 4.57.6 still honors it with a FutureWarning that says it "will be removed in v5 of Transformers". In 5.17.0 it is gone: no warning, the directory is never created, and downloads land in the default cache, so the disk keeps filling up while you think you moved it.
Fix 2: Delete Old Models with hf cache rm
If you cannot move the cache (shared server, no second disk), delete what you no longer need. In huggingface_hub 1.x the CLI is hf. The old huggingface-cli entry point is still installed but only prints "huggingface-cli is deprecated and no longer works. Use hf instead." The hf cache subcommands in 1.33.0 are list (alias ls), rm, prune and verify; there is no scan.
pip install -U huggingface_hub
hf cache ls
hf cache rm model/openai-community/gpt2 # asks for confirmation
hf cache rm model/openai-community/gpt2 -y # no prompt, for scripts
hf cache rm model/openai-community/gpt2 --dry-run
Real output of the -y run:
About to delete 1 repo(s) totalling 551.0M.
- model/openai-community/gpt2 (entire repo)
Cache deletion done. Saved 551.0M.
ā Deleted 1 repo(s) and 1 revision(s); freed 551.0M.
df confirmed it: the 2 GB tmpfs went back to 120K used. hf cache rm also accepts revision hashes and hf:// file URIs if you only want to drop one revision or file. Both rm and prune take --cache-dir when the cache is not in the default place.
On huggingface_hub 0.x the old commands still run: huggingface-cli scan-cache on 0.36.2 printed "'huggingface-cli scan-cache' is deprecated. Use 'hf cache scan' instead." and then worked, and huggingface-cli delete-cache still offered its --disable-tui and --sort options.
Fix 3: Clean Up Leftovers with hf cache prune
hf cache prune removes detached revisions, incomplete downloads and, on 1.33.0, shared blobs no model folder points to any more. Run it with --dry-run first:
hf cache prune --dry-run
hf cache prune -y
After a download killed partway (plain HTTP), the dry run found the partial file:
About to delete 1 incomplete download(s) (21.0M total).
ā Dry run: no files were deleted.
and the real run removed it: "Deleted 1 incomplete download(s); freed 21.0M." A killed Xet download left a 0-byte .incomplete file instead, and prune removed that too.
The bigger reason to know prune is manual deletion. Deleting a model folder by hand used to be enough. On 1.33.0 with a Xet download, it was not:
rm -rf ~/.cache/huggingface/hub/models--openai-community--gpt2
df -h # before and after in my run: 523M used both times
The 548 MB file survived in hub/blobs/ because that hard link was still there, and hf cache ls said "No results found." so nothing looked wrong. hf cache prune found it:
About to delete 1 unreferenced shared blob(s) (548.1M total).
Cache deletion done. Saved 0.0.
ā Deleted 1 unreferenced shared blob(s); freed 548.1M.
(The "Saved 0.0." line is what 1.33.0 printed; the last line and df, back to 76K used, show the space really came back.) So: use hf cache rm to delete models, and if you already deleted folders by hand, run hf cache prune afterwards.
Fix 4: Symlink the Cache to Another Drive
When you cannot set environment variables (a fixed entrypoint, a tool that launches its own processes) but do have a bigger disk, move the directory and leave a symlink:
mkdir -p /data
mv ~/.cache/huggingface /data/huggingface
ln -s /data/huggingface ~/.cache/huggingface
I tested this with no HF_HOME set: huggingface_hub still reported /root/.cache/huggingface/hub as its cache, and the downloaded config.json ended up under the symlink target on the other filesystem.
In Docker, mount the big disk and point HF_HOME at it, which is how every test in this post was run:
docker run --gpus all \
-v /data/huggingface:/hf -e HF_HOME=/hf \
my-ml-image python train.py
Fix 5: Per-Call Cache Override with cache_dir
For one script, pass cache_dir. It overrides HF_HOME for that call only:
from transformers import AutoConfig
AutoConfig.from_pretrained("openai-community/gpt2", cache_dir="/big/cd")
# /big/cd now contains: .locks, CACHEDIR.TAG, models--openai-community--gpt2
AutoModel*.from_pretrained, AutoTokenizer.from_pretrained, hf_hub_download and snapshot_download take the same argument. If you use it in several places, read the path from one config value or an environment variable so moving the cache is a one-line change.
Fix 6: Download Only the Files You Need
If you use snapshot_download(), filter the files. One trap: in allow_patterns, * also matches /. My first try with allow_patterns=["*.json", "*.safetensors"] on gpt2 also pulled the onnx/ subfolder's JSON files. Naming top-level files explicitly gave exactly what a safetensors load needs:
from huggingface_hub import snapshot_download
import os
path = snapshot_download(
"openai-community/gpt2",
allow_patterns=["*.safetensors", "config.json", "generation_config.json",
"tokenizer*", "vocab.json", "merges.txt"],
)
print(sorted(os.listdir(path)))
# ['config.json', 'generation_config.json', 'merges.txt', 'model.safetensors',
# 'tokenizer.json', 'tokenizer_config.json', 'vocab.json']
That came to 551.0M in hf cache ls, against about 5.6 GB for the whole repo. ignore_patterns works the other way round, for example ignore_patterns=["*.bin", "*.h5", "*.msgpack", "*.ot", "onnx/*"].
Quick-Reference: Which Fix to Use
| Situation | Best fix | Effort |
|---|---|---|
| You have a second disk or data partition | export HF_HOME=/data/huggingface in .bashrc |
One-time, 2 commands |
| Old models you no longer need | hf cache rm model/<repo_id> |
Minutes |
| Deleted folders by hand but space did not come back, or a killed download | hf cache prune |
One command |
| Cannot change env vars | Symlink ~/.cache/huggingface to a larger partition |
Three commands |
| One script, no system changes | from_pretrained(..., cache_dir="/data/hf_cache") |
One line of code |
snapshot_download pulling every weight format |
allow_patterns / ignore_patterns |
One argument |
Which one should you actually do
If you have any access to a bigger disk, set HF_HOME once and stop thinking about it. It fixes the problem at the source instead of every few weeks when the cache fills again. For cleanup, stick to hf cache rm and hf cache prune rather than rm -rf: on current huggingface_hub the files on disk are no longer laid out one folder per model, and prune is what knows where the space went.