Skip to content

torch.cuda.OutOfMemoryError: How to Fix GPU Out of Memory in PyTorch

Tested with: PyTorch 2.4.1 (CUDA 12.1), Python 3.12, NVIDIA GeForce GTX 1070 8 GB, driver 580. Last run 2026-09-27.

TL;DR, in the order to try them:
  1. If the OOM happens during validation or inference, wrap it in torch.no_grad(). On the test model below this cut peak memory from 4,621 MiB to 781 MiB.
  2. Lower the batch size and use gradient accumulation to keep the effective batch size. Batch 64 as 4 micro-batches of 16: 4,896 MiB down to 1,249 MiB.
  3. Train with mixed precision via torch.autocast("cuda") and torch.amp.GradScaler("cuda"): 4,896 MiB down to 3,367 MiB.
  4. Gradient checkpointing, with use_reentrant=False and no closure over a loop variable: 4,896 MiB down to 3,104 MiB.
torch.cuda.empty_cache() is not a fix for an OOM inside your training step. See Fix 2.

On PyTorch 2.4 the error looks like this (real output from the reproduction below):

torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 1024.00 MiB.
GPU 0 has a total capacity of 7.91 GiB of which 631.00 MiB is free.
Including non-PyTorch memory, this process has 7.14 GiB memory in use.
Of the allocated memory 7.05 GiB is allocated by PyTorch, and 856.50 KiB
is reserved by PyTorch but unallocated. If reserved but unallocated memory
is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to
avoid fragmentation.

The class prints as torch.OutOfMemoryError; torch.cuda.OutOfMemoryError is the same class, so except torch.cuda.OutOfMemoryError still catches it. Older releases printed a different message ending in a max_split_size_mb hint.

Read the last two numbers before changing anything. If "allocated by PyTorch" is close to the total, your tensors really do not fit and you need one of the fixes below. If "reserved but unallocated" is large (hundreds of MiB or more), the memory is fragmented, and PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True is the setting to try. In the run above it was under 1 MiB, so fragmentation was not the problem.

A crashed process does not keep the GPU: after the OOM above, nvidia-smi showed 159 MiB in use on the card, the same as before the run. If memory stays taken after a crash, something is still alive, typically a Jupyter kernel or a DataLoader worker process.


Reproduced on Real Hardware

Test model: a stem conv plus 8 blocks of Conv2d(64, 64, 3) + BatchNorm2d + ReLU, random 3×128×128 inputs, 10 classes, Adam, 3 training steps. Each configuration ran in a fresh process, and the number is torch.cuda.max_memory_allocated().

Configuration (batch 64 unless noted)Peak allocatedSaving
train, batch 161,239 MiB
train, batch 64 (baseline)4,896 MiB
train, batch 967,334 MiB
train, batch 256OOM (the error above)
train, 4 × 16 gradient accumulation1,249 MiB-74%
train, autocast float16 + GradScaler3,367 MiB-31%
train, checkpoint every block3,104 MiB-37%
forward in eval() without no_grad4,621 MiB
forward in eval() with no_grad781 MiB-83%

What the numbers show:

  • Activation memory scales with batch size almost linearly: 16, 64 and 96 cost 1,239, 4,896 and 7,334 MiB. That makes it easy to estimate the largest batch that fits from two small runs.
  • Mixed precision saved 31%, not the "halves memory" often quoted, because weights, gradients and Adam state stay in float32 and some ops (such as the loss) run in float32 under autocast. On the GTX 1070, which has no tensor cores, it also did not make training faster.
  • Leaving out torch.no_grad() in evaluation cost 5.9x the memory here. It is the cheapest fix on this list.

Why GPU OOM Happens

PyTorch allocates GPU memory for:

  • Model parameters and their gradients (2x parameter count in float32)
  • Optimizer state: Adam stores two additional tensors per parameter (momentum and variance), which is 2x the parameter memory
  • Activations from the forward pass, held in memory until the backward pass completes
  • The current batch of input data and labels

A model with 100M parameters in float32 uses 400 MB for weights, 400 MB for gradients, and 800 MB for Adam state, before a single batch touches GPU. Large batches or deep networks with many intermediate activations push past GPU limits.


Fix 1: Reduce Batch Size

Activation memory grows with batch size almost linearly (1,239 MiB at batch 16, 4,896 MiB at batch 64 in the test above), so halving the batch roughly halves the activation part.

from torch.utils.data import DataLoader

# Before: OOM (on your GPU; batch 64 fit the 8 GB test card, batch 256 did not)
train_loader = DataLoader(dataset, batch_size=64, shuffle=True)

# Fix: reduce batch size
train_loader = DataLoader(dataset, batch_size=16, shuffle=True)

If you need the effective batch size to stay large for training stability, use gradient accumulation: run multiple small batches before calling optimizer.step():

accumulation_steps = 4  # effective batch size = batch_size * accumulation_steps

optimizer.zero_grad()

for i, (inputs, labels) in enumerate(train_loader):
    inputs, labels = inputs.cuda(), labels.cuda()
    outputs = model(inputs)
    loss = criterion(outputs, labels)

    # Scale loss to account for accumulation
    loss = loss / accumulation_steps
    loss.backward()

    if (i + 1) % accumulation_steps == 0:
        optimizer.step()
        optimizer.zero_grad()

Fix 2: Clear the Cache Correctly

torch.cuda.empty_cache() is widely misunderstood. It does not free memory that PyTorch is still using; it only releases memory from PyTorch's internal allocator cache back to CUDA. It will not fix an OOM that happens during a forward pass.

import torch
import gc

# After a crash or between training runs, do both:
gc.collect()
torch.cuda.empty_cache()

# Check what is actually using memory
print(torch.cuda.memory_summary())

What it actually does, measured after deleting a 1 GiB tensor:

after del:         allocated 65 MiB   reserved 1098 MiB
after empty_cache: allocated 65 MiB   reserved 70 MiB

Allocated memory did not change. empty_cache() only gave the cached 1 GiB back to the driver, which helps other processes on the GPU (and makes nvidia-smi look better) but gives your own process nothing it could not already reuse.

A common cause of unexpectedly high memory is holding references to tensors across iterations. This pattern accumulates memory every step:

# BAD: loss is a tensor on GPU, appending it keeps the computation graph alive
losses = []
for batch in train_loader:
    loss = compute_loss(batch)
    losses.append(loss)  # holds graph in memory!

# GOOD: extract the scalar value
losses = []
for batch in train_loader:
    loss = compute_loss(batch)
    losses.append(loss.item())  # detaches from graph

Measured, the leak only happens when the graph has not been freed by backward(). In a training loop, backward() releases the saved activations, so appending the tensor grew memory by 20.5 MiB over 20 steps, the same as .item() (20.4 MiB, all of it Adam state created on the first step). In a validation loop without torch.no_grad(), appending the loss tensor kept each batch's full graph alive: 2,184 MiB after 4 batches of 8, against 8 MiB with .item(). Use .item() anyway, but if you are hunting a leak, look at your evaluation code first.


Fix 3: Use torch.no_grad() During Evaluation

During inference and validation, you do not need gradients. Without torch.no_grad(), PyTorch builds the full computation graph and stores all intermediate activations for no reason. On the test model, the same forward pass took 4,621 MiB without it and 781 MiB with it.

import torch

model.eval()

# BAD: still builds computation graph
for inputs, labels in val_loader:
    outputs = model(inputs.cuda())
    loss = criterion(outputs, labels.cuda())

# GOOD: disables gradient tracking entirely
with torch.no_grad():
    for inputs, labels in val_loader:
        outputs = model(inputs.cuda())
        loss = criterion(outputs, labels.cuda())

torch.no_grad() can also be used as a decorator on inference functions:

@torch.no_grad()
def predict(model, inputs):
    model.eval()
    return model(inputs.cuda())

Fix 4: Gradient Checkpointing

Gradient checkpointing trades compute for memory. Instead of storing all intermediate activations during the forward pass (needed for backprop), it stores only the inputs of each checkpointed segment and recomputes the rest during the backward pass, which costs roughly one extra forward pass. On the test model, checkpointing each of the 8 blocks cut peak memory from 4,896 MiB to 3,104 MiB.

A common version of this code is broken. An earlier version of this post used checkpoint.checkpoint(lambda inp: self.relu(layer(inp)), x) inside the for layer in self.layers loop. The lambda looks up layer when it is called, and checkpointing calls it again during backward(), after the loop has finished, so every segment is recomputed with the last layer. On PyTorch 2.4.1 with a 12-layer model, only layer 11 received a gradient, and that gradient was wrong. No error is raised; 11 of 12 layers silently stop training. It also warns that use_reentrant should be passed explicitly.

The fixed version binds the layer as a default argument and passes use_reentrant=False, the variant PyTorch recommends. Its gradients matched the non-checkpointed model:

import torch
import torch.utils.checkpoint as checkpoint

class CheckpointedModel(torch.nn.Module):
    def __init__(self):
        super().__init__()
        self.layers = torch.nn.ModuleList([
            torch.nn.Linear(512, 512) for _ in range(12)
        ])
        self.relu = torch.nn.ReLU()

    def forward(self, x):
        for layer in self.layers:
            # l=layer binds the current layer; without it every
            # recomputation during backward uses the last layer
            x = checkpoint.checkpoint(
                lambda inp, l=layer: self.relu(l(inp)), x, use_reentrant=False
            )
        return x

If each block is already an nn.Module, pass it directly and skip the lambda: x = checkpoint.checkpoint(block, x, use_reentrant=False). That is what the measured test model does.

For HuggingFace Transformers models, gradient checkpointing is a single flag:

from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased")
model.gradient_checkpointing_enable()  # one line
model.cuda()

Fix 5: Mixed Precision Training (torch.amp)

Float32 tensors use 4 bytes per element. Float16 uses 2 bytes. Mixed precision training runs most of the forward pass in float16, which shrinks the stored activations, while weights and the optimizer update stay in float32. On the test model it saved 31% (4,896 MiB to 3,367 MiB), less than half, because only part of the memory is activations.

The old torch.cuda.amp imports still work but are deprecated. On PyTorch 2.4.1 they emit FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead., and the same for GradScaler. The current form:

import torch

model = model.cuda()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-4)
scaler = torch.amp.GradScaler("cuda")  # avoids float16 gradient underflow

for inputs, labels in train_loader:
    inputs, labels = inputs.cuda(), labels.cuda()
    optimizer.zero_grad()

    # Forward pass in float16
    with torch.autocast("cuda", dtype=torch.float16):
        outputs = model(inputs)
        loss = criterion(outputs, labels)

    # Backward pass with scaled gradients
    scaler.scale(loss).backward()
    scaler.step(optimizer)
    scaler.update()

Mixed precision can be combined with gradient checkpointing and gradient accumulation. With autocast alone, batch 96 went from 7,334 MiB (barely fitting on 8 GB) to 5,040 MiB.


Profiling Memory Usage

To see where the memory goes in your own model, measure one training step:

import torch

# Snapshot before and after a training step
torch.cuda.reset_peak_memory_stats()

# ... run your training step ...

print(torch.cuda.memory_summary(device=None, abbreviated=True))

# Key stats:
print(f"Allocated: {torch.cuda.memory_allocated() / 1e9:.2f} GB")
print(f"Reserved:  {torch.cuda.memory_reserved() / 1e9:.2f} GB")
print(f"Peak:      {torch.cuda.max_memory_allocated() / 1e9:.2f} GB")

To find which line of code triggers the OOM, use PyTorch's memory snapshot tool. The underscore-prefixed functions are not a stable API, but the code below ran as written on 2.4.1 and wrote the snapshot file after a caught OOM:

import torch

torch.cuda.memory._record_memory_history(max_entries=100000)

try:
    # ... your training code ...
    pass
except torch.cuda.OutOfMemoryError:
    torch.cuda.memory._dump_snapshot("oom_snapshot.pickle")
    raise

Load the snapshot in the PyTorch memory visualizer at pytorch.org/memory_viz to see a timeline of every allocation.


Fix Priority Order

Use the order in the TL;DR at the top. Before any of them, check for tensors accumulated without .item() in evaluation code; that is a bug, not a capacity problem, and no amount of batch-size tuning fixes it. After enabling checkpointing, confirm every parameter still gets a gradient: [n for n, p in model.named_parameters() if p.grad is None] should be empty after backward().

For vector search and embedding workloads running alongside training, see FAISS index errors and memory patterns for GPU/CPU memory separation strategies.