Fix CUDA error: device-side assert triggered in PyTorch
- Look at stderr just above the traceback. PyTorch prints the kernel that failed, for example
Loss.cu:250: nll_loss_forward_reduce_cuda_kernel_2d ... Assertion `t >= 0 && t < n_classes` failed.That line alone usually names the cause. - Rerun with
CUDA_LAUNCH_BLOCKING=1so the traceback points at the real call, or run the same batch on CPU to getIndexError: Target 10 is out of bounds.orIndexError: index out of range in self. - Fix the data: class labels must be in
[0, num_classes)(useignore_indexfor 255 or -1), and token IDs must be< num_embeddings. - Restart the Python process (or Jupyter kernel). After the assert every later CUDA call in that process fails with the same error.
This is the real output on PyTorch 2.4.1 for a CrossEntropyLoss call where
one label is 10 and the model has 10 classes:
../aten/src/ATen/native/cuda/Loss.cu:250: nll_loss_forward_reduce_cuda_kernel_2d: block: [0,0,0], thread: [2,0,0] Assertion `t >= 0 && t < n_classes` failed.
Traceback (most recent call last):
File "/tmp/cdsa/cdsa.py", line 17, in <module>
print("loss", loss.item()) # line B
^^^^^^^^^^^
RuntimeError: CUDA error: device-side assert triggered
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
The traceback blames loss.item() on line 17. The bad call was
the loss on line 15, which returned without complaint (the script printed
"loss call returned" between the two). The first line, printed by the
kernel itself, is the useful one: thread 2 of the loss kernel saw a target
outside [0, n_classes). In this small batch the thread number
matched the position of the bad label (index 2 here, index 1 in the -1
test), which is a handy hint but not guaranteed for large batches.
TORCH_USE_CUDA_DSA is a build-time flag for PyTorch itself.
Setting it as an environment variable on a pip-installed wheel does nothing,
so skip that hint.
Reproduced on Real Hardware
Each case ran as a fresh process with small random tensors (a Linear(64, 10) classifier, a batch of 4, an Embedding(1000, 64)).
| What I ran | GPU (default) | CPU | Fix that worked |
|---|---|---|---|
CrossEntropyLoss, labels [3, 7, 10, 1], 10 classes | Assertion t >= 0 && t < n_classes failed, then device-side assert triggered at loss.item() | IndexError: Target 10 is out of bounds. | Labels [3, 7, 9, 1]: loss 2.1810 |
CrossEntropyLoss, label -1 | Same assertion, same error | IndexError: Target -1 is out of bounds. | ignore_index set to the padding value |
NLLLoss on log_softmax, label 10 | Same assertion (same kernel) | IndexError: Target 10 is out of bounds. | Same as above |
Embedding(1000, 64), token ID 1000 | 64 lines of Indexing.cu:1231: indexSelectSmallIndex ... Assertion `srcIndex < srcSelectDimSize` failed., then the assert at the next sync | IndexError: index out of range in self | ID 999: output shape [1, 4, 64] |
Either case with CUDA_LAUNCH_BLOCKING=1 | Traceback moves to the loss or F.embedding call | n/a | Use it to locate, then fix the data |
| Any CUDA op after catching the assert | RuntimeError: CUDA error: device-side assert triggered on torch.zeros(1, device="cuda") | n/a | Restart the process; a new process ran fine |
Logits containing inf and nan | No assert, loss is nan | Loss is nan | Not this error, see Fix 3 |
| Float labels from a NumPy array | "nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Double' | expected scalar type Long but found Double | .long(): loss 1.4583. Not this error, see Fix 4 |
So on current PyTorch this assert comes from index bounds checks. NaN values
and wrong label dtypes, which older guides (including an earlier version of
this one) list as causes, produced a nan loss or an ordinary
error in these runs, not a device-side assert.
Why the Traceback Points at the Wrong Line
PyTorch queues CUDA kernels and returns to Python without waiting for them.
When a kernel's bounds check fails, the error is only reported at the next
call that waits for the GPU, such as .item(),
copying a tensor to CPU, or
torch.cuda.synchronize(). In the run
above that was loss.item(), two lines after the loss call.
In a training loop it is often loss.backward() or a logging
line.
The process is broken after the assert
A device-side assert is not a normal exception you can catch and move on
from. I wrapped the bad loss call in try/except RuntimeError,
caught the error, and then ran torch.zeros(1, device="cuda").
It failed with the same RuntimeError: CUDA error: device-side assert
triggered. The CUDA context of that process stays in the error
state, so every later GPU call fails. The only recovery is a new process:
restart the script, or restart the kernel in Jupyter. A fresh process on the
same GPU ran normally right afterwards.
Fix 0: Find the Real Failing Line (Do This First)
Before you change any model code, get an error that names the bad call and the bad value. There are two ways: make the GPU run synchronously, or run the same batch on CPU, where the bounds checks raise ordinary Python exceptions.
Option A: Set CUDA_LAUNCH_BLOCKING=1
Run your script with the environment variable set. This forces every CUDA kernel launch to block until the kernel completes, which makes GPU errors surface at the correct call site.
CUDA_LAUNCH_BLOCKING=1 python train.py
Or inside a Python script before any CUDA calls:
import os
os.environ["CUDA_LAUNCH_BLOCKING"] = "1"
import torch
# ... rest of your code
Set it before the first CUDA call; changing it later in a running process
has no effect. With blocking enabled, the same label-10 script failed on
line 15 (the loss call) instead of line 17, and the stack went through
loss.py and functional.py down to
torch._C._nn.cross_entropy_loss. For the embedding case it
ended in F.embedding. The error text itself stays
RuntimeError: CUDA error: device-side assert triggered, so
you still need the kernel line or a CPU run to see which value was bad.
Blocking mode slows training, so use it only while debugging.
Option B: Move Model and Data to CPU
CPU kernels check the same bounds and raise a normal Python exception with the bad value in the message. Move the model and one failing batch to CPU for a single forward pass:
import torch
import torch.nn as nn
# Assume you have these from your normal training setup
model = MyModel()
inputs = next(iter(train_loader)) # whatever your loader returns
# Move to CPU for diagnosis
model_cpu = model.cpu()
# Unpack your batch; adjust to your actual data format
x, labels = inputs
x = x.cpu()
labels = labels.cpu()
# Single forward + loss on CPU, the real error will appear here
logits = model_cpu(x)
loss = criterion(logits, labels)
print("Forward pass succeeded on CPU, so check the other batches")
On CPU the label case raised
IndexError: Target 10 is out of bounds.
and the embedding case raised
IndexError: index out of range in self
(both real output from the runs above). The sections below fix each one.
Fix 1: Label Index Out of Range (CrossEntropyLoss)
This is the case the kernel line above describes, and the one I would check first.
nn.CrossEntropyLoss expects each label
to be an integer in the range
[0, num_classes). If any label is equal
to num_classes or higher, or is negative
and not the ignore_index value
(default -100), the CUDA kernel triggers the device-side
assert. nn.NLLLoss uses the same kernel
and fails the same way; a label of -1 gave
IndexError: Target -1 is out of bounds. on CPU.
Broken code
import torch
import torch.nn as nn
num_classes = 10
model = nn.Linear(64, num_classes).cuda()
criterion = nn.CrossEntropyLoss()
# Batch of 4 samples, labels should be 0..9
logits = model(torch.randn(4, 64).cuda())
# BUG: one label equals num_classes (10), which is out of range [0, 10)
labels = torch.tensor([3, 7, 10, 1]).cuda() # 10 is invalid
# On GPU: RuntimeError: CUDA error: device-side assert triggered
# On CPU: IndexError: Target 10 is out of bounds.
loss = criterion(logits, labels)
How to check
def check_labels(labels, num_classes, ignore_index=-100):
mask = labels != ignore_index
valid = labels[mask]
if valid.min() < 0:
raise ValueError(
f"Label contains negative value {valid.min().item()} "
f"(not equal to ignore_index={ignore_index})"
)
if valid.max() >= num_classes:
raise ValueError(
f"Label {valid.max().item()} is out of range "
f"[0, {num_classes}). Did you accidentally include num_classes as a label?"
)
print(f"Labels OK: min={valid.min().item()}, max={valid.max().item()}, "
f"num_classes={num_classes}")
check_labels(labels, num_classes=10)
Fixed code
import torch
import torch.nn as nn
num_classes = 10
model = nn.Linear(64, num_classes).cuda()
criterion = nn.CrossEntropyLoss()
logits = model(torch.randn(4, 64).cuda())
# Correct: labels in [0, num_classes) = [0, 9]
labels = torch.tensor([3, 7, 9, 1]).cuda()
loss = criterion(logits, labels)
print(f"loss = {loss.item():.4f}")
# loss = 2.1810 (with torch.manual_seed(0))
Run on the bad labels [3, 7, 10, 1], the
check_labels function above raises
ValueError: Label 10 is out of range [0, 10). Did you accidentally include num_classes as a label?
before any kernel runs, so the process stays usable.
1-indexed labels and ignore values
A common version is an off-by-one. Datasets that are 1-indexed (1
through N) instead of 0-indexed (0 through N-1) will pass label
N when the model has
N output classes, triggering the assert.
Subtract 1 from every label when loading data from such a dataset:
labels = labels - 1 # convert 1-indexed to 0-indexed
Also watch for datasets that use a special "background" or "ignore" class
encoded as 255 in segmentation tasks.
Pass that value as
ignore_index=255 (tested: labels
[3, 255, 9, 1] gave a normal loss of 2.0084) to
nn.CrossEntropyLoss rather than leaving
it as a real label:
criterion = nn.CrossEntropyLoss(ignore_index=255)
Fix 2: Embedding Index Out of Range
nn.Embedding(vocab_size, embedding_dim)
creates a lookup table with indices
0 through
vocab_size - 1. Passing an index that
equals or exceeds vocab_size triggers the
same device-side assert.
What it looks like
import torch
import torch.nn as nn
vocab_size = 1000
embedding = nn.Embedding(vocab_size, 64).cuda()
# Token IDs in a tokenized sentence; one ID exceeds the vocabulary
token_ids = torch.tensor([[12, 45, 1000, 7]]).cuda()
# 1000 is out of range: valid indices are 0..999
# On GPU: RuntimeError: CUDA error: device-side assert triggered
# On CPU: IndexError: index out of range in self
out = embedding(token_ids)
Catching it before it hits the kernel
def check_embedding_indices(token_ids, vocab_size):
if token_ids.min() < 0:
raise ValueError(
f"Embedding index contains negative value: {token_ids.min().item()}"
)
if token_ids.max() >= vocab_size:
bad_ids = token_ids[token_ids >= vocab_size].unique().tolist()
raise ValueError(
f"Embedding index out of range. "
f"vocab_size={vocab_size}, offending indices: {bad_ids}"
)
print(f"Embedding indices OK: min={token_ids.min().item()}, "
f"max={token_ids.max().item()}, vocab_size={vocab_size}")
check_embedding_indices(token_ids, vocab_size=1000)
Fixed code
import torch
import torch.nn as nn
vocab_size = 1000
embedding = nn.Embedding(vocab_size, 64).cuda()
# All indices in [0, vocab_size)
token_ids = torch.tensor([[12, 45, 999, 7]]).cuda()
out = embedding(token_ids)
print(f"Embedding output shape: {out.shape}")
This prints Embedding output shape: torch.Size([1, 4, 64]). On
the bad input, check_embedding_indices reports the offending
IDs ([1000]) instead of crashing the CUDA context. The GPU
version of this bug prints one
Indexing.cu:1231: indexSelectSmallIndex ... Assertion `srcIndex < srcSelectDimSize` failed.
line per GPU thread (64 in my run), so look
for srcIndex in a long log.
Common sources of out-of-range embedding indices
- Tokenizer vocabulary mismatch. The tokenizer has more tokens than the model's embedding table, so any token past the table size (often a rare one that only shows up in the test set) triggers the assert.
-
Special tokens not added to the model. You added
[PAD],[CLS],[SEP]to the tokenizer but did not callmodel.resize_token_embeddings(len(tokenizer)). -
Data pre-processing bug. An integer encoding step maps an
unknown category to an ID equal to the number of categories rather than a
reserved
UNKindex inside the range.
# After adding special tokens to a HuggingFace tokenizer:
tokenizer.add_special_tokens({"pad_token": "[PAD]"})
model.resize_token_embeddings(len(tokenizer)) # expand embedding table
Fix 3: NaN or Inf in Inputs (a different failure)
NaN or Inf logits are often listed as a cause of this assert. In my test
they were not: cross-entropy on logits containing
inf and
nan returned a loss of
nan on both GPU and CPU, with no assert.
The real symptom of NaN is a loss that turns into nan and
stays there. It is still worth catching early, because a NaN that reaches an
index computation (for example a value cast to an integer and used as a
label or token ID) can become an out-of-range index and then trigger the
assert.
What it looks like
import torch
import torch.nn as nn
# Simulate logits that have gone to infinity (e.g., after gradient explosion)
logits = torch.tensor([[1.0, float("inf"), -2.0],
[0.5, 0.1, float("nan")]]).cuda()
labels = torch.tensor([1, 2]).cuda()
criterion = nn.CrossEntropyLoss()
# No assert on torch 2.4.1: prints loss nan
loss = criterion(logits, labels)
print("loss", loss.item())
How to check
def check_finite(tensor, name="tensor"):
if torch.isnan(tensor).any():
nan_count = torch.isnan(tensor).sum().item()
raise ValueError(f"{name} contains {nan_count} NaN value(s)")
if torch.isinf(tensor).any():
inf_count = torch.isinf(tensor).sum().item()
raise ValueError(f"{name} contains {inf_count} Inf value(s)")
print(f"{name}: finite, min={tensor.min().item():.4f}, "
f"max={tensor.max().item():.4f}")
check_finite(logits, "logits")
# ValueError: logits contains 1 NaN value(s) (for the tensor above)
Add these checks immediately before the loss call when debugging. You can
also register a forward hook on any module to check all its outputs
automatically during a diagnostic run. Fed an all-NaN input, it raised
RuntimeError: NaN/Inf detected in output of Linear:
def nan_hook(module, inputs, output):
if isinstance(output, torch.Tensor):
if torch.isnan(output).any() or torch.isinf(output).any():
raise RuntimeError(
f"NaN/Inf detected in output of {module.__class__.__name__}"
)
# Register on every submodule
for name, module in model.named_modules():
module.register_forward_hook(nan_hook)
If the loss does go to nan, the usual suspects are a learning
rate that is too high (clip with
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)),
a division or torch.log without an
epsilon, or float16 overflow above
65,504 under AMP. For AMP, use
torch.amp.GradScaler("cuda") and watch
scaler.get_scale(): a scale that keeps
shrinking means overflowing steps are being skipped. This loop ran on the
GTX 1070 and printed loss=1.5897, amp_scale=65536.0 on its first
step with a small test model:
import torch
# torch.cuda.amp.autocast / GradScaler still work on 2.4.1 but print a
# FutureWarning; this is the current API
scaler = torch.amp.GradScaler("cuda")
for batch in train_loader:
optimizer.zero_grad()
with torch.autocast("cuda", dtype=torch.float16):
logits = model(batch["input_ids"].cuda())
loss = criterion(logits, batch["labels"].cuda())
scaler.scale(loss).backward()
# Clip gradients BEFORE the optimizer step
scaler.unscale_(optimizer)
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
scaler.step(optimizer)
scaler.update()
print(f"loss={loss.item():.4f}, amp_scale={scaler.get_scale():.1f}")
Fix 4: Wrong Label dtype (a different error)
nn.CrossEntropyLoss and other
classification losses require labels to be
torch.long (64-bit integer) when you
pass class indices. Float labels happen when you load them from a NumPy
float array or a pandas column without casting. This does not trigger the
device-side assert. PyTorch rejects the dtype before the kernel runs, with
a normal error that leaves the GPU usable.
Broken code
import torch
import torch.nn as nn
import numpy as np
num_classes = 5
model = nn.Linear(32, num_classes).cuda()
criterion = nn.CrossEntropyLoss()
logits = model(torch.randn(8, 32).cuda())
# Labels come from a numpy float array (common when reading a CSV)
raw_labels = np.array([0.0, 2.0, 4.0, 1.0, 3.0, 0.0, 2.0, 1.0])
labels = torch.tensor(raw_labels).cuda() # dtype = torch.float64
# On GPU: RuntimeError: "nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Double'
# On CPU: RuntimeError: expected scalar type Long but found Double
loss = criterion(logits, labels)
A NumPy float array becomes torch.float64, hence 'Double'.
With torch.float32 labels the messages end in 'Float' instead.
How to check
def check_label_dtype(labels):
if labels.dtype != torch.long:
raise TypeError(
f"Labels have dtype {labels.dtype}. "
f"CrossEntropyLoss requires torch.long (torch.int64). "
f"Fix: labels = labels.long()"
)
print(f"Label dtype OK: {labels.dtype}")
check_label_dtype(labels)
Fixed code
import torch
import torch.nn as nn
import numpy as np
num_classes = 5
model = nn.Linear(32, num_classes).cuda()
criterion = nn.CrossEntropyLoss()
logits = model(torch.randn(8, 32).cuda())
raw_labels = np.array([0.0, 2.0, 4.0, 1.0, 3.0, 0.0, 2.0, 1.0])
# Cast explicitly to long before creating the tensor
labels = torch.tensor(raw_labels, dtype=torch.long).cuda()
# Or, if you already have a float tensor: labels = labels.long()
loss = criterion(logits, labels)
print(f"loss = {loss.item():.4f}") # loss = 1.4583 in my run
A similar dtype issue appears with
nn.BCELoss for binary classification:
it expects labels to be
torch.float32, not
torch.long. Passing long labels to
nn.BCEWithLogitsLoss raised
RuntimeError: result type Float can't be cast to the desired output type Long
on both GPU and CPU. If you switch between
multi-class and binary setups, double-check both the loss function and the
label dtype.
# Multi-class: nn.CrossEntropyLoss, labels must be torch.long
# Binary: nn.BCELoss / nn.BCEWithLogitsLoss, labels must be torch.float
# Binary example
bce = nn.BCEWithLogitsLoss()
logits_binary = model_binary(x).squeeze(1) # shape [B]
labels_binary = labels_01.float() # cast to float32
loss = bce(logits_binary, labels_binary)
Diagnostic Checklist
When you hit
RuntimeError: CUDA error: device-side assert triggered,
work through this list in order:
| # | Step | What to look for |
|---|---|---|
| 0 | Read the Assertion ... failed line on stderr, then run with CUDA_LAUNCH_BLOCKING=1 or on CPU |
n_classes means labels, srcIndex means an indexing op such as nn.Embedding |
| 1 | Check label range | labels.min() >= 0 and labels.max() < num_classes, ignoring ignore_index |
| 2 | Check embedding indices | token_ids.max() < vocab_size |
| 3 | Check for NaN / Inf (causes a nan loss, not this assert) |
torch.isnan(x).any() and torch.isinf(x).any() |
| 4 | Check label dtype (causes a dtype error, not this assert) | labels.dtype == torch.long for CrossEntropyLoss |
| 5 | Restart the process after fixing | The CUDA context stays broken after the assert. For tensors on different devices, a separate error, see the PyTorch device mismatch fix |
Run these checks as assertions at the top of your training loop during debugging.
DEBUG = True # set False for production
def debug_check_batch(inputs, labels, logits, num_classes):
if not DEBUG:
return
assert not torch.isnan(inputs).any(), "NaN in inputs"
assert not torch.isinf(inputs).any(), "Inf in inputs"
assert labels.dtype == torch.long, f"Expected torch.long, got {labels.dtype}"
assert labels.min() >= 0, f"Negative label: {labels.min().item()}"
assert labels.max() < num_classes, (
f"Label {labels.max().item()} out of range [0, {num_classes})"
)
assert not torch.isnan(logits).any(), "NaN in logits"
assert not torch.isinf(logits).any(), "Inf in logits"
# Inside training loop:
for batch in train_loader:
inputs, labels = batch
inputs = inputs.cuda()
labels = labels.cuda()
logits = model(inputs)
debug_check_batch(inputs, labels, logits, num_classes=NUM_CLASSES)
loss = criterion(logits, labels)
loss.backward()
optimizer.step()
optimizer.zero_grad()
Short Version
The assert comes from a bounds check in your data, found late because the
GPU runs asynchronously. The Assertion ... failed line on
stderr, CUDA_LAUNCH_BLOCKING=1, or a CPU
run each get you from the misleading traceback to the real call.
In every run for this article the cause was an
index out of bounds: a class label outside
[0, num_classes) or a token ID past the embedding table. NaN
inputs and float labels are real bugs too, but they show up as a
nan loss or a dtype error rather than this assert. Put the
corresponding assertion from the checklist above at the top of your
training loop while you debug, confirm it actually catches the bad batch,
then go fix wherever the data pipeline produced it in the first place, and
restart the process before you rerun.
Related articles
Where to next
New to PyTorch ยท step 5 of 7