Skip to content

Fix CUDA error: device-side assert triggered in PyTorch

Tested with: PyTorch 2.4.1 (CUDA 12.1), Python 3.12, NumPy 2.4, NVIDIA GeForce GTX 1070 8 GB, driver 580. Last run 2026-09-27.

TL;DR:
  1. Look at stderr just above the traceback. PyTorch prints the kernel that failed, for example Loss.cu:250: nll_loss_forward_reduce_cuda_kernel_2d ... Assertion `t >= 0 && t < n_classes` failed. That line alone usually names the cause.
  2. Rerun with CUDA_LAUNCH_BLOCKING=1 so the traceback points at the real call, or run the same batch on CPU to get IndexError: Target 10 is out of bounds. or IndexError: index out of range in self.
  3. Fix the data: class labels must be in [0, num_classes) (use ignore_index for 255 or -1), and token IDs must be < num_embeddings.
  4. Restart the Python process (or Jupyter kernel). After the assert every later CUDA call in that process fails with the same error.

This is the real output on PyTorch 2.4.1 for a CrossEntropyLoss call where one label is 10 and the model has 10 classes:

../aten/src/ATen/native/cuda/Loss.cu:250: nll_loss_forward_reduce_cuda_kernel_2d: block: [0,0,0], thread: [2,0,0] Assertion `t >= 0 && t < n_classes` failed.
Traceback (most recent call last):
  File "/tmp/cdsa/cdsa.py", line 17, in <module>
    print("loss", loss.item())             # line B
                  ^^^^^^^^^^^
RuntimeError: CUDA error: device-side assert triggered
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.

The traceback blames loss.item() on line 17. The bad call was the loss on line 15, which returned without complaint (the script printed "loss call returned" between the two). The first line, printed by the kernel itself, is the useful one: thread 2 of the loss kernel saw a target outside [0, n_classes). In this small batch the thread number matched the position of the bad label (index 2 here, index 1 in the -1 test), which is a handy hint but not guaranteed for large batches.

TORCH_USE_CUDA_DSA is a build-time flag for PyTorch itself. Setting it as an environment variable on a pip-installed wheel does nothing, so skip that hint.


Reproduced on Real Hardware

Each case ran as a fresh process with small random tensors (a Linear(64, 10) classifier, a batch of 4, an Embedding(1000, 64)).

What I ranGPU (default)CPUFix that worked
CrossEntropyLoss, labels [3, 7, 10, 1], 10 classesAssertion t >= 0 && t < n_classes failed, then device-side assert triggered at loss.item()IndexError: Target 10 is out of bounds.Labels [3, 7, 9, 1]: loss 2.1810
CrossEntropyLoss, label -1Same assertion, same errorIndexError: Target -1 is out of bounds.ignore_index set to the padding value
NLLLoss on log_softmax, label 10Same assertion (same kernel)IndexError: Target 10 is out of bounds.Same as above
Embedding(1000, 64), token ID 100064 lines of Indexing.cu:1231: indexSelectSmallIndex ... Assertion `srcIndex < srcSelectDimSize` failed., then the assert at the next syncIndexError: index out of range in selfID 999: output shape [1, 4, 64]
Either case with CUDA_LAUNCH_BLOCKING=1Traceback moves to the loss or F.embedding calln/aUse it to locate, then fix the data
Any CUDA op after catching the assertRuntimeError: CUDA error: device-side assert triggered on torch.zeros(1, device="cuda")n/aRestart the process; a new process ran fine
Logits containing inf and nanNo assert, loss is nanLoss is nanNot this error, see Fix 3
Float labels from a NumPy array"nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Double'expected scalar type Long but found Double.long(): loss 1.4583. Not this error, see Fix 4

So on current PyTorch this assert comes from index bounds checks. NaN values and wrong label dtypes, which older guides (including an earlier version of this one) list as causes, produced a nan loss or an ordinary error in these runs, not a device-side assert.


Why the Traceback Points at the Wrong Line

PyTorch queues CUDA kernels and returns to Python without waiting for them. When a kernel's bounds check fails, the error is only reported at the next call that waits for the GPU, such as .item(), copying a tensor to CPU, or torch.cuda.synchronize(). In the run above that was loss.item(), two lines after the loss call. In a training loop it is often loss.backward() or a logging line.

The process is broken after the assert

A device-side assert is not a normal exception you can catch and move on from. I wrapped the bad loss call in try/except RuntimeError, caught the error, and then ran torch.zeros(1, device="cuda"). It failed with the same RuntimeError: CUDA error: device-side assert triggered. The CUDA context of that process stays in the error state, so every later GPU call fails. The only recovery is a new process: restart the script, or restart the kernel in Jupyter. A fresh process on the same GPU ran normally right afterwards.


Fix 0: Find the Real Failing Line (Do This First)

Before you change any model code, get an error that names the bad call and the bad value. There are two ways: make the GPU run synchronously, or run the same batch on CPU, where the bounds checks raise ordinary Python exceptions.

Option A: Set CUDA_LAUNCH_BLOCKING=1

Run your script with the environment variable set. This forces every CUDA kernel launch to block until the kernel completes, which makes GPU errors surface at the correct call site.

CUDA_LAUNCH_BLOCKING=1 python train.py

Or inside a Python script before any CUDA calls:

import os
os.environ["CUDA_LAUNCH_BLOCKING"] = "1"

import torch
# ... rest of your code

Set it before the first CUDA call; changing it later in a running process has no effect. With blocking enabled, the same label-10 script failed on line 15 (the loss call) instead of line 17, and the stack went through loss.py and functional.py down to torch._C._nn.cross_entropy_loss. For the embedding case it ended in F.embedding. The error text itself stays RuntimeError: CUDA error: device-side assert triggered, so you still need the kernel line or a CPU run to see which value was bad. Blocking mode slows training, so use it only while debugging.

Option B: Move Model and Data to CPU

CPU kernels check the same bounds and raise a normal Python exception with the bad value in the message. Move the model and one failing batch to CPU for a single forward pass:

import torch
import torch.nn as nn

# Assume you have these from your normal training setup
model = MyModel()
inputs = next(iter(train_loader))   # whatever your loader returns

# Move to CPU for diagnosis
model_cpu = model.cpu()

# Unpack your batch; adjust to your actual data format
x, labels = inputs
x = x.cpu()
labels = labels.cpu()

# Single forward + loss on CPU, the real error will appear here
logits = model_cpu(x)
loss = criterion(logits, labels)
print("Forward pass succeeded on CPU, so check the other batches")

On CPU the label case raised IndexError: Target 10 is out of bounds. and the embedding case raised IndexError: index out of range in self (both real output from the runs above). The sections below fix each one.


Fix 1: Label Index Out of Range (CrossEntropyLoss)

This is the case the kernel line above describes, and the one I would check first.

nn.CrossEntropyLoss expects each label to be an integer in the range [0, num_classes). If any label is equal to num_classes or higher, or is negative and not the ignore_index value (default -100), the CUDA kernel triggers the device-side assert. nn.NLLLoss uses the same kernel and fails the same way; a label of -1 gave IndexError: Target -1 is out of bounds. on CPU.

Broken code

import torch
import torch.nn as nn

num_classes = 10
model = nn.Linear(64, num_classes).cuda()
criterion = nn.CrossEntropyLoss()

# Batch of 4 samples, labels should be 0..9
logits = model(torch.randn(4, 64).cuda())

# BUG: one label equals num_classes (10), which is out of range [0, 10)
labels = torch.tensor([3, 7, 10, 1]).cuda()  # 10 is invalid

# On GPU: RuntimeError: CUDA error: device-side assert triggered
# On CPU: IndexError: Target 10 is out of bounds.
loss = criterion(logits, labels)

How to check

def check_labels(labels, num_classes, ignore_index=-100):
    mask = labels != ignore_index
    valid = labels[mask]
    if valid.min() < 0:
        raise ValueError(
            f"Label contains negative value {valid.min().item()} "
            f"(not equal to ignore_index={ignore_index})"
        )
    if valid.max() >= num_classes:
        raise ValueError(
            f"Label {valid.max().item()} is out of range "
            f"[0, {num_classes}). Did you accidentally include num_classes as a label?"
        )
    print(f"Labels OK: min={valid.min().item()}, max={valid.max().item()}, "
          f"num_classes={num_classes}")

check_labels(labels, num_classes=10)

Fixed code

import torch
import torch.nn as nn

num_classes = 10
model = nn.Linear(64, num_classes).cuda()
criterion = nn.CrossEntropyLoss()

logits = model(torch.randn(4, 64).cuda())

# Correct: labels in [0, num_classes) = [0, 9]
labels = torch.tensor([3, 7, 9, 1]).cuda()

loss = criterion(logits, labels)
print(f"loss = {loss.item():.4f}")
# loss = 2.1810  (with torch.manual_seed(0))

Run on the bad labels [3, 7, 10, 1], the check_labels function above raises ValueError: Label 10 is out of range [0, 10). Did you accidentally include num_classes as a label? before any kernel runs, so the process stays usable.

1-indexed labels and ignore values

A common version is an off-by-one. Datasets that are 1-indexed (1 through N) instead of 0-indexed (0 through N-1) will pass label N when the model has N output classes, triggering the assert. Subtract 1 from every label when loading data from such a dataset:

labels = labels - 1   # convert 1-indexed to 0-indexed

Also watch for datasets that use a special "background" or "ignore" class encoded as 255 in segmentation tasks. Pass that value as ignore_index=255 (tested: labels [3, 255, 9, 1] gave a normal loss of 2.0084) to nn.CrossEntropyLoss rather than leaving it as a real label:

criterion = nn.CrossEntropyLoss(ignore_index=255)

Fix 2: Embedding Index Out of Range

nn.Embedding(vocab_size, embedding_dim) creates a lookup table with indices 0 through vocab_size - 1. Passing an index that equals or exceeds vocab_size triggers the same device-side assert.

What it looks like

import torch
import torch.nn as nn

vocab_size = 1000
embedding = nn.Embedding(vocab_size, 64).cuda()

# Token IDs in a tokenized sentence; one ID exceeds the vocabulary
token_ids = torch.tensor([[12, 45, 1000, 7]]).cuda()
# 1000 is out of range: valid indices are 0..999

# On GPU: RuntimeError: CUDA error: device-side assert triggered
# On CPU: IndexError: index out of range in self
out = embedding(token_ids)

Catching it before it hits the kernel

def check_embedding_indices(token_ids, vocab_size):
    if token_ids.min() < 0:
        raise ValueError(
            f"Embedding index contains negative value: {token_ids.min().item()}"
        )
    if token_ids.max() >= vocab_size:
        bad_ids = token_ids[token_ids >= vocab_size].unique().tolist()
        raise ValueError(
            f"Embedding index out of range. "
            f"vocab_size={vocab_size}, offending indices: {bad_ids}"
        )
    print(f"Embedding indices OK: min={token_ids.min().item()}, "
          f"max={token_ids.max().item()}, vocab_size={vocab_size}")

check_embedding_indices(token_ids, vocab_size=1000)

Fixed code

import torch
import torch.nn as nn

vocab_size = 1000
embedding = nn.Embedding(vocab_size, 64).cuda()

# All indices in [0, vocab_size)
token_ids = torch.tensor([[12, 45, 999, 7]]).cuda()

out = embedding(token_ids)
print(f"Embedding output shape: {out.shape}")

This prints Embedding output shape: torch.Size([1, 4, 64]). On the bad input, check_embedding_indices reports the offending IDs ([1000]) instead of crashing the CUDA context. The GPU version of this bug prints one Indexing.cu:1231: indexSelectSmallIndex ... Assertion `srcIndex < srcSelectDimSize` failed. line per GPU thread (64 in my run), so look for srcIndex in a long log.

Common sources of out-of-range embedding indices

  • Tokenizer vocabulary mismatch. The tokenizer has more tokens than the model's embedding table, so any token past the table size (often a rare one that only shows up in the test set) triggers the assert.
  • Special tokens not added to the model. You added [PAD], [CLS], [SEP] to the tokenizer but did not call model.resize_token_embeddings(len(tokenizer)).
  • Data pre-processing bug. An integer encoding step maps an unknown category to an ID equal to the number of categories rather than a reserved UNK index inside the range.
# After adding special tokens to a HuggingFace tokenizer:
tokenizer.add_special_tokens({"pad_token": "[PAD]"})
model.resize_token_embeddings(len(tokenizer))  # expand embedding table

Fix 3: NaN or Inf in Inputs (a different failure)

NaN or Inf logits are often listed as a cause of this assert. In my test they were not: cross-entropy on logits containing inf and nan returned a loss of nan on both GPU and CPU, with no assert. The real symptom of NaN is a loss that turns into nan and stays there. It is still worth catching early, because a NaN that reaches an index computation (for example a value cast to an integer and used as a label or token ID) can become an out-of-range index and then trigger the assert.

What it looks like

import torch
import torch.nn as nn

# Simulate logits that have gone to infinity (e.g., after gradient explosion)
logits = torch.tensor([[1.0, float("inf"), -2.0],
                       [0.5, 0.1,          float("nan")]]).cuda()

labels = torch.tensor([1, 2]).cuda()
criterion = nn.CrossEntropyLoss()

# No assert on torch 2.4.1: prints loss nan
loss = criterion(logits, labels)
print("loss", loss.item())

How to check

def check_finite(tensor, name="tensor"):
    if torch.isnan(tensor).any():
        nan_count = torch.isnan(tensor).sum().item()
        raise ValueError(f"{name} contains {nan_count} NaN value(s)")
    if torch.isinf(tensor).any():
        inf_count = torch.isinf(tensor).sum().item()
        raise ValueError(f"{name} contains {inf_count} Inf value(s)")
    print(f"{name}: finite, min={tensor.min().item():.4f}, "
          f"max={tensor.max().item():.4f}")

check_finite(logits, "logits")
# ValueError: logits contains 1 NaN value(s)  (for the tensor above)

Add these checks immediately before the loss call when debugging. You can also register a forward hook on any module to check all its outputs automatically during a diagnostic run. Fed an all-NaN input, it raised RuntimeError: NaN/Inf detected in output of Linear:

def nan_hook(module, inputs, output):
    if isinstance(output, torch.Tensor):
        if torch.isnan(output).any() or torch.isinf(output).any():
            raise RuntimeError(
                f"NaN/Inf detected in output of {module.__class__.__name__}"
            )

# Register on every submodule
for name, module in model.named_modules():
    module.register_forward_hook(nan_hook)

If the loss does go to nan, the usual suspects are a learning rate that is too high (clip with torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)), a division or torch.log without an epsilon, or float16 overflow above 65,504 under AMP. For AMP, use torch.amp.GradScaler("cuda") and watch scaler.get_scale(): a scale that keeps shrinking means overflowing steps are being skipped. This loop ran on the GTX 1070 and printed loss=1.5897, amp_scale=65536.0 on its first step with a small test model:

import torch

# torch.cuda.amp.autocast / GradScaler still work on 2.4.1 but print a
# FutureWarning; this is the current API
scaler = torch.amp.GradScaler("cuda")

for batch in train_loader:
    optimizer.zero_grad()
    with torch.autocast("cuda", dtype=torch.float16):
        logits = model(batch["input_ids"].cuda())
        loss = criterion(logits, batch["labels"].cuda())

    scaler.scale(loss).backward()
    # Clip gradients BEFORE the optimizer step
    scaler.unscale_(optimizer)
    torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
    scaler.step(optimizer)
    scaler.update()

    print(f"loss={loss.item():.4f}, amp_scale={scaler.get_scale():.1f}")

Fix 4: Wrong Label dtype (a different error)

nn.CrossEntropyLoss and other classification losses require labels to be torch.long (64-bit integer) when you pass class indices. Float labels happen when you load them from a NumPy float array or a pandas column without casting. This does not trigger the device-side assert. PyTorch rejects the dtype before the kernel runs, with a normal error that leaves the GPU usable.

Broken code

import torch
import torch.nn as nn
import numpy as np

num_classes = 5
model = nn.Linear(32, num_classes).cuda()
criterion = nn.CrossEntropyLoss()

logits = model(torch.randn(8, 32).cuda())

# Labels come from a numpy float array (common when reading a CSV)
raw_labels = np.array([0.0, 2.0, 4.0, 1.0, 3.0, 0.0, 2.0, 1.0])
labels = torch.tensor(raw_labels).cuda()   # dtype = torch.float64

# On GPU: RuntimeError: "nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Double'
# On CPU: RuntimeError: expected scalar type Long but found Double
loss = criterion(logits, labels)

A NumPy float array becomes torch.float64, hence 'Double'. With torch.float32 labels the messages end in 'Float' instead.

How to check

def check_label_dtype(labels):
    if labels.dtype != torch.long:
        raise TypeError(
            f"Labels have dtype {labels.dtype}. "
            f"CrossEntropyLoss requires torch.long (torch.int64). "
            f"Fix: labels = labels.long()"
        )
    print(f"Label dtype OK: {labels.dtype}")

check_label_dtype(labels)

Fixed code

import torch
import torch.nn as nn
import numpy as np

num_classes = 5
model = nn.Linear(32, num_classes).cuda()
criterion = nn.CrossEntropyLoss()

logits = model(torch.randn(8, 32).cuda())

raw_labels = np.array([0.0, 2.0, 4.0, 1.0, 3.0, 0.0, 2.0, 1.0])
# Cast explicitly to long before creating the tensor
labels = torch.tensor(raw_labels, dtype=torch.long).cuda()
# Or, if you already have a float tensor: labels = labels.long()

loss = criterion(logits, labels)
print(f"loss = {loss.item():.4f}")   # loss = 1.4583 in my run

A similar dtype issue appears with nn.BCELoss for binary classification: it expects labels to be torch.float32, not torch.long. Passing long labels to nn.BCEWithLogitsLoss raised RuntimeError: result type Float can't be cast to the desired output type Long on both GPU and CPU. If you switch between multi-class and binary setups, double-check both the loss function and the label dtype.

# Multi-class: nn.CrossEntropyLoss, labels must be torch.long
# Binary:      nn.BCELoss / nn.BCEWithLogitsLoss, labels must be torch.float

# Binary example
bce = nn.BCEWithLogitsLoss()
logits_binary = model_binary(x).squeeze(1)      # shape [B]
labels_binary  = labels_01.float()              # cast to float32
loss = bce(logits_binary, labels_binary)

Diagnostic Checklist

When you hit RuntimeError: CUDA error: device-side assert triggered, work through this list in order:

# Step What to look for
0 Read the Assertion ... failed line on stderr, then run with CUDA_LAUNCH_BLOCKING=1 or on CPU n_classes means labels, srcIndex means an indexing op such as nn.Embedding
1 Check label range labels.min() >= 0 and labels.max() < num_classes, ignoring ignore_index
2 Check embedding indices token_ids.max() < vocab_size
3 Check for NaN / Inf (causes a nan loss, not this assert) torch.isnan(x).any() and torch.isinf(x).any()
4 Check label dtype (causes a dtype error, not this assert) labels.dtype == torch.long for CrossEntropyLoss
5 Restart the process after fixing The CUDA context stays broken after the assert. For tensors on different devices, a separate error, see the PyTorch device mismatch fix

Run these checks as assertions at the top of your training loop during debugging.

DEBUG = True  # set False for production

def debug_check_batch(inputs, labels, logits, num_classes):
    if not DEBUG:
        return
    assert not torch.isnan(inputs).any(), "NaN in inputs"
    assert not torch.isinf(inputs).any(), "Inf in inputs"
    assert labels.dtype == torch.long, f"Expected torch.long, got {labels.dtype}"
    assert labels.min() >= 0, f"Negative label: {labels.min().item()}"
    assert labels.max() < num_classes, (
        f"Label {labels.max().item()} out of range [0, {num_classes})"
    )
    assert not torch.isnan(logits).any(), "NaN in logits"
    assert not torch.isinf(logits).any(), "Inf in logits"

# Inside training loop:
for batch in train_loader:
    inputs, labels = batch
    inputs  = inputs.cuda()
    labels  = labels.cuda()
    logits  = model(inputs)
    debug_check_batch(inputs, labels, logits, num_classes=NUM_CLASSES)
    loss = criterion(logits, labels)
    loss.backward()
    optimizer.step()
    optimizer.zero_grad()

Short Version

The assert comes from a bounds check in your data, found late because the GPU runs asynchronously. The Assertion ... failed line on stderr, CUDA_LAUNCH_BLOCKING=1, or a CPU run each get you from the misleading traceback to the real call.

In every run for this article the cause was an index out of bounds: a class label outside [0, num_classes) or a token ID past the embedding table. NaN inputs and float labels are real bugs too, but they show up as a nan loss or a dtype error rather than this assert. Put the corresponding assertion from the checklist above at the top of your training loop while you debug, confirm it actually catches the bad batch, then go fix wherever the data pipeline produced it in the first place, and restart the process before you rerun.