Skip to content

PyTorch RuntimeError: Expected All Tensors to Be on the Same Device

Tested with: PyTorch 2.4.1 (CUDA 12.1), Python 3.12, NVIDIA GeForce GTX 1070 8 GB, driver 580, Ubuntu 24.04. Last run 2026-09-27.

TL;DR:
  1. Read the end of the message. (when checking argument for argument mat1 in method wrapper_CUDA_addmm) means the input to a Linear layer is on the wrong device; argument target ... nll_loss_forward means your labels; argument index ... index_select means the IDs passed to an Embedding. Elementwise ops (+, *, torch.where, MSELoss) print no hint, so check the last frame of the traceback instead.
  2. Pick one device, device = torch.device("cuda" if torch.cuda.is_available() else "cpu"), and call .to(device) on the model once and on every input and label every batch.
  3. Inside forward(), create tensors with device=x.device. Store constant tensors with register_buffer and sub-layers in nn.ModuleList, so that model.to() actually moves them.
  4. If the error comes from adam.py, the optimizer holds state from before you moved the model. Build the optimizer after model.to(device), or reload its state_dict.

This is the full message PyTorch 2.4.1 prints when a model on the GPU gets a CPU input:

RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument mat1 in method wrapper_CUDA_addmm)

model.cuda() looks like it should cover everything, but it only moves the model's registered parameters and buffers. Input tensors, label tensors, anything created inside forward(), plain tensor attributes and the optimizer's state all stay where they were, usually on the CPU. The device order in the message can flip (cpu and cuda:0 appeared in two of the runs below), so don't read anything into which device is named first.


Reproduced on Real Hardware

Every row below ran on the GTX 1070 in one script with small random tensors. "same device" is short for the full Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!, and the text in parentheses is what PyTorch appended.

What I ranError or outputFix that worked
Linear on GPU, input on CPUsame device (when checking argument for argument mat1 in method wrapper_CUDA_addmm)x.to(device): output on cuda:0
CrossEntropyLoss, labels on CPUsame device (when checking argument for argument target in method wrapper_CUDA_nll_loss_forward)labels.to(device): loss 1.0460
MSELoss, target on CPUsame device, no argument hint, last frame mse_lossMove the target
torch.zeros(...) inside forward(), added to the outputsame device, no argument hintdevice=x.device or torch.zeros_like(x)
torch.arange positions fed to a GPU Embeddingsame device (when checking argument for argument index in method wrapper_CUDA__index_select)torch.arange(n, device=x.device)
Plain tensor attribute self.scale = torch.ones(4), then model.to("cuda")same device, no argument hintself.register_buffer("scale", torch.ones(4))
Layers stored in a Python list, then model.to("cuda")same device, but reported as cpu and cuda:0, (... argument mat1 in method wrapper_CUDA_addmm)nn.ModuleList
Adam stepped on CPU, then model moved to GPU and stepped againsame device, raised in _multi_tensor_adam (_single_tensor_adam with foreach=False)New optimizer after the move, or reload its state_dict
gather / index_select on a GPU tensor with a CPU indexsame device (when checking argument for argument index in method wrapper_CUDA_gather) or wrapper_CUDA__index_selectindex.to(src.device)
CPU tensor indexed with a CUDA index tensorindices should be either on cpu or on the same device as the indexed tensor (cpu)Index with a CPU tensor
GPU tensor indexed with a CPU index tensorNo error, result on cuda:0Not needed
GPU tensor times a 0-dim CPU tensor torch.tensor(2.0)No error, result on cuda:0Not needed
torch.cat, @, + and torch.where mixing devicessame device (argument tensors ... wrapper_CUDA_cat, argument mat2 ... wrapper_CUDA_mm, no hint for the last two)Move one operand
[t1, t2].to("cuda")AttributeError: 'list' object has no attribute 'to' (a different error)[t.to(device) for t in batch]
DataParallel (1 GPU) with a CPU inputNo error, output on cuda:0Not needed
DataParallel, one sub-layer moved back to CPU after wrappingmodule must have its parameters and buffers on device cuda:0 (device_ids[0]) but found one of them on device: cpuMove the whole wrapped model with .to(device)
torch.load of a GPU checkpoint, no map_location, into a CPU modelNo error: tensors load on cuda:0, load_state_dict copies them onto the CPU paramsNothing to fix for this path
Same checkpoint tensor used directly with a CPU tensorsame device (when checking argument for argument mat2 in method wrapper_CUDA_mm)torch.load(path, map_location=device)

Older guides (including an earlier version of this one) quote a second wording, Tensors must be on the same device to be operated on. PyTorch 2.4.1 did not print it in any of these runs, so search for the Expected all tensors text instead.


Cause 1: Inputs Left on CPU While Model Is on CUDA

import torch
import torch.nn as nn

class SimpleNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc = nn.Linear(10, 2)

    def forward(self, x):
        return self.fc(x)

model = SimpleNet()
model.cuda()  # model weights are now on cuda:0

# Input tensor is still on CPU (default)
x = torch.randn(32, 10)  # cpu

# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu! (when checking argument for argument mat1
# in method wrapper_CUDA_addmm)
output = model(x)

model.cuda() moves the model's registered parameters and buffers to the GPU. Every tensor you create with torch.randn() or torch.zeros(), or get from a DataLoader, starts on the CPU. nn.Linear runs addmm, and mat1 in the message is its first matrix argument: your input.

Define a device variable once and move every tensor explicitly with .to(device). Avoid hardcoding .cuda(). On a machine where no GPU is visible it fails with RuntimeError: No CUDA GPUs are available (tested with CUDA_VISIBLE_DEVICES="").

import torch
import torch.nn as nn
from torch.utils.data import DataLoader, TensorDataset

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using device: {device}")

class SimpleNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc = nn.Linear(10, 2)

    def forward(self, x):
        return self.fc(x)

model = SimpleNet().to(device)  # move model

X = torch.randn(200, 10)
y = torch.randint(0, 2, (200,))
dataset = TensorDataset(X, y)
loader = DataLoader(dataset, batch_size=32)

criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

for inputs, labels in loader:
    inputs = inputs.to(device)   # move inputs
    labels = labels.to(device)   # move labels

    optimizer.zero_grad()
    outputs = model(inputs)
    loss = criterion(outputs, labels)
    loss.backward()
    optimizer.step()

When a tensor is already on the target device, .to(device) returns the same object without copying (t.to(device) is t printed True), so it is safe to call on every batch.


Cause 2: Label Tensor on CPU During Loss Computation

import torch
import torch.nn as nn

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = nn.Linear(10, 3).to(device)
criterion = nn.CrossEntropyLoss()

inputs = torch.randn(16, 10).to(device)  # correctly moved
labels = torch.randint(0, 3, (16,))      # forgotten, still on CPU

outputs = model(inputs)  # outputs is on cuda:0
loss = criterion(outputs, labels)
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu! (when checking argument for argument
# target in method wrapper_CUDA_nll_loss_forward)

Here the forward pass succeeds and the loss call fails. With CrossEntropyLoss the message names argument target. With MSELoss it names nothing, and the traceback's last frame is mse_loss in torch/nn/functional.py. Move both tensors:

import torch
import torch.nn as nn

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = nn.Linear(10, 3).to(device)
criterion = nn.CrossEntropyLoss()

inputs = torch.randn(16, 10)
labels = torch.randint(0, 3, (16,))

inputs, labels = inputs.to(device), labels.to(device)

outputs = model(inputs)
loss = criterion(outputs, labels)  # both on the same device, works
print(f"Loss: {loss.item():.4f}")

Batches that are dicts or nested lists are where a tensor usually gets missed. A list has no .to() method ([t1, t2].to("cuda") raises AttributeError: 'list' object has no attribute 'to'), so walk the structure:

def to_device(batch, device):
    """Recursively move a batch (tensor, list, dict, tuple) to device."""
    if isinstance(batch, torch.Tensor):
        return batch.to(device)
    if isinstance(batch, (list, tuple)):
        return type(batch)(to_device(b, device) for b in batch)
    if isinstance(batch, dict):
        return {k: to_device(v, device) for k, v in batch.items()}
    return batch

for batch in loader:
    batch = to_device(batch, device)
    inputs, labels = batch

Cause 3: Tensors Created Inside forward() Default to CPU

import torch
import torch.nn as nn

class AttentionModel(nn.Module):
    def __init__(self, hidden_size):
        super().__init__()
        self.hidden_size = hidden_size
        self.linear = nn.Linear(hidden_size, hidden_size)

    def forward(self, x):
        batch_size = x.size(0)

        # BUG: without device=, torch.zeros creates a CPU tensor
        # even when x is on cuda:0
        mask = torch.zeros(batch_size, self.hidden_size)

        # RuntimeError: Expected all tensors to be on the same device, but
        # found at least two devices, cuda:0 and cpu!
        return self.linear(x) + mask

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = AttentionModel(64).to(device)
x = torch.randn(8, 64).to(device)
output = model(x)  # crashes on GPU

Factory functions like torch.zeros(), torch.ones(), torch.randn() and torch.arange() use the default device (the CPU unless you changed it) no matter where the model lives. Position IDs hit the same problem: self.pos_emb(torch.arange(seq_len)) on a GPU model fails with (when checking argument for argument index in method wrapper_CUDA__index_select), because nn.Embedding calls index_select.

Pass device=x.device so the new tensor follows the input. That keeps working on CPU-only runs too:

import torch
import torch.nn as nn

class AttentionModel(nn.Module):
    def __init__(self, hidden_size):
        super().__init__()
        self.hidden_size = hidden_size
        self.linear = nn.Linear(hidden_size, hidden_size)

    def forward(self, x):
        batch_size = x.size(0)

        # Preferred: pass device= explicitly, taken from the input
        mask = torch.zeros(batch_size, self.hidden_size, device=x.device)

        # Also works, with an extra copy:
        # mask = torch.zeros(batch_size, self.hidden_size).to(x.device)

        # When the shape matches x, zeros_like copies device and dtype:
        # mask = torch.zeros_like(x)

        return self.linear(x) + mask

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = AttentionModel(64).to(device)
x = torch.randn(8, 64).to(device)
output = model(x)  # works
print(output.device)  # cuda:0

Cause 4: Tensors and Layers That model.to() Can't See

model.to(device) only moves parameters, registered buffers and submodules. A tensor saved as a plain attribute, or layers kept in a regular Python list, stay on the CPU:

import torch
import torch.nn as nn

class Scaled(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc = nn.Linear(4, 4)
        self.scale = torch.ones(4)           # BUG: plain attribute, not moved
        self.blocks = [nn.Linear(4, 4)]      # BUG: plain list, not moved

    def forward(self, x):
        x = self.fc(x) * self.scale          # fails here first
        for block in self.blocks:
            x = block(x)
        return x

model = Scaled().to("cuda")
model(torch.randn(2, 4, device="cuda"))
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu!

Register the tensor as a buffer and wrap the layers in nn.ModuleList. Both then move with the model and are included in state_dict() (pass persistent=False to register_buffer to leave the buffer out of checkpoints):

class Scaled(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc = nn.Linear(4, 4)
        self.register_buffer("scale", torch.ones(4))
        self.blocks = nn.ModuleList([nn.Linear(4, 4)])

    def forward(self, x):
        x = self.fc(x) * self.scale
        for block in self.blocks:
            x = block(x)
        return x

model = Scaled().to("cuda")
print(model(torch.randn(2, 4, device="cuda")).device)  # cuda:0

Cause 5: Optimizer State Created Before the Model Was Moved

This one shows up when resuming training: you build the model and optimizer on the CPU, take at least one step (or load an optimizer checkpoint), and then move the model to the GPU. Adam's exp_avg and exp_avg_sq tensors stay on the CPU, and the next optimizer.step() fails with the same message. The traceback ends in torch/optim/adam.py, in _multi_tensor_adam (the default on CUDA) or _single_tensor_adam with foreach=False.

import torch
import torch.nn as nn

model = nn.Linear(4, 2)
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
model(torch.randn(3, 4)).sum().backward()
optimizer.step()             # creates CPU state
optimizer.zero_grad()

model.to("cuda")             # params move, optimizer state does not
model(torch.randn(3, 4, device="cuda")).sum().backward()
optimizer.step()
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu!

The simplest fix is ordering: call model.to(device) before you create the optimizer. If you already have state, rebuild the optimizer after the move and load the old state into it, since load_state_dict casts the state to each parameter's device:

state = optimizer.state_dict()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)  # after model.to()
optimizer.load_state_dict(state)
model(torch.randn(3, 4, device="cuda")).sum().backward()
optimizer.step()  # works: exp_avg and exp_avg_sq are now on cuda:0

Adam's step counter stays a CPU tensor after this. That is expected: PyTorch keeps it on the CPU unless you pass capturable=True. Plain SGD (no momentum) has no state, so moving the model after creating it caused no error.


Cause 6: Index Tensors on the Wrong Device

Indexing ops check devices too, though not all in the same way:

import torch

src = torch.randn(3, 5, device="cuda")
idx = torch.zeros(3, 1, dtype=torch.long)      # CPU

src.gather(1, idx)
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu! (when checking argument for argument
# index in method wrapper_CUDA_gather)

src.gather(1, idx.to(src.device))               # works
src[torch.tensor([0, 1])]                       # also works: CPU index on a GPU tensor

cpu_src = torch.randn(5)
cpu_src[torch.tensor([0, 1], device="cuda")]
# RuntimeError: indices should be either on cpu or on the same device as the
# indexed tensor (cpu)

index_select and nn.Embedding with CPU IDs fail the way gather does, naming wrapper_CUDA__index_select. Keep index tensors on the same device as the tensor being indexed and you won't have to remember which ops are lenient.

Checkpoints are a related trap. torch.load without map_location puts saved GPU tensors back on cuda:0. Passing that state dict to load_state_dict on a CPU model worked fine in testing, because the values are copied into the existing parameters. The error only appeared when a loaded tensor was used directly in math with a CPU tensor (argument mat2 in method wrapper_CUDA_mm). torch.load(path, map_location=device) avoids that.


Cause 7: DataParallel and Multiple GPUs

nn.DataParallel is more forgiving than older guides suggest. On the single-GPU test machine a DataParallel model accepted a CPU input and returned output on cuda:0, since it scatters the batch itself. What it does check is the model: every parameter and buffer must be on device_ids[0]. Moving part of the wrapped model elsewhere produced a different error:

import torch
import torch.nn as nn

model = nn.DataParallel(nn.Sequential(nn.Linear(10, 2), nn.Linear(2, 2)))
model.module[1].cpu()   # e.g. a helper that moves a submodule
model(torch.randn(4, 10, device="cuda"))
# RuntimeError: module must have its parameters and buffers on device cuda:0
# (device_ids[0]) but found one of them on device: cpu

Call model.to("cuda:0") (or whatever device_ids[0] is) on the whole wrapped model after any change like that. With several GPUs, the "same device" error can also read cuda:0 and cuda:1 when tensors from two GPUs meet. That needs a second GPU and was not reproduced here.

For new multi-GPU code, torch.nn.parallel.DistributedDataParallel (DDP) is the better choice. Each process owns exactly one GPU, so device placement is explicit. This snippet ran as a single process on the test machine; the commented lines are what a real torchrun job adds:

import os
import torch
import torch.nn as nn
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP

# torchrun sets LOCAL_RANK for each worker process
local_rank = int(os.environ.get("LOCAL_RANK", 0))
torch.cuda.set_device(local_rank)
device = torch.device(f"cuda:{local_rank}")

model = nn.Linear(10, 2).to(device)

# dist.init_process_group would be called before this in real usage
# model = DDP(model, device_ids=[local_rank])

x = torch.randn(16, 10).to(device)  # matches this process's GPU
output = model(x)

Find Which Tensor Is on the Wrong Device

Start with the hint at the end of the message and the last frame of the traceback. If the model is large and that isn't enough, list where everything lives:

import torch
import torch.nn as nn

def audit_model_devices(model):
    print("=== Model parameter devices ===")
    for name, param in model.named_parameters():
        print(f"  {name}: {param.device} (dtype={param.dtype})")
    print("=== Model buffer devices ===")
    for name, buf in model.named_buffers():
        print(f"  {name}: {buf.device} (dtype={buf.dtype})")

model = nn.Sequential(
    nn.Linear(10, 20),
    nn.ReLU(),
    nn.Linear(20, 2)
)

# Move only part of the model accidentally
model[0] = model[0].cuda()

audit_model_devices(model)
# === Model parameter devices ===
#   0.weight: cuda:0 (dtype=torch.float32)
#   0.bias: cuda:0 (dtype=torch.float32)
#   2.weight: cpu (dtype=torch.float32)
#   2.bias: cpu (dtype=torch.float32)
# === Model buffer devices ===

Running that half-moved model on a GPU input fails at the second Linear, with the devices reported as cpu and cuda:0. Plain tensor attributes from Cause 4 never show up in named_parameters() or named_buffers(). If a tensor your forward() uses is missing from this listing, model.to() won't move it.

import torch

def audit_batch_devices(batch):
    if isinstance(batch, torch.Tensor):
        print(f"tensor: shape={tuple(batch.shape)}, device={batch.device}, dtype={batch.dtype}")
    elif isinstance(batch, (list, tuple)):
        for i, item in enumerate(batch):
            print(f"  [{i}]", end=" ")
            audit_batch_devices(item)
    elif isinstance(batch, dict):
        for k, v in batch.items():
            print(f"  '{k}':", end=" ")
            audit_batch_devices(v)

# Example: NLP batch with mixed devices
batch = {
    "input_ids":      torch.randint(0, 1000, (8, 128)).cuda(),
    "attention_mask": torch.ones(8, 128),           # still on CPU
    "labels":         torch.randint(0, 2, (8,)),    # still on CPU
}

audit_batch_devices(batch)
#   'input_ids': tensor: shape=(8, 128), device=cuda:0, dtype=torch.int64
#   'attention_mask': tensor: shape=(8, 128), device=cpu, dtype=torch.float32
#   'labels': tensor: shape=(8,), device=cpu, dtype=torch.int64

If you only keep one habit from this page, make it creating the optimizer after model.to(device) and moving the batch at the top of the loop. Those two cover the input, label and optimizer rows in the table above; the rest come from tensors created or stored inside the model.