PyTorch RuntimeError: Expected All Tensors to Be on the Same Device
- Read the end of the message.
(when checking argument for argument mat1 in method wrapper_CUDA_addmm)means the input to aLinearlayer is on the wrong device;argument target ... nll_loss_forwardmeans your labels;argument index ... index_selectmeans the IDs passed to anEmbedding. Elementwise ops (+,*,torch.where,MSELoss) print no hint, so check the last frame of the traceback instead. - Pick one device,
device = torch.device("cuda" if torch.cuda.is_available() else "cpu"), and call.to(device)on the model once and on every input and label every batch. - Inside
forward(), create tensors withdevice=x.device. Store constant tensors withregister_bufferand sub-layers innn.ModuleList, so thatmodel.to()actually moves them. - If the error comes from
adam.py, the optimizer holds state from before you moved the model. Build the optimizer aftermodel.to(device), or reload itsstate_dict.
This is the full message PyTorch 2.4.1 prints when a model on the GPU gets a CPU input:
RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument mat1 in method wrapper_CUDA_addmm)
model.cuda() looks like it should cover everything, but it only moves the model's registered parameters and buffers. Input tensors, label tensors, anything created inside forward(), plain tensor attributes and the optimizer's state all stay where they were, usually on the CPU. The device order in the message can flip (cpu and cuda:0 appeared in two of the runs below), so don't read anything into which device is named first.
Reproduced on Real Hardware
Every row below ran on the GTX 1070 in one script with small random tensors. "same device" is short for the full Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu!, and the text in parentheses is what PyTorch appended.
| What I ran | Error or output | Fix that worked |
|---|---|---|
Linear on GPU, input on CPU | same device (when checking argument for argument mat1 in method wrapper_CUDA_addmm) | x.to(device): output on cuda:0 |
CrossEntropyLoss, labels on CPU | same device (when checking argument for argument target in method wrapper_CUDA_nll_loss_forward) | labels.to(device): loss 1.0460 |
MSELoss, target on CPU | same device, no argument hint, last frame mse_loss | Move the target |
torch.zeros(...) inside forward(), added to the output | same device, no argument hint | device=x.device or torch.zeros_like(x) |
torch.arange positions fed to a GPU Embedding | same device (when checking argument for argument index in method wrapper_CUDA__index_select) | torch.arange(n, device=x.device) |
Plain tensor attribute self.scale = torch.ones(4), then model.to("cuda") | same device, no argument hint | self.register_buffer("scale", torch.ones(4)) |
Layers stored in a Python list, then model.to("cuda") | same device, but reported as cpu and cuda:0, (... argument mat1 in method wrapper_CUDA_addmm) | nn.ModuleList |
| Adam stepped on CPU, then model moved to GPU and stepped again | same device, raised in _multi_tensor_adam (_single_tensor_adam with foreach=False) | New optimizer after the move, or reload its state_dict |
gather / index_select on a GPU tensor with a CPU index | same device (when checking argument for argument index in method wrapper_CUDA_gather) or wrapper_CUDA__index_select | index.to(src.device) |
| CPU tensor indexed with a CUDA index tensor | indices should be either on cpu or on the same device as the indexed tensor (cpu) | Index with a CPU tensor |
| GPU tensor indexed with a CPU index tensor | No error, result on cuda:0 | Not needed |
GPU tensor times a 0-dim CPU tensor torch.tensor(2.0) | No error, result on cuda:0 | Not needed |
torch.cat, @, + and torch.where mixing devices | same device (argument tensors ... wrapper_CUDA_cat, argument mat2 ... wrapper_CUDA_mm, no hint for the last two) | Move one operand |
[t1, t2].to("cuda") | AttributeError: 'list' object has no attribute 'to' (a different error) | [t.to(device) for t in batch] |
DataParallel (1 GPU) with a CPU input | No error, output on cuda:0 | Not needed |
DataParallel, one sub-layer moved back to CPU after wrapping | module must have its parameters and buffers on device cuda:0 (device_ids[0]) but found one of them on device: cpu | Move the whole wrapped model with .to(device) |
torch.load of a GPU checkpoint, no map_location, into a CPU model | No error: tensors load on cuda:0, load_state_dict copies them onto the CPU params | Nothing to fix for this path |
| Same checkpoint tensor used directly with a CPU tensor | same device (when checking argument for argument mat2 in method wrapper_CUDA_mm) | torch.load(path, map_location=device) |
Older guides (including an earlier version of this one) quote a second wording, Tensors must be on the same device to be operated on. PyTorch 2.4.1 did not print it in any of these runs, so search for the Expected all tensors text instead.
Cause 1: Inputs Left on CPU While Model Is on CUDA
import torch
import torch.nn as nn
class SimpleNet(nn.Module):
def __init__(self):
super().__init__()
self.fc = nn.Linear(10, 2)
def forward(self, x):
return self.fc(x)
model = SimpleNet()
model.cuda() # model weights are now on cuda:0
# Input tensor is still on CPU (default)
x = torch.randn(32, 10) # cpu
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu! (when checking argument for argument mat1
# in method wrapper_CUDA_addmm)
output = model(x)
model.cuda() moves the model's registered parameters and buffers to the GPU. Every tensor you create with torch.randn() or torch.zeros(), or get from a DataLoader, starts on the CPU. nn.Linear runs addmm, and mat1 in the message is its first matrix argument: your input.
Define a device variable once and move every tensor explicitly with .to(device). Avoid hardcoding .cuda(). On a machine where no GPU is visible it fails with RuntimeError: No CUDA GPUs are available (tested with CUDA_VISIBLE_DEVICES="").
import torch
import torch.nn as nn
from torch.utils.data import DataLoader, TensorDataset
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using device: {device}")
class SimpleNet(nn.Module):
def __init__(self):
super().__init__()
self.fc = nn.Linear(10, 2)
def forward(self, x):
return self.fc(x)
model = SimpleNet().to(device) # move model
X = torch.randn(200, 10)
y = torch.randint(0, 2, (200,))
dataset = TensorDataset(X, y)
loader = DataLoader(dataset, batch_size=32)
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
for inputs, labels in loader:
inputs = inputs.to(device) # move inputs
labels = labels.to(device) # move labels
optimizer.zero_grad()
outputs = model(inputs)
loss = criterion(outputs, labels)
loss.backward()
optimizer.step()
When a tensor is already on the target device, .to(device) returns the same object without copying (t.to(device) is t printed True), so it is safe to call on every batch.
Cause 2: Label Tensor on CPU During Loss Computation
import torch
import torch.nn as nn
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = nn.Linear(10, 3).to(device)
criterion = nn.CrossEntropyLoss()
inputs = torch.randn(16, 10).to(device) # correctly moved
labels = torch.randint(0, 3, (16,)) # forgotten, still on CPU
outputs = model(inputs) # outputs is on cuda:0
loss = criterion(outputs, labels)
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu! (when checking argument for argument
# target in method wrapper_CUDA_nll_loss_forward)
Here the forward pass succeeds and the loss call fails. With CrossEntropyLoss the message names argument target. With MSELoss it names nothing, and the traceback's last frame is mse_loss in torch/nn/functional.py. Move both tensors:
import torch
import torch.nn as nn
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = nn.Linear(10, 3).to(device)
criterion = nn.CrossEntropyLoss()
inputs = torch.randn(16, 10)
labels = torch.randint(0, 3, (16,))
inputs, labels = inputs.to(device), labels.to(device)
outputs = model(inputs)
loss = criterion(outputs, labels) # both on the same device, works
print(f"Loss: {loss.item():.4f}")
Batches that are dicts or nested lists are where a tensor usually gets missed. A list has no .to() method ([t1, t2].to("cuda") raises AttributeError: 'list' object has no attribute 'to'), so walk the structure:
def to_device(batch, device):
"""Recursively move a batch (tensor, list, dict, tuple) to device."""
if isinstance(batch, torch.Tensor):
return batch.to(device)
if isinstance(batch, (list, tuple)):
return type(batch)(to_device(b, device) for b in batch)
if isinstance(batch, dict):
return {k: to_device(v, device) for k, v in batch.items()}
return batch
for batch in loader:
batch = to_device(batch, device)
inputs, labels = batch
Cause 3: Tensors Created Inside forward() Default to CPU
import torch
import torch.nn as nn
class AttentionModel(nn.Module):
def __init__(self, hidden_size):
super().__init__()
self.hidden_size = hidden_size
self.linear = nn.Linear(hidden_size, hidden_size)
def forward(self, x):
batch_size = x.size(0)
# BUG: without device=, torch.zeros creates a CPU tensor
# even when x is on cuda:0
mask = torch.zeros(batch_size, self.hidden_size)
# RuntimeError: Expected all tensors to be on the same device, but
# found at least two devices, cuda:0 and cpu!
return self.linear(x) + mask
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = AttentionModel(64).to(device)
x = torch.randn(8, 64).to(device)
output = model(x) # crashes on GPU
Factory functions like torch.zeros(), torch.ones(), torch.randn() and torch.arange() use the default device (the CPU unless you changed it) no matter where the model lives. Position IDs hit the same problem: self.pos_emb(torch.arange(seq_len)) on a GPU model fails with (when checking argument for argument index in method wrapper_CUDA__index_select), because nn.Embedding calls index_select.
Pass device=x.device so the new tensor follows the input. That keeps working on CPU-only runs too:
import torch
import torch.nn as nn
class AttentionModel(nn.Module):
def __init__(self, hidden_size):
super().__init__()
self.hidden_size = hidden_size
self.linear = nn.Linear(hidden_size, hidden_size)
def forward(self, x):
batch_size = x.size(0)
# Preferred: pass device= explicitly, taken from the input
mask = torch.zeros(batch_size, self.hidden_size, device=x.device)
# Also works, with an extra copy:
# mask = torch.zeros(batch_size, self.hidden_size).to(x.device)
# When the shape matches x, zeros_like copies device and dtype:
# mask = torch.zeros_like(x)
return self.linear(x) + mask
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = AttentionModel(64).to(device)
x = torch.randn(8, 64).to(device)
output = model(x) # works
print(output.device) # cuda:0
Cause 4: Tensors and Layers That model.to() Can't See
model.to(device) only moves parameters, registered buffers and submodules. A tensor saved as a plain attribute, or layers kept in a regular Python list, stay on the CPU:
import torch
import torch.nn as nn
class Scaled(nn.Module):
def __init__(self):
super().__init__()
self.fc = nn.Linear(4, 4)
self.scale = torch.ones(4) # BUG: plain attribute, not moved
self.blocks = [nn.Linear(4, 4)] # BUG: plain list, not moved
def forward(self, x):
x = self.fc(x) * self.scale # fails here first
for block in self.blocks:
x = block(x)
return x
model = Scaled().to("cuda")
model(torch.randn(2, 4, device="cuda"))
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu!
Register the tensor as a buffer and wrap the layers in nn.ModuleList. Both then move with the model and are included in state_dict() (pass persistent=False to register_buffer to leave the buffer out of checkpoints):
class Scaled(nn.Module):
def __init__(self):
super().__init__()
self.fc = nn.Linear(4, 4)
self.register_buffer("scale", torch.ones(4))
self.blocks = nn.ModuleList([nn.Linear(4, 4)])
def forward(self, x):
x = self.fc(x) * self.scale
for block in self.blocks:
x = block(x)
return x
model = Scaled().to("cuda")
print(model(torch.randn(2, 4, device="cuda")).device) # cuda:0
Cause 5: Optimizer State Created Before the Model Was Moved
This one shows up when resuming training: you build the model and optimizer on the CPU, take at least one step (or load an optimizer checkpoint), and then move the model to the GPU. Adam's exp_avg and exp_avg_sq tensors stay on the CPU, and the next optimizer.step() fails with the same message. The traceback ends in torch/optim/adam.py, in _multi_tensor_adam (the default on CUDA) or _single_tensor_adam with foreach=False.
import torch
import torch.nn as nn
model = nn.Linear(4, 2)
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
model(torch.randn(3, 4)).sum().backward()
optimizer.step() # creates CPU state
optimizer.zero_grad()
model.to("cuda") # params move, optimizer state does not
model(torch.randn(3, 4, device="cuda")).sum().backward()
optimizer.step()
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu!
The simplest fix is ordering: call model.to(device) before you create the optimizer. If you already have state, rebuild the optimizer after the move and load the old state into it, since load_state_dict casts the state to each parameter's device:
state = optimizer.state_dict()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3) # after model.to()
optimizer.load_state_dict(state)
model(torch.randn(3, 4, device="cuda")).sum().backward()
optimizer.step() # works: exp_avg and exp_avg_sq are now on cuda:0
Adam's step counter stays a CPU tensor after this. That is expected: PyTorch keeps it on the CPU unless you pass capturable=True. Plain SGD (no momentum) has no state, so moving the model after creating it caused no error.
Cause 6: Index Tensors on the Wrong Device
Indexing ops check devices too, though not all in the same way:
import torch
src = torch.randn(3, 5, device="cuda")
idx = torch.zeros(3, 1, dtype=torch.long) # CPU
src.gather(1, idx)
# RuntimeError: Expected all tensors to be on the same device, but found at
# least two devices, cuda:0 and cpu! (when checking argument for argument
# index in method wrapper_CUDA_gather)
src.gather(1, idx.to(src.device)) # works
src[torch.tensor([0, 1])] # also works: CPU index on a GPU tensor
cpu_src = torch.randn(5)
cpu_src[torch.tensor([0, 1], device="cuda")]
# RuntimeError: indices should be either on cpu or on the same device as the
# indexed tensor (cpu)
index_select and nn.Embedding with CPU IDs fail the way gather does, naming wrapper_CUDA__index_select. Keep index tensors on the same device as the tensor being indexed and you won't have to remember which ops are lenient.
Checkpoints are a related trap. torch.load without map_location puts saved GPU tensors back on cuda:0. Passing that state dict to load_state_dict on a CPU model worked fine in testing, because the values are copied into the existing parameters. The error only appeared when a loaded tensor was used directly in math with a CPU tensor (argument mat2 in method wrapper_CUDA_mm). torch.load(path, map_location=device) avoids that.
Cause 7: DataParallel and Multiple GPUs
nn.DataParallel is more forgiving than older guides suggest. On the single-GPU test machine a DataParallel model accepted a CPU input and returned output on cuda:0, since it scatters the batch itself. What it does check is the model: every parameter and buffer must be on device_ids[0]. Moving part of the wrapped model elsewhere produced a different error:
import torch
import torch.nn as nn
model = nn.DataParallel(nn.Sequential(nn.Linear(10, 2), nn.Linear(2, 2)))
model.module[1].cpu() # e.g. a helper that moves a submodule
model(torch.randn(4, 10, device="cuda"))
# RuntimeError: module must have its parameters and buffers on device cuda:0
# (device_ids[0]) but found one of them on device: cpu
Call model.to("cuda:0") (or whatever device_ids[0] is) on the whole wrapped model after any change like that. With several GPUs, the "same device" error can also read cuda:0 and cuda:1 when tensors from two GPUs meet. That needs a second GPU and was not reproduced here.
For new multi-GPU code, torch.nn.parallel.DistributedDataParallel (DDP) is the better choice. Each process owns exactly one GPU, so device placement is explicit. This snippet ran as a single process on the test machine; the commented lines are what a real torchrun job adds:
import os
import torch
import torch.nn as nn
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP
# torchrun sets LOCAL_RANK for each worker process
local_rank = int(os.environ.get("LOCAL_RANK", 0))
torch.cuda.set_device(local_rank)
device = torch.device(f"cuda:{local_rank}")
model = nn.Linear(10, 2).to(device)
# dist.init_process_group would be called before this in real usage
# model = DDP(model, device_ids=[local_rank])
x = torch.randn(16, 10).to(device) # matches this process's GPU
output = model(x)
Find Which Tensor Is on the Wrong Device
Start with the hint at the end of the message and the last frame of the traceback. If the model is large and that isn't enough, list where everything lives:
import torch
import torch.nn as nn
def audit_model_devices(model):
print("=== Model parameter devices ===")
for name, param in model.named_parameters():
print(f" {name}: {param.device} (dtype={param.dtype})")
print("=== Model buffer devices ===")
for name, buf in model.named_buffers():
print(f" {name}: {buf.device} (dtype={buf.dtype})")
model = nn.Sequential(
nn.Linear(10, 20),
nn.ReLU(),
nn.Linear(20, 2)
)
# Move only part of the model accidentally
model[0] = model[0].cuda()
audit_model_devices(model)
# === Model parameter devices ===
# 0.weight: cuda:0 (dtype=torch.float32)
# 0.bias: cuda:0 (dtype=torch.float32)
# 2.weight: cpu (dtype=torch.float32)
# 2.bias: cpu (dtype=torch.float32)
# === Model buffer devices ===
Running that half-moved model on a GPU input fails at the second Linear, with the devices reported as cpu and cuda:0. Plain tensor attributes from Cause 4 never show up in named_parameters() or named_buffers(). If a tensor your forward() uses is missing from this listing, model.to() won't move it.
import torch
def audit_batch_devices(batch):
if isinstance(batch, torch.Tensor):
print(f"tensor: shape={tuple(batch.shape)}, device={batch.device}, dtype={batch.dtype}")
elif isinstance(batch, (list, tuple)):
for i, item in enumerate(batch):
print(f" [{i}]", end=" ")
audit_batch_devices(item)
elif isinstance(batch, dict):
for k, v in batch.items():
print(f" '{k}':", end=" ")
audit_batch_devices(v)
# Example: NLP batch with mixed devices
batch = {
"input_ids": torch.randint(0, 1000, (8, 128)).cuda(),
"attention_mask": torch.ones(8, 128), # still on CPU
"labels": torch.randint(0, 2, (8,)), # still on CPU
}
audit_batch_devices(batch)
# 'input_ids': tensor: shape=(8, 128), device=cuda:0, dtype=torch.int64
# 'attention_mask': tensor: shape=(8, 128), device=cpu, dtype=torch.float32
# 'labels': tensor: shape=(8,), device=cpu, dtype=torch.int64
If you only keep one habit from this page, make it creating the optimizer
after model.to(device) and moving the batch at the top of the
loop. Those two cover the input, label and optimizer rows in the table above;
the rest come from tensors created or stored inside the model.