PyTorch RuntimeError: Expected Scalar Type Long but Found Float
- Your class labels are not
torch.int64. Cast them before the loss:loss = criterion(outputs, labels.long()), or build them astorch.tensor(y, dtype=torch.long)in yourDataset. - On the GPU the same bug prints a different message:
"nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Float'. Same fix. int32labels fail too (found Int), and so dofloat64labels from NumPy (found Double). Onlyint64works.BCELoss,BCEWithLogitsLossandMSELossgo the other way: they want float targets, andBCELosswants exactlyfloat32. For inputs,mat1 and mat2 must have the same dtypemeans afloat64NumPy array reached afloat32layer: use.float().
Training a classification model, and the loop crashes with
RuntimeError: expected scalar type Long but found Float. The cause is nearly always
the same: CrossEntropyLoss (or NLLLoss) wants integer class indices
and got floating-point values. Below is exactly what PyTorch 2.4.1 prints on CPU and on CUDA,
where the float labels usually come from, and the related dtype errors that show up once this
one is fixed.
Contents
- Reproduced on real hardware
- The exact error message (CPU and CUDA)
- Root cause: CrossEntropyLoss needs torch.long
- The one-line fix: labels.long()
- BCELoss vs CrossEntropyLoss dtype requirements
- Embedding layer: a different error for the same mistake
- Dtype table for common loss functions
- Debugging snippet: print dtypes before the forward pass
- Related fixes
Reproduced on Real Hardware
Every row ran in one script with small random tensors, once on the CPU and once on the GTX 1070. Where the two wordings differ, both are quoted.
| What I ran | Error or output (CPU / CUDA) | Fix that worked |
|---|---|---|
CrossEntropyLoss, labels float32 | CPU: expected scalar type Long but found FloatCUDA: "nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Float' | labels.long() |
Labels float64 (from NumPy) | CPU: ... but found DoubleCUDA: ... not implemented for 'Double' | labels.long() |
Labels int32 (np.int32 or .int()) | CPU: ... but found IntCUDA: ... not implemented for 'Int' | labels.long() |
F.nll_loss, float targets | Same messages as CrossEntropyLoss | target.long() |
Segmentation: logits [2, 3, 4, 4], mask float32 [2, 4, 4] | expected scalar type Long but found Float on both CPU and CUDA | mask.long() |
CrossEntropyLoss with a float one-hot target, same shape as the logits | No error: treated as class probabilities, same loss as the index labels | Not needed |
BCEWithLogitsLoss, target int64 | result type Float can't be cast to the desired output type Long (both) | target.float() |
BCELoss, target int64 / float64 | Found dtype Long but expected Float / Found dtype Double but expected Float (both) | target.float() |
BCEWithLogitsLoss, target float64 | No error | Not needed |
MSELoss, target int64 or float64 | Forward works; .backward() raises Found dtype Long but expected Float / Found dtype Double but expected Float | target.float() |
nn.Embedding, float32 indices | Expected tensor for argument #1 'indices' to have one of the following scalar types: Long, Int; but got torch.FloatTensor instead (while checking arguments for embedding) (torch.cuda.FloatTensor on GPU) | .long(); int32 also works |
nn.Linear (float32) fed torch.from_numpy of a float64 array | mat1 and mat2 must have the same dtype, but got Double and Float (both) | x.float() |
model.double(), input float32 | mat1 and mat2 must have the same dtype, but got Float and Double | Match the input to the model |
float32 @ float64 with @ or torch.mm | CPU: expected m1 and m2 to have the same dtype, but got: float != doubleCUDA: expected mat1 and mat2 ... | Cast one side |
Conv2d fed a float64 image | Input type (double) and bias type (float) should be the same | x.float() |
AMP: float16 output from autocast used in a float32 Linear outside the block | mat1 and mat2 must have the same dtype, but got Half and Float | Keep the op inside autocast, or .float() |
AMP: CrossEntropyLoss on float16 logits with int64 labels inside autocast | No error, loss is float32 | Not needed |
If you search for the CPU wording but train on a GPU, you won't find your error: the CUDA kernel
reports the unsupported dtype instead. And some wrong dtypes do not raise at all, or only raise
in backward().
1. The Exact Error Message (CPU and CUDA)
This is the full traceback PyTorch 2.4.1 prints for float labels on the CPU (site-packages path shortened, caret lines removed):
Traceback (most recent call last):
File "/tmp/scalar/repro2.py", line 7, in <module>
loss = ce(outputs, labels)
File ".../torch/nn/modules/module.py", line 1553, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File ".../torch/nn/modules/module.py", line 1562, in _call_impl
return forward_call(*args, **kwargs)
File ".../torch/nn/modules/loss.py", line 1188, in forward
return F.cross_entropy(input, target, weight=self.weight,
File ".../torch/nn/functional.py", line 3104, in cross_entropy
return torch._C._nn.cross_entropy_loss(input, target, weight, _Reduction.get_enum(reduction), ignore_index, label_smoothing)
RuntimeError: expected scalar type Long but found Float
The same code with outputs and labels on cuda goes through the same frames and ends with:
RuntimeError: "nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Float'
The last word tells you what you passed: Float is float32,
Double is float64, Int is int32. The key line
is loss = ce(outputs, labels): labels needs to be
torch.int64 (also called torch.long). Older guides quote a longer
wording, Expected object of scalar type Long but got scalar type Float for argument #2
'target', from PyTorch 1.x. I did not run a 1.x build, and 2.4.1 did not print it.
2. Root Cause: CrossEntropyLoss Expects Class Indices as torch.long
With a target of shape [N], nn.CrossEntropyLoss computes
log_softmax and then nll_loss, which uses each target value as an
index into the class dimension. The kernel is written for int64 indices only, so
float32, float64 and even int32 are rejected.
Where the wrong dtype usually comes from, each checked on torch 2.4.1 and NumPy 2.4.4:
- Labels loaded with NumPy.
np.loadtxtreturnsfloat64by default, andtorch.from_numpykeeps that dtype, givingfound Double. Dataset.__getitem__returnsnp.float32labels. The default DataLoader collate turns them into afloat32batch. Returning a plain Pythonintgivesint64.- An explicit
np.int32array. It becomes atorch.int32tensor and fails withfound Int. - True division.
torch.tensor([0, 1, 2]) / 2isfloat32;//keepsint64. - A float in the list.
torch.tensor([0, 1, 2])is alreadyint64, buttorch.tensor([0, 1, 2.])isfloat32.
MSELoss with a float64 or
int64 target runs its forward pass and only fails in backward(), and
BCEWithLogitsLoss accepts a float64 target without complaint.
3. The One-Line Fix: labels.long()
Call .long() on your target tensor before passing it to the loss function.
import torch
import torch.nn as nn
criterion = nn.CrossEntropyLoss()
# Labels that arrived as float (common when loaded from CSV/NumPy)
labels = torch.tensor([0, 2, 1, 3], dtype=torch.float32)
print(labels.dtype) # torch.float32
labels = labels.long()
print(labels.dtype) # torch.int64
outputs = torch.randn(4, 4) # batch_size=4, num_classes=4
loss = criterion(outputs, labels) # no error
print(loss.item()) # a positive float; the exact value depends on the random outputs
You can also apply the cast in your training loop:
for inputs, labels in dataloader:
inputs = inputs.to(device)
labels = labels.long().to(device) # <-- cast here, before loss
optimizer.zero_grad()
outputs = model(inputs)
loss = criterion(outputs, labels)
loss.backward()
optimizer.step()
Or fix it at the source, inside your Dataset. This also takes care of
float64 NumPy features, which would otherwise hit
mat1 and mat2 must have the same dtype, but got Double and Float in the first
Linear layer:
import numpy as np
import torch
from torch.utils.data import Dataset
class MyDataset(Dataset):
def __init__(self, X, y):
self.X = torch.tensor(X, dtype=torch.float32) # match the model's float32 weights
self.y = torch.tensor(y, dtype=torch.long) # class indices as int64
def __len__(self):
return len(self.y)
def __getitem__(self, idx):
return self.X[idx], self.y[idx]
torch.tensor(y, dtype=torch.long) converts whole-number floats cleanly: a
float64 array [0., 2., 1.] becomes tensor([0, 2, 1]).
.long() and .to(torch.int64) keep the tensor on
its device. .type(torch.LongTensor) does not: on a CUDA tensor it returned a
CPU tensor in my run, which then triggers a device mismatch error. Use
.long().
4. BCELoss vs CrossEntropyLoss Dtype Requirements
If you move a model from binary to multi-class classification (or back), the target dtype has to flip too:
- CrossEntropyLoss: targets are
torch.longclass indices, shape[batch_size]. (A float target of the same shape as the logits is also accepted and treated as class probabilities.) - BCELoss / BCEWithLogitsLoss: targets are float 0/1 values with the same shape as the output. Integer targets fail with
Found dtype Long but expected Float(BCELoss) orresult type Float can't be cast to the desired output type Long(BCEWithLogitsLoss).
import torch
import torch.nn as nn
batch_size = 8
num_classes = 5
# Multi-class: CrossEntropyLoss
logits_ce = torch.randn(batch_size, num_classes)
targets_ce = torch.randint(0, num_classes, (batch_size,)) # torch.int64 by default
criterion_ce = nn.CrossEntropyLoss()
loss_ce = criterion_ce(logits_ce, targets_ce) # OK
# Binary: BCEWithLogitsLoss
logits_bce = torch.randn(batch_size, 1)
targets_bce = torch.randint(0, 2, (batch_size, 1)).float() # must be float
criterion_bce = nn.BCEWithLogitsLoss()
loss_bce = criterion_bce(logits_bce, targets_bce) # OK
# Common mistake: float labels with CrossEntropyLoss
bad_targets = torch.randint(0, num_classes, (batch_size,)).float()
# criterion_ce(logits_ce, bad_targets) # RuntimeError: expected scalar type Long but found Float
criterion_ce(logits_ce, bad_targets.long()) # fixed
Watch the shape too. A [N] target with [N, 1] logits in
BCEWithLogitsLoss raises
ValueError: Target size (torch.Size([4])) must be the same as input size (torch.Size([4, 1])).
Use targets.float().unsqueeze(1).
5. Embedding Layer: A Different Error for the Same Mistake
nn.Embedding also needs integer indices, but in PyTorch 2.4.1 float indices do
not produce the "expected scalar type Long" message. They produce this, on both CPU and
CUDA (the GPU run says torch.cuda.FloatTensor):
RuntimeError: Expected tensor for argument #1 'indices' to have one of the following scalar types: Long, Int; but got torch.FloatTensor instead (while checking arguments for embedding)
import torch
import torch.nn as nn
vocab_size = 1000
embed_dim = 64
embedding = nn.Embedding(vocab_size, embed_dim)
indices_wrong = torch.tensor([4, 17, 42, 0], dtype=torch.float32)
# embedding(indices_wrong) # RuntimeError: Expected tensor for argument #1 'indices' ...
output = embedding(indices_wrong.long())
print(output.shape) # torch.Size([4, 64])
As the message says, int32 indices are fine here: embedding on an
int32 tensor returned torch.Size([4, 64]) on CPU and GPU. So an
np.int32 token array works for the embedding but fails if the same array is used
as CrossEntropyLoss labels. Other indexing ops are stricter:
gather with a float index raised gather(): Expected dtype int64 for index,
and F.one_hot on floats raised one_hot is only applicable to index tensor.
6. Dtype Table for Common PyTorch Loss Functions
What each loss accepted for the target in PyTorch 2.4.1, with float32 predictions.
"Backward fails" means the forward pass returned a loss and .backward() raised.
| Loss Function | Input (predictions) dtype | Target dtype | Notes |
|---|---|---|---|
nn.CrossEntropyLoss |
float32 | int64 (long) | Class indices, shape [N] (or [N, H, W] for segmentation). int32, float32 and float64 indices all fail. A float target of shape [N, C] is read as class probabilities. |
nn.NLLLoss |
float32 | int64 (long) | Input should be log-probabilities (apply log_softmax first). Same errors as CrossEntropyLoss. |
nn.BCELoss |
float32 | float32 | Input in [0, 1] (apply sigmoid first). int64 and float64 targets fail with Found dtype ... but expected Float. |
nn.BCEWithLogitsLoss |
float32 | float32 | Sigmoid + BCE in one stable step. int64 target fails; float64 target ran, including backward. |
nn.MSELoss |
float32 | float32 | Same shape as input. int64 and float64 targets: forward runs, backward fails. |
nn.L1Loss |
float32 | float32 | Mean absolute error. An int64 target ran forward and backward in my test, but float32 is the safe choice. |
The table splits by what the loss is doing: CrossEntropyLoss and NLLLoss index into discrete
classes, so they need long. The others compare continuous values, so give them
float32 targets that match the predictions.
7. Debugging Snippet: Print Dtype of Every Tensor Before the Forward Pass
When you are not sure which tensor has the wrong dtype, print the name, shape and dtype of everything you are about to use:
def debug_dtypes(**tensors):
"""Print name, shape, and dtype for each tensor. Call before forward pass."""
print("=" * 55)
for name, t in tensors.items():
if hasattr(t, "dtype"):
print(f" {name:20s} shape={str(tuple(t.shape)):20s} dtype={t.dtype}")
else:
print(f" {name:20s} (not a tensor, type={type(t).__name__})")
print("=" * 55)
# Usage inside your training loop:
for inputs, labels in dataloader:
debug_dtypes(inputs=inputs, labels=labels)
# After confirming dtypes, comment out the line above to speed up training
labels = labels.long()
outputs = model(inputs)
loss = criterion(outputs, labels)
Output from a DataLoader whose Dataset returns np.float32 labels
(batch of 32 images, 3x32x32):
=======================================================
inputs shape=(32, 3, 32, 32) dtype=torch.float32
labels shape=(32,) dtype=torch.float32
=======================================================
After labels.long():
=======================================================
inputs shape=(32, 3, 32, 32) dtype=torch.float32
labels shape=(32,) dtype=torch.int64
=======================================================
Or use assertions at the top of the loop while debugging:
assert inputs.dtype == torch.float32, f"inputs dtype mismatch: {inputs.dtype}"
assert labels.dtype == torch.long, f"labels dtype mismatch: {labels.dtype}"
They fail on the first bad batch with a message that names the tensor, which is quicker to read
than a traceback that ends inside torch._C._nn.cross_entropy_loss.
Other dtype errors you may hit next
-
NumPy features:
torch.from_numpyon a default NumPy float array givesfloat64, and afloat32Linearlayer then raisesmat1 and mat2 must have the same dtype, but got Double and Float.Conv2dsaysInput type (double) and bias type (float) should be the same. Fix withx.float(). -
Mixed precision (
torch.autocast): insideautocast("cuda", dtype=torch.float16)aLinearoutput isfloat16, andCrossEntropyLosson it withint64labels returned afloat32loss; aGradScalerstep then ran fine. Labels still have to belong: float labels gave the samenot implemented for 'Float'error under autocast. The trap is using thefloat16output outside theautocastblock, which raisedmat1 and mat2 must have the same dtype, but got Half and Float. -
Double precision models: after
model.double(), afloat32input fails withgot Float and Double. Inputs must befloat64too. Labels still staylong. -
Segmentation masks:
CrossEntropyLosstakes 2D targets of shape[N, H, W]. A float mask gaveexpected scalar type Long but found Floaton both CPU and GPU. Usemask.long(). -
One-hot targets: a float one-hot tensor with the same shape as the logits
does not raise.
CrossEntropyLossreads it as class probabilities and returned the same loss as the index labels. If you would rather pass indices, uselabels = one_hot_labels.argmax(dim=1), which is alreadyint64.
8. Related Fixes
Once the dtypes are right, the next crash in a new training script is often a device error or an out-of-memory error:
Found a mistake or have a question? Every error message on this page was produced by PyTorch 2.4.1 on CPU and on a GTX 1070 (CUDA 12.1). Other versions may word some of them differently.