Skip to content

PyTorch RuntimeError: Expected Scalar Type Long but Found Float

Tested with: PyTorch 2.4.1 (CUDA 12.1), NumPy 2.4.4, Python 3.12.3, NVIDIA GeForce GTX 1070 8 GB, driver 580, Ubuntu 24.04. Last run 2026-09-27.

TL;DR:
  1. Your class labels are not torch.int64. Cast them before the loss: loss = criterion(outputs, labels.long()), or build them as torch.tensor(y, dtype=torch.long) in your Dataset.
  2. On the GPU the same bug prints a different message: "nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Float'. Same fix.
  3. int32 labels fail too (found Int), and so do float64 labels from NumPy (found Double). Only int64 works.
  4. BCELoss, BCEWithLogitsLoss and MSELoss go the other way: they want float targets, and BCELoss wants exactly float32. For inputs, mat1 and mat2 must have the same dtype means a float64 NumPy array reached a float32 layer: use .float().

Training a classification model, and the loop crashes with RuntimeError: expected scalar type Long but found Float. The cause is nearly always the same: CrossEntropyLoss (or NLLLoss) wants integer class indices and got floating-point values. Below is exactly what PyTorch 2.4.1 prints on CPU and on CUDA, where the float labels usually come from, and the related dtype errors that show up once this one is fixed.

Reproduced on Real Hardware

Every row ran in one script with small random tensors, once on the CPU and once on the GTX 1070. Where the two wordings differ, both are quoted.

What I ranError or output (CPU / CUDA)Fix that worked
CrossEntropyLoss, labels float32CPU: expected scalar type Long but found Float
CUDA: "nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Float'
labels.long()
Labels float64 (from NumPy)CPU: ... but found Double
CUDA: ... not implemented for 'Double'
labels.long()
Labels int32 (np.int32 or .int())CPU: ... but found Int
CUDA: ... not implemented for 'Int'
labels.long()
F.nll_loss, float targetsSame messages as CrossEntropyLosstarget.long()
Segmentation: logits [2, 3, 4, 4], mask float32 [2, 4, 4]expected scalar type Long but found Float on both CPU and CUDAmask.long()
CrossEntropyLoss with a float one-hot target, same shape as the logitsNo error: treated as class probabilities, same loss as the index labelsNot needed
BCEWithLogitsLoss, target int64result type Float can't be cast to the desired output type Long (both)target.float()
BCELoss, target int64 / float64Found dtype Long but expected Float / Found dtype Double but expected Float (both)target.float()
BCEWithLogitsLoss, target float64No errorNot needed
MSELoss, target int64 or float64Forward works; .backward() raises Found dtype Long but expected Float / Found dtype Double but expected Floattarget.float()
nn.Embedding, float32 indicesExpected tensor for argument #1 'indices' to have one of the following scalar types: Long, Int; but got torch.FloatTensor instead (while checking arguments for embedding) (torch.cuda.FloatTensor on GPU).long(); int32 also works
nn.Linear (float32) fed torch.from_numpy of a float64 arraymat1 and mat2 must have the same dtype, but got Double and Float (both)x.float()
model.double(), input float32mat1 and mat2 must have the same dtype, but got Float and DoubleMatch the input to the model
float32 @ float64 with @ or torch.mmCPU: expected m1 and m2 to have the same dtype, but got: float != double
CUDA: expected mat1 and mat2 ...
Cast one side
Conv2d fed a float64 imageInput type (double) and bias type (float) should be the samex.float()
AMP: float16 output from autocast used in a float32 Linear outside the blockmat1 and mat2 must have the same dtype, but got Half and FloatKeep the op inside autocast, or .float()
AMP: CrossEntropyLoss on float16 logits with int64 labels inside autocastNo error, loss is float32Not needed

If you search for the CPU wording but train on a GPU, you won't find your error: the CUDA kernel reports the unsupported dtype instead. And some wrong dtypes do not raise at all, or only raise in backward().

1. The Exact Error Message (CPU and CUDA)

This is the full traceback PyTorch 2.4.1 prints for float labels on the CPU (site-packages path shortened, caret lines removed):

Traceback (most recent call last):
  File "/tmp/scalar/repro2.py", line 7, in <module>
    loss = ce(outputs, labels)
  File ".../torch/nn/modules/module.py", line 1553, in _wrapped_call_impl
    return self._call_impl(*args, **kwargs)
  File ".../torch/nn/modules/module.py", line 1562, in _call_impl
    return forward_call(*args, **kwargs)
  File ".../torch/nn/modules/loss.py", line 1188, in forward
    return F.cross_entropy(input, target, weight=self.weight,
  File ".../torch/nn/functional.py", line 3104, in cross_entropy
    return torch._C._nn.cross_entropy_loss(input, target, weight, _Reduction.get_enum(reduction), ignore_index, label_smoothing)
RuntimeError: expected scalar type Long but found Float

The same code with outputs and labels on cuda goes through the same frames and ends with:

RuntimeError: "nll_loss_forward_reduce_cuda_kernel_2d_index" not implemented for 'Float'

The last word tells you what you passed: Float is float32, Double is float64, Int is int32. The key line is loss = ce(outputs, labels): labels needs to be torch.int64 (also called torch.long). Older guides quote a longer wording, Expected object of scalar type Long but got scalar type Float for argument #2 'target', from PyTorch 1.x. I did not run a 1.x build, and 2.4.1 did not print it.

2. Root Cause: CrossEntropyLoss Expects Class Indices as torch.long

With a target of shape [N], nn.CrossEntropyLoss computes log_softmax and then nll_loss, which uses each target value as an index into the class dimension. The kernel is written for int64 indices only, so float32, float64 and even int32 are rejected.

Where the wrong dtype usually comes from, each checked on torch 2.4.1 and NumPy 2.4.4:

  • Labels loaded with NumPy. np.loadtxt returns float64 by default, and torch.from_numpy keeps that dtype, giving found Double.
  • Dataset.__getitem__ returns np.float32 labels. The default DataLoader collate turns them into a float32 batch. Returning a plain Python int gives int64.
  • An explicit np.int32 array. It becomes a torch.int32 tensor and fails with found Int.
  • True division. torch.tensor([0, 1, 2]) / 2 is float32; // keeps int64.
  • A float in the list. torch.tensor([0, 1, 2]) is already int64, but torch.tensor([0, 1, 2.]) is float32.
Important: PyTorch does not cast class labels for you. But other losses are less strict than you might expect: MSELoss with a float64 or int64 target runs its forward pass and only fails in backward(), and BCEWithLogitsLoss accepts a float64 target without complaint.

3. The One-Line Fix: labels.long()

Call .long() on your target tensor before passing it to the loss function.

import torch
import torch.nn as nn

criterion = nn.CrossEntropyLoss()

# Labels that arrived as float (common when loaded from CSV/NumPy)
labels = torch.tensor([0, 2, 1, 3], dtype=torch.float32)
print(labels.dtype)  # torch.float32

labels = labels.long()
print(labels.dtype)  # torch.int64

outputs = torch.randn(4, 4)  # batch_size=4, num_classes=4
loss = criterion(outputs, labels)  # no error
print(loss.item())   # a positive float; the exact value depends on the random outputs

You can also apply the cast in your training loop:

for inputs, labels in dataloader:
    inputs  = inputs.to(device)
    labels  = labels.long().to(device)   # <-- cast here, before loss

    optimizer.zero_grad()
    outputs = model(inputs)
    loss    = criterion(outputs, labels)
    loss.backward()
    optimizer.step()

Or fix it at the source, inside your Dataset. This also takes care of float64 NumPy features, which would otherwise hit mat1 and mat2 must have the same dtype, but got Double and Float in the first Linear layer:

import numpy as np
import torch
from torch.utils.data import Dataset

class MyDataset(Dataset):
    def __init__(self, X, y):
        self.X = torch.tensor(X, dtype=torch.float32)  # match the model's float32 weights
        self.y = torch.tensor(y, dtype=torch.long)     # class indices as int64

    def __len__(self):
        return len(self.y)

    def __getitem__(self, idx):
        return self.X[idx], self.y[idx]

torch.tensor(y, dtype=torch.long) converts whole-number floats cleanly: a float64 array [0., 2., 1.] becomes tensor([0, 2, 1]).

Tip: .long() and .to(torch.int64) keep the tensor on its device. .type(torch.LongTensor) does not: on a CUDA tensor it returned a CPU tensor in my run, which then triggers a device mismatch error. Use .long().

4. BCELoss vs CrossEntropyLoss Dtype Requirements

If you move a model from binary to multi-class classification (or back), the target dtype has to flip too:

  • CrossEntropyLoss: targets are torch.long class indices, shape [batch_size]. (A float target of the same shape as the logits is also accepted and treated as class probabilities.)
  • BCELoss / BCEWithLogitsLoss: targets are float 0/1 values with the same shape as the output. Integer targets fail with Found dtype Long but expected Float (BCELoss) or result type Float can't be cast to the desired output type Long (BCEWithLogitsLoss).
import torch
import torch.nn as nn

batch_size = 8
num_classes = 5

# Multi-class: CrossEntropyLoss
logits_ce = torch.randn(batch_size, num_classes)
targets_ce = torch.randint(0, num_classes, (batch_size,))  # torch.int64 by default
criterion_ce = nn.CrossEntropyLoss()
loss_ce = criterion_ce(logits_ce, targets_ce)  # OK

# Binary: BCEWithLogitsLoss
logits_bce = torch.randn(batch_size, 1)
targets_bce = torch.randint(0, 2, (batch_size, 1)).float()  # must be float
criterion_bce = nn.BCEWithLogitsLoss()
loss_bce = criterion_bce(logits_bce, targets_bce)  # OK

# Common mistake: float labels with CrossEntropyLoss
bad_targets = torch.randint(0, num_classes, (batch_size,)).float()
# criterion_ce(logits_ce, bad_targets)  # RuntimeError: expected scalar type Long but found Float
criterion_ce(logits_ce, bad_targets.long())  # fixed

Watch the shape too. A [N] target with [N, 1] logits in BCEWithLogitsLoss raises ValueError: Target size (torch.Size([4])) must be the same as input size (torch.Size([4, 1])). Use targets.float().unsqueeze(1).

5. Embedding Layer: A Different Error for the Same Mistake

nn.Embedding also needs integer indices, but in PyTorch 2.4.1 float indices do not produce the "expected scalar type Long" message. They produce this, on both CPU and CUDA (the GPU run says torch.cuda.FloatTensor):

RuntimeError: Expected tensor for argument #1 'indices' to have one of the following scalar types: Long, Int; but got torch.FloatTensor instead (while checking arguments for embedding)
import torch
import torch.nn as nn

vocab_size = 1000
embed_dim  = 64
embedding  = nn.Embedding(vocab_size, embed_dim)

indices_wrong = torch.tensor([4, 17, 42, 0], dtype=torch.float32)
# embedding(indices_wrong)  # RuntimeError: Expected tensor for argument #1 'indices' ...

output = embedding(indices_wrong.long())
print(output.shape)  # torch.Size([4, 64])

As the message says, int32 indices are fine here: embedding on an int32 tensor returned torch.Size([4, 64]) on CPU and GPU. So an np.int32 token array works for the embedding but fails if the same array is used as CrossEntropyLoss labels. Other indexing ops are stricter: gather with a float index raised gather(): Expected dtype int64 for index, and F.one_hot on floats raised one_hot is only applicable to index tensor.

6. Dtype Table for Common PyTorch Loss Functions

What each loss accepted for the target in PyTorch 2.4.1, with float32 predictions. "Backward fails" means the forward pass returned a loss and .backward() raised.

Loss Function Input (predictions) dtype Target dtype Notes
nn.CrossEntropyLoss float32 int64 (long) Class indices, shape [N] (or [N, H, W] for segmentation). int32, float32 and float64 indices all fail. A float target of shape [N, C] is read as class probabilities.
nn.NLLLoss float32 int64 (long) Input should be log-probabilities (apply log_softmax first). Same errors as CrossEntropyLoss.
nn.BCELoss float32 float32 Input in [0, 1] (apply sigmoid first). int64 and float64 targets fail with Found dtype ... but expected Float.
nn.BCEWithLogitsLoss float32 float32 Sigmoid + BCE in one stable step. int64 target fails; float64 target ran, including backward.
nn.MSELoss float32 float32 Same shape as input. int64 and float64 targets: forward runs, backward fails.
nn.L1Loss float32 float32 Mean absolute error. An int64 target ran forward and backward in my test, but float32 is the safe choice.

The table splits by what the loss is doing: CrossEntropyLoss and NLLLoss index into discrete classes, so they need long. The others compare continuous values, so give them float32 targets that match the predictions.

7. Debugging Snippet: Print Dtype of Every Tensor Before the Forward Pass

When you are not sure which tensor has the wrong dtype, print the name, shape and dtype of everything you are about to use:

def debug_dtypes(**tensors):
    """Print name, shape, and dtype for each tensor. Call before forward pass."""
    print("=" * 55)
    for name, t in tensors.items():
        if hasattr(t, "dtype"):
            print(f"  {name:20s}  shape={str(tuple(t.shape)):20s}  dtype={t.dtype}")
        else:
            print(f"  {name:20s}  (not a tensor, type={type(t).__name__})")
    print("=" * 55)


# Usage inside your training loop:
for inputs, labels in dataloader:
    debug_dtypes(inputs=inputs, labels=labels)

    # After confirming dtypes, comment out the line above to speed up training
    labels = labels.long()
    outputs = model(inputs)
    loss = criterion(outputs, labels)

Output from a DataLoader whose Dataset returns np.float32 labels (batch of 32 images, 3x32x32):

=======================================================
  inputs                shape=(32, 3, 32, 32)       dtype=torch.float32
  labels                shape=(32,)                 dtype=torch.float32
=======================================================

After labels.long():

=======================================================
  inputs                shape=(32, 3, 32, 32)       dtype=torch.float32
  labels                shape=(32,)                 dtype=torch.int64
=======================================================

Or use assertions at the top of the loop while debugging:

assert inputs.dtype == torch.float32, f"inputs dtype mismatch: {inputs.dtype}"
assert labels.dtype == torch.long,    f"labels dtype mismatch: {labels.dtype}"

They fail on the first bad batch with a message that names the tensor, which is quicker to read than a traceback that ends inside torch._C._nn.cross_entropy_loss.

Other dtype errors you may hit next

  • NumPy features: torch.from_numpy on a default NumPy float array gives float64, and a float32 Linear layer then raises mat1 and mat2 must have the same dtype, but got Double and Float. Conv2d says Input type (double) and bias type (float) should be the same. Fix with x.float().
  • Mixed precision (torch.autocast): inside autocast("cuda", dtype=torch.float16) a Linear output is float16, and CrossEntropyLoss on it with int64 labels returned a float32 loss; a GradScaler step then ran fine. Labels still have to be long: float labels gave the same not implemented for 'Float' error under autocast. The trap is using the float16 output outside the autocast block, which raised mat1 and mat2 must have the same dtype, but got Half and Float.
  • Double precision models: after model.double(), a float32 input fails with got Float and Double. Inputs must be float64 too. Labels still stay long.
  • Segmentation masks: CrossEntropyLoss takes 2D targets of shape [N, H, W]. A float mask gave expected scalar type Long but found Float on both CPU and GPU. Use mask.long().
  • One-hot targets: a float one-hot tensor with the same shape as the logits does not raise. CrossEntropyLoss reads it as class probabilities and returned the same loss as the index labels. If you would rather pass indices, use labels = one_hot_labels.argmax(dim=1), which is already int64.

Once the dtypes are right, the next crash in a new training script is often a device error or an out-of-memory error:


Found a mistake or have a question? Every error message on this page was produced by PyTorch 2.4.1 on CPU and on a GTX 1070 (CUDA 12.1). Other versions may word some of them differently.