Skip to content

HuggingFace Tokenizer Errors: Padding, Truncation, and the Left/Right Trap

Tested with: transformers 5.17.0, tokenizers 0.23.2, torch 2.14.0+cpu, Python 3.12.3 (venv), Ubuntu 24.04, CPU only. Tokenizers and models: gpt2, bert-base-uncased, HuggingFaceTB/SmolLM2-135M (a small Llama-architecture model, no login needed). Last run 2026-09-27.

TL;DR:
  1. Asking to pad but the tokenizer does not have a padding token: gpt2 and SmolLM2 load with pad_token=None. Fix: tokenizer.pad_token = tokenizer.eos_token.
  2. Unable to create tensor, you should probably activate truncation and/or padding: you asked for return_tensors="pt" from sequences of different lengths. Fix: padding=True, truncation=True.
  3. `max_length` is ignored when `padding`=`True`: without truncation=True, long inputs are not cut, and gpt2 then fails with IndexError: index out of range in self. Pass truncation=True with max_length.
  4. A decoder-only architecture is being used, but right-padding was detected!: batched generate() output for shorter prompts is wrong. Fix: tokenizer.padding_side = "left" and always pass attention_mask.
  5. Added a new [PAD] token? Call model.resize_token_embeddings(len(tokenizer)) or the forward pass raises IndexError: index out of range in self.

Single-sequence tokenization almost never breaks. The moment you batch more than one input through a HuggingFace tokenizer, you run into a handful of padding and truncation issues. Some raise a clear error at the tokenizer call. Others only print a warning, or print nothing at all, and hand you wrong output.

pip install transformers torch

Reproduced

Every row was run in a fresh venv with the versions above. The defaults the tokenizers loaded with: gpt2 has pad_token=None, padding_side="right", model_max_length=1024; SmolLM2-135M has pad_token=None, padding_side="right", model_max_length=8192; bert-base-uncased has pad_token="[PAD]", padding_side="right".

What I ranExact resultFix that worked
gpt2, tok(texts, padding=True, return_tensors="pt")ValueError: Asking to pad but the tokenizer does not have a padding token. Please select a token to use as `pad_token` `(tokenizer.pad_token = tokenizer.eos_token e.g.)` or add a new pad token via `tokenizer.add_special_tokens({'pad_token': '[PAD]'})`.tok.pad_token = tok.eos_token (pad id 50256, attention mask [1, 1, 0, 0, 0, 0] on the short row)
gpt2 and bert-base-uncased, tok(texts, return_tensors="pt"), two lengthsValueError: Unable to create tensor, you should probably activate truncation and/or padding with 'padding=True' 'truncation=True' to have batched tensors with the same length. Perhaps your features (`input_ids` in this case) have excessive nesting (inputs type `list` where type `int` is expected).padding=True, truncation=True
torch.stack or default_collate on per-example encodingsRuntimeError: stack expects each tensor to be equal size, but got [2] at entry 0 and [6] at entry 1tokenizer.pad(features, padding=True, return_tensors="pt") in the collate function
gpt2, 2001-token input, max_length=16, padding=True, no truncationUserWarning: `max_length` is ignored when `padding`=`True` and there is no truncation strategy. To pad to max length, use `padding='max_length'`. Output shape [2, 2001]Add truncation=True: shape [2, 16]
Same input, padding="max_length", no truncationThe Unable to create tensor error aboveAdd truncation=True
gpt2 model forward on the 2001-token inputTokenizer logs Token indices sequence length is longer than the specified maximum sequence length for this model (2001 > 1024), then IndexError: index out of range in selfTruncate to 1024 or less
SmolLM2-135M, batched greedy generate(), right paddingA decoder-only architecture is being used, but right-padding was detected! For correct generation results, please set `padding_side='left'` when initializing the tokenizer. Short prompt continued as 'The capital of France is the capital of France. The capital'padding_side = "left": ' the capital of the country.\n\nThe capital of France', same as unbatched
Left padding, generate(input_ids) without attention_maskThe attention mask is not set with a batched input, and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Output changed to ' the capital of the department of Alsace-Lor'Pass **inputs so the mask goes in
SmolLM2-135M, add_special_tokens({"pad_token": "[PAD]"}) without resizePad id 49152, embedding rows 49152: IndexError: index out of range in selfmodel.resize_token_embeddings(len(tokenizer)): 49153 rows, forward runs

"Asking to pad but the tokenizer does not have a padding token"

You get this from decoder-only tokenizers. gpt2 and SmolLM2-135M both load with pad_token set to None, while an encoder tokenizer like bert-base-uncased ships with [PAD].

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("gpt2")
batch = tokenizer(["short text", "a somewhat longer piece of text"], padding=True, return_tensors="pt")
ValueError: Asking to pad but the tokenizer does not have a padding token. Please select a token to use as `pad_token` `(tokenizer.pad_token = tokenizer.eos_token e.g.)` or add a new pad token via `tokenizer.add_special_tokens({'pad_token': '[PAD]'})`.

The error message names both fixes. The first is the one most people reach for:

tokenizer = AutoTokenizer.from_pretrained("gpt2")
tokenizer.pad_token = tokenizer.eos_token

batch = tokenizer(["short text", "a somewhat longer piece of text"], padding=True, return_tensors="pt")
print(batch)
{'input_ids': tensor([[19509,  2420, 50256, 50256, 50256, 50256],
        [   64,  6454,  2392,  3704,   286,  2420]]), 'attention_mask': tensor([[1, 1, 0, 0, 0, 0],
        [1, 1, 1, 1, 1, 1]])}

Reusing eos_token as the pad token works for inference and most fine-tuning, because the attention mask tells the model which positions are real. The case where it bites: DataCollatorForLanguageModeling sets the label of every pad-id position to -100, and with pad = eos that includes your real end-of-text token. With two examples that each end in <|endoftext|> (id 50256), the collator produced these labels:

# pad_token = eos_token: the real eos at the end of row 2 is masked too
[[19509, 2420, -100, -100, -100, -100, -100], [64, 6454, 2392, 3704, 286, 2420, -100]]

# separate [PAD] token (id 50257): eos keeps label 50256
[[19509, 2420, 50256, -100, -100, -100, -100], [64, 6454, 2392, 3704, 286, 2420, 50256]]

In the first case the model is never trained to emit eos, so it never learns when to stop. If that matters for your fine-tune, add a separate token:

tokenizer.add_special_tokens({"pad_token": "[PAD]"})
model.resize_token_embeddings(len(tokenizer))

Don't skip resize_token_embeddings. On SmolLM2-135M the new pad token got id 49152 while the embedding matrix had 49152 rows (ids 0 to 49151), and the first forward pass on a padded batch failed with IndexError: index out of range in self. After the resize the matrix had 49153 rows and the same batch ran. transformers 5.17.0 also logs that the new row is initialized from the mean and covariance of the old embeddings (pass mean_resizing=False to turn that off).


Batched generation gives different output for shorter prompts

No exception here, only a warning that is easy to miss in a notebook. You fix the pad-token error, batch prompts of different lengths through model.generate(), and the shorter prompts get worse output than when you run them one at a time.

The cause is padding side. tokenizer.padding_side defaults to "right", which is fine for encoder models such as BERT classification, but wrong for decoder-only generation. A causal LM appends new tokens after the last position in each row. With right padding, the last positions of a shorter prompt are pad tokens, so the model continues from padding instead of from your prompt. transformers detects this and prints:

A decoder-only architecture is being used, but right-padding was detected! For correct generation results, please set `padding_side='left'` when initializing the tokenizer.

Greedy decoding with SmolLM2-135M, max_new_tokens=12, prompts "The capital of France is" and "My favourite programming language is Python because it", new tokens only:

SetupShort prompt continuation
One prompt at a time (no padding)' the capital of the country.\n\nThe capital of France'
Batched, right padding'The capital of France is the capital of France. The capital'
Batched, left padding, with attention_mask' the capital of the country.\n\nThe capital of France'
Batched, left padding, attention_mask not passed' the capital of the department of Alsace-Lor'

The longer prompt needed no padding and produced ' is so easy to learn and it is a great language to' in all four setups. Only the padded row changed.

tokenizer.padding_side = "left"
tokenizer.pad_token = tokenizer.eos_token  # still needed

inputs = tokenizer(
    ["The capital of France is", "My favourite programming language is Python because it"],
    padding=True, return_tensors="pt"
)
outputs = model.generate(**inputs, max_new_tokens=12, do_sample=False,
                         pad_token_id=tokenizer.pad_token_id)

Pass **inputs, not just inputs["input_ids"]. When pad and eos share an id, generate() cannot rebuild the mask from the ids, warns The attention mask is not set with a batched input, and cannot be inferred from input because pad token is same as eos token., and attends to the padding, which is how the last row of the table went wrong. The same thing shows up in a plain forward pass: on the left-padded short row (3 pad positions), the final-position logits differed from the unpadded run by at most 0.25 with the mask and by up to 3.59 without it.

Left padding matters for generate() with decoder-only models. For classification and encoder models, keep the default right padding.


"Unable to create tensor, you should probably activate truncation and/or padding"

If you ask for tensors without padding, the tokenizer itself refuses to build a ragged batch:

batch = tokenizer(["short text", "a somewhat longer piece of text"], return_tensors="pt")
ValueError: Unable to create tensor, you should probably activate truncation and/or padding with 'padding=True' 'truncation=True' to have batched tensors with the same length. Perhaps your features (`input_ids` in this case) have excessive nesting (inputs type `list` where type `int` is expected).

The same text came from gpt2 and bert-base-uncased, so it is not a decoder-only issue. The fix is in the message:

batch = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")

The harder case is when you tokenize per example inside a Dataset and let a DataLoader batch them. The default collate function stacks tensors, and the traceback then points at PyTorch rather than the tokenizer:

RuntimeError: stack expects each tensor to be equal size, but got [2] at entry 0 and [6] at entry 1

The numbers are the sequence lengths of the two examples. Pad in the collate function with the tokenizer's own pad(), so pad ids and attention masks stay consistent:

def collate_fn(batch):
    return tokenizer.pad(batch, padding=True, return_tensors="pt")

On the two gpt2 encodings above (with pad_token = eos_token) this returned a [2, 6] batch with the same ids and mask as tokenizing both texts together.


"`max_length` is ignored when `padding`=`True`": long inputs are not truncated

This one is a warning, not an exception, which is why it causes problems downstream instead of at the tokenizer call:

long_text = "hello " * 2000
batch = tokenizer([long_text, "hi"], max_length=16, padding=True, return_tensors="pt")
print(batch["input_ids"].shape)
UserWarning: `max_length` is ignored when `padding`=`True` and there is no truncation strategy. To pad to max length, use `padding='max_length'`.
torch.Size([2, 2001])

With padding=True and no truncation argument, max_length does nothing: the 2001-token input comes through whole. gpt2 has 1024 position embeddings, so the next step fails. The tokenizer logs Token indices sequence length is longer than the specified maximum sequence length for this model (2001 > 1024). Running this sequence through the model will result in indexing errors, and the forward pass on CPU raises IndexError: index out of range in self, far from the tokenizer call that caused it.

The rules in transformers 5.17.0, from running each case:

  • max_length alone, padding left at its default: truncates (1 sequence came back with 16 tokens), with no warning.
  • max_length + padding=True, no truncation: the warning above, no truncation.
  • max_length + padding="max_length", no truncation: the Unable to create tensor error, because the long row stays at 2001 while the short row is padded to 16.
  • truncation=True without max_length: truncates to model_max_length (1024 for gpt2).
# max_length is ignored, long rows pass through
batch = tokenizer(texts, max_length=512, padding=True, return_tensors="pt")

# Correct: shape was [2, 16] for max_length=16
batch = tokenizer(texts, max_length=512, padding=True, truncation=True, return_tensors="pt")

Older posts quote a Truncation was not explicitly activated warning. That string is not in transformers 5.17.0; the max_length-only case now truncates silently. Either way, treat max_length and truncation=True as a pair.


A tokenize helper that avoids all of these

def safe_tokenize(tokenizer, texts, max_length=512, for_generation=False):
    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token

    if for_generation:
        tokenizer.padding_side = "left"

    return tokenizer(
        texts,
        padding=True,
        truncation=True,
        max_length=max_length,
        return_tensors="pt",
    )

With SmolLM2-135M, safe_tokenize(tok, ["The capital of France is", "hello " * 3000], for_generation=True) returned a [2, 512] left-padded batch with no warnings, and model.generate(**enc, max_new_tokens=6, do_sample=False) continued the short prompt with ' the capital of the country.', the same start as the unbatched run. Remember to pass **enc so the attention mask reaches generate().

Most of what shows up above comes from two assumptions: that a decoder-only tokenizer ships configured for batching (gpt2 and SmolLM2 do not), and that padding and truncation are independent flags. Of the two, the second is the one that fails late: the tokenizer call succeeds and the crash comes a few calls later in the model.