Skip to content

[BUG]: pipeline sequence classification picks the wrong token for left-padded inputs (Llama, Qwen2, OPT) #6457

Description

@Arthur031221

Is there an existing issue for this bug?

  • I have searched the existing issues

The bug has not been fixed in the latest main branch

  • I have checked the latest main branch

Do you feel comfortable sharing a concise (minimal) script that reproduces the error? :)

Yes, I will share a minimal reproducible script.

Describe the bug

With pipeline parallelism, LlamaForSequenceClassification, Qwen2ForSequenceClassification and OPTForSequenceClassification take the logits (and the loss) of a left-padded row from the wrong token, so the pipelined model trains and evaluates on different outputs than the same model without pipeline parallelism.

The pipeline forwards in colossalai/shardformer/modeling/llama.py, qwen2.py and opt.py find the pooled token with

sequence_lengths = (torch.ne(input_ids, self.config.pad_token_id).sum(-1) - 1).to(logits.device)

That is the index of the last real token only when the padding is on the right. For the row [0, 0, 5, 6, 7, 8] with pad id 0 it gives 3 instead of 5. transformers 4.51.3 (the version pinned in requirements/requirements.txt) and the Qwen3 pipeline forward in this repo take the rightmost non-pad token instead. LlamaTokenizerFast pads on the left by default in transformers 4.51.3.

Minimal script on main (55929b0), one GPU. It compares the pipeline forward (one stage holding all layers) with the model's own forward:

import torch
from transformers import LlamaConfig, LlamaForSequenceClassification

import colossalai
from colossalai.cluster import ProcessGroupMesh
from colossalai.pipeline.stage_manager import PipelineStageManager
from colossalai.shardformer import ShardConfig
from colossalai.shardformer.modeling.llama import LlamaPipelineForwards

colossalai.launch(rank=0, world_size=1, host="localhost", port=29512, backend="nccl")
stage_manager = PipelineStageManager(ProcessGroupMesh(1), pipeline_axis=0)
shard_config = ShardConfig(enable_tensor_parallelism=False, pipeline_stage_manager=stage_manager)

config = LlamaConfig(
    vocab_size=64,
    hidden_size=32,
    intermediate_size=64,
    num_hidden_layers=2,
    num_attention_heads=4,
    num_key_value_heads=4,
    pad_token_id=0,
    num_labels=3,
)
torch.manual_seed(0)
model = LlamaForSequenceClassification(config).cuda().eval()

# first row is left-padded, second row has no padding
input_ids = torch.tensor([[0, 0, 5, 6, 7, 8], [9, 10, 11, 12, 13, 14]]).cuda()
attention_mask = (input_ids != config.pad_token_id).long()

with torch.no_grad():
    expected = model(input_ids=input_ids, attention_mask=attention_mask).logits
    output = LlamaPipelineForwards.llama_for_sequence_classification_forward(
        model,
        input_ids=input_ids,
        attention_mask=attention_mask,
        stage_manager=stage_manager,
        stage_index=[0, config.num_hidden_layers],
        shard_config=shard_config,
    ).logits

print("row matches model.forward:", torch.isclose(output, expected, atol=1e-5).all(-1).tolist())

Output on main:

row matches model.forward: [False, True]

The unpadded row matches and the left-padded row does not. Expected [True, True].

I have a fix with a test and will open a PR for it.

Environment

ColossalAI main at 55929b0, torch 2.8.0+cu128, transformers 4.51.3, Python 3.12.13, one RTX 5090.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions