Is there an existing issue for this bug?
The bug has not been fixed in the latest main branch
Do you feel comfortable sharing a concise (minimal) script that reproduces the error? :)
Yes, I will share a minimal reproducible script.
Describe the bug
With pipeline parallelism, LlamaForSequenceClassification, Qwen2ForSequenceClassification and OPTForSequenceClassification take the logits (and the loss) of a left-padded row from the wrong token, so the pipelined model trains and evaluates on different outputs than the same model without pipeline parallelism.
The pipeline forwards in colossalai/shardformer/modeling/llama.py, qwen2.py and opt.py find the pooled token with
sequence_lengths = (torch.ne(input_ids, self.config.pad_token_id).sum(-1) - 1).to(logits.device)
That is the index of the last real token only when the padding is on the right. For the row [0, 0, 5, 6, 7, 8] with pad id 0 it gives 3 instead of 5. transformers 4.51.3 (the version pinned in requirements/requirements.txt) and the Qwen3 pipeline forward in this repo take the rightmost non-pad token instead. LlamaTokenizerFast pads on the left by default in transformers 4.51.3.
Minimal script on main (55929b0), one GPU. It compares the pipeline forward (one stage holding all layers) with the model's own forward:
import torch
from transformers import LlamaConfig, LlamaForSequenceClassification
import colossalai
from colossalai.cluster import ProcessGroupMesh
from colossalai.pipeline.stage_manager import PipelineStageManager
from colossalai.shardformer import ShardConfig
from colossalai.shardformer.modeling.llama import LlamaPipelineForwards
colossalai.launch(rank=0, world_size=1, host="localhost", port=29512, backend="nccl")
stage_manager = PipelineStageManager(ProcessGroupMesh(1), pipeline_axis=0)
shard_config = ShardConfig(enable_tensor_parallelism=False, pipeline_stage_manager=stage_manager)
config = LlamaConfig(
vocab_size=64,
hidden_size=32,
intermediate_size=64,
num_hidden_layers=2,
num_attention_heads=4,
num_key_value_heads=4,
pad_token_id=0,
num_labels=3,
)
torch.manual_seed(0)
model = LlamaForSequenceClassification(config).cuda().eval()
# first row is left-padded, second row has no padding
input_ids = torch.tensor([[0, 0, 5, 6, 7, 8], [9, 10, 11, 12, 13, 14]]).cuda()
attention_mask = (input_ids != config.pad_token_id).long()
with torch.no_grad():
expected = model(input_ids=input_ids, attention_mask=attention_mask).logits
output = LlamaPipelineForwards.llama_for_sequence_classification_forward(
model,
input_ids=input_ids,
attention_mask=attention_mask,
stage_manager=stage_manager,
stage_index=[0, config.num_hidden_layers],
shard_config=shard_config,
).logits
print("row matches model.forward:", torch.isclose(output, expected, atol=1e-5).all(-1).tolist())
Output on main:
row matches model.forward: [False, True]
The unpadded row matches and the left-padded row does not. Expected [True, True].
I have a fix with a test and will open a PR for it.
Environment
ColossalAI main at 55929b0, torch 2.8.0+cu128, transformers 4.51.3, Python 3.12.13, one RTX 5090.
Is there an existing issue for this bug?
The bug has not been fixed in the latest main branch
Do you feel comfortable sharing a concise (minimal) script that reproduces the error? :)
Yes, I will share a minimal reproducible script.
Describe the bug
With pipeline parallelism,
LlamaForSequenceClassification,Qwen2ForSequenceClassificationandOPTForSequenceClassificationtake the logits (and the loss) of a left-padded row from the wrong token, so the pipelined model trains and evaluates on different outputs than the same model without pipeline parallelism.The pipeline forwards in
colossalai/shardformer/modeling/llama.py,qwen2.pyandopt.pyfind the pooled token withThat is the index of the last real token only when the padding is on the right. For the row
[0, 0, 5, 6, 7, 8]with pad id 0 it gives 3 instead of 5. transformers 4.51.3 (the version pinned inrequirements/requirements.txt) and the Qwen3 pipeline forward in this repo take the rightmost non-pad token instead.LlamaTokenizerFastpads on the left by default in transformers 4.51.3.Minimal script on main (55929b0), one GPU. It compares the pipeline forward (one stage holding all layers) with the model's own
forward:Output on main:
The unpadded row matches and the left-padded row does not. Expected
[True, True].I have a fix with a test and will open a PR for it.
Environment
ColossalAI main at 55929b0, torch 2.8.0+cu128, transformers 4.51.3, Python 3.12.13, one RTX 5090.