Skip to content

Add grid-stride + ROCm cap to grad_mean kernel - #6078

Open
q10 wants to merge 1 commit into
pytorch:mainfrom
q10:export-D113351689
Open

Add grid-stride + ROCm cap to grad_mean kernel#6078
q10 wants to merge 1 commit into
pytorch:mainfrom
q10:export-D113351689

Conversation

@q10

@q10 q10 commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Summary:
The MEAN-pooling TBE backward launches grad_mean{,_vbe}_kernel with
grid = div_round_up(total_B, kMaxThreads / grad_mean_warp_size) and
block = dim3(grad_mean_warp_size, kMaxThreads / grad_mean_warp_size)
(total_B = B * T), so total threads ~= total_B * grad_mean_warp_size exceeds
the HIP 2^32 threads-per-launch limit on ROCm for large total_B.

Cap the launch with utils::cuda::cap_grid_dim_x(..., OverflowOnly) and add a
ROCm grid-stride loop over b_t to the kernel (no internal early-returns; single
guard converted to the loop bound). No-op on CUDA.

Reviewed By: cthi

Differential Revision: D113351689

Summary:
The MEAN-pooling TBE backward launches grad_mean{,_vbe}_kernel with
grid = div_round_up(total_B, kMaxThreads / grad_mean_warp_size) and
block = dim3(grad_mean_warp_size, kMaxThreads / grad_mean_warp_size)
(total_B = B * T), so total threads ~= total_B * grad_mean_warp_size exceeds
the HIP 2^32 threads-per-launch limit on ROCm for large total_B.

Cap the launch with utils::cuda::cap_grid_dim_x(..., OverflowOnly) and add a
ROCm grid-stride loop over b_t to the kernel (no internal early-returns; single
guard converted to the loop bound). No-op on CUDA.

Reviewed By: cthi

Differential Revision: D113351689
@meta-codesync

meta-codesync Bot commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

@q10 has exported this pull request. If you are a Meta employee, you can view the originating Diff in D113351689.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant