-
Notifications
You must be signed in to change notification settings - Fork 766
Pull requests: pytorch/FBGEMM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
mx4 quantize: compute per-group scale via exact FP32 bit-construction instead of FP64 exp2
cla signed
meta-exported
#6083
opened Jul 28, 2026 by
qchip
Contributor
Loading…
Fold grid.y/z into permute cap_grid_dim_x to fix HIP 2^32 overflow
cla signed
meta-exported
#6082
opened Jul 28, 2026 by
q10
Contributor
Loading…
replace raw copy threads with a persistent copy_executor_
cla signed
meta-exported
#6079
opened Jul 27, 2026 by
FriedCosey
Loading…
Add grid-stride + ROCm cap to grad_mean kernel
ciflow/rocm
cla signed
meta-exported
module: rocm
#6078
opened Jul 27, 2026 by
q10
Contributor
Loading…
Replace raw consumer threads + polled queue with a push executor
cla signed
meta-exported
#6076
opened Jul 27, 2026 by
FriedCosey
Loading…
Make RES chunk size + ship/copy thread counts config-tunable
cla signed
meta-exported
#6075
opened Jul 25, 2026 by
FriedCosey
Loading…
Un-nest tensor_copy_chunk dispatch lambdas (#6074)
cla signed
meta-exported
#6074
opened Jul 25, 2026 by
FriedCosey
Loading…
Bump setuptools from 78.1.1 to 83.0.0 in /fbgemm_gpu
cla signed
dependencies
Pull requests that update a dependency file
python
Pull requests that update python code
#6070
opened Jul 25, 2026 by
dependabot
Bot
Loading…
Drain the stream queue across N consumer threads in raw_embedding_streamer (#6069)
cla signed
meta-exported
#6069
opened Jul 25, 2026 by
FriedCosey
Loading…
Chunk the D2H copy across N copy threads in raw_embedding_streamer
cla signed
meta-exported
#6068
opened Jul 24, 2026 by
FriedCosey
Loading…
Drop redundant output pre-zero in quantized TBE CPU pooled forward (#6065)
cla signed
meta-exported
#6065
opened Jul 24, 2026 by
knoebelja
Loading…
Add CPU nbit-forward tests for empty/pruned-bag zero-fill (#6064)
cla signed
meta-exported
#6064
opened Jul 24, 2026 by
knoebelja
Loading…
Drop redundant uint32_t narrowing at cap_grid_dim_x in repeat_arange/get_source_mask
cla signed
meta-exported
#6063
opened Jul 24, 2026 by
q10
Contributor
Loading…
Migrate TBE benchmark type imports to fbgemm_gpu.tbe.* leaves (Phase 3.7)
cla signed
meta-exported
#6060
opened Jul 24, 2026 by
q10
Contributor
Loading…
Scale cross-commit benchmark stats analysis to full sweeps
cla signed
meta-exported
#6059
opened Jul 24, 2026 by
q10
Contributor
Loading…
Switch sparse_permute_1d to OverflowOnly cap (follow-up to D104903707)
cla signed
meta-exported
#6057
opened Jul 23, 2026 by
q10
Contributor
Loading…
Fix HIP 2^32 grid-overflow in batch_index_select_dim0 forward kernels (#6053)
cla signed
meta-exported
#6053
opened Jul 23, 2026 by
cx-yin
Loading…
Enable Pyrefly in fbcode/deeplearning/fbgemm
cla signed
meta-exported
#6048
opened Jul 22, 2026 by
maggiemoss
Loading…
Make empty PT2 wrapper TUs byte-distinct per optimizer
cla signed
#6044
opened Jul 22, 2026 by
jithunnair-amd
Contributor
•
Draft
Skip codegen of empty PT2 wrappers for optimizers without backend support
cla signed
#6043
opened Jul 22, 2026 by
jithunnair-amd
Contributor
•
Draft
Bounds-assert table_idx and row byte range in embedding_inplace_update kernels (#6035)
cla signed
meta-exported
#6035
opened Jul 21, 2026 by
iuliur-meta
Loading…
Apply ROCm grid-overflow cap hygiene to _float_to_fusednbitrowwise_cuda_kernel
ciflow/rocm
cla signed
meta-exported
module: rocm
#6026
opened Jul 16, 2026 by
q10
Contributor
Loading…
Skip jagged_softmax large-grid tests on ROCm (exceed HIP 2^32 launch limit)
ciflow/rocm
cla signed
meta-exported
module: rocm
#6022
opened Jul 16, 2026 by
gchalump
Contributor
Loading…
[ROCm] Add mixed precision support to the HIP optimized warp per row TBE backward kernel
ciflow/rocm
cla signed
module: rocm
#6020
opened Jul 16, 2026 by
avbokovoy
Contributor
Loading…
Backout "Revert D95859348" (#6015)
ci-no-td
cla signed
meta-exported
#6015
opened Jul 15, 2026 by
pklund
Loading…
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.