Skip to content

fix(patcher): resolve MP+DDP device_map memory through the active accelerator - #10292

Open
li-lizhe wants to merge 2 commits into
modelscope:mainfrom
li-lizhe:fix/mp-ddp-max-memory-active-accelerator
Open

li-lizhe wants to merge 2 commits into
modelscope:mainfrom
li-lizhe:fix/mp-ddp-max-memory-active-accelerator

Conversation

@li-lizhe

@li-lizhe li-lizhe commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

PR type

  • Bug Fix
  • New Feature
  • Document Updates
  • More Models or Datasets Support

PR information

swift/model/patcher.py::_get_max_memory() (called from patch_mp_ddp() for the device_map + DDP path) pins the device handling to the CUDA namespace:

  • it warms up each device with a bare integer device id — torch.tensor([0], device=i) — which only resolves under CUDA, and
  • it reads per-device memory with torch.cuda.mem_get_info(i).

On a non-CUDA accelerator (Ascend NPU / XPU / MUSA) those calls abort inside _infer_auto_device_map_patch before the model is sharded, so MP + device_map cannot run there:

AssertionError: Torch not compiled with CUDA enabled

This resolves the accelerator namespace through helpers the package already exposes — get_torch_device() and get_device(i) — mirroring swift/utils/torch_utils.py::get_max_reserved_memory(), which already dispatches via get_torch_device(). CUDA behaviour is unchanged: get_device(i) returns cuda:i and get_torch_device() returns torch.cuda, i.e. the same calls as before. A unit test is added next to the analogous tests/utils/test_max_reserved_memory.py.

Experiment results

Real Ascend NPU (910, CANN 9.2.0-beta.2, torch 2.15.0.dev20260917+cpu, torch_npu 2.15.0.dev20260917+gitec69335, 1 card). The real _get_max_memory body was extracted from the source via AST and executed with get_device_count mapped to torch.npu.device_count:

_get_max_memory([0])
upstream AssertionError: Torch not compiled with CUDA enabled
patched {0: 31285399552, 'cpu': 2140501688320} — max_memory[0] equals torch.npu.mem_get_info(0)[0]

_get_max_memory([]) still returns {0: 0, 'cpu': ...} (device not in this shard ⇒ 0).

Unit test tests/general/test_mp_ddp_max_memory.py: passes against the patched source; against the upstream body it fails on CPU-only torch (RuntimeError: Cannot access accelerator device when none is available.). Formatting follows the repo gates (flake8, isort, yapf 0.43.0 — all clean on the touched files).

Not tested: CUDA / XPU / MUSA paths (no such hardware here); a full end-to-end MP + device_map torchrun run was not exercised — only the changed helper.

Notes

  • No linked issue.

…elerator

`_get_max_memory()` (used by `patch_mp_ddp` for the device_map + DDP path)
queried per-device memory with `torch.cuda.mem_get_info` and warmed the devices
up with a bare integer device id. On a non-CUDA accelerator (Ascend NPU, XPU,
MUSA) the CUDA-only call raises `AssertionError: Torch not compiled with CUDA
enabled`, so `infer_auto_device_map` aborts before the model is sharded.

Resolve the accelerator namespace with the helpers this package already exposes
(`get_torch_device()` / `get_device()`), the same way
`torch_utils.get_max_reserved_memory` does, and cover it with a unit test.

Signed-off-by: li-lizhe <147392333@qq.com>
… real device ordinal

Signed-off-by: li-lizhe <147392333@qq.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant