Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
53 changes: 40 additions & 13 deletions .github/workflows/cuda_ext_check_before_merge.yml
Original file line number Diff line number Diff line change
Expand Up @@ -30,25 +30,52 @@ jobs:
matrix: ${{fromJson(needs.matrix_preparation.outputs.matrix)}}
container:
image: ${{ matrix.build.cuda_image }}
options: --gpus all --rm
options: --rm
env:
CUDA_VISIBLE_DEVICES: ''
NVIDIA_VISIBLE_DEVICES: void
FORCE_CUDA: '1'
COLOSSAL_CPU_ARCH: x86-64
MAX_JOBS: '2'
steps:
- uses: actions/checkout@v2

- name: Install PyTorch
run: eval ${{ matrix.build.torch_command }}

- name: Download cub for CUDA 10.2
- name: Build wheel without a GPU
run: |
CUDA_VERSION=$(nvcc -V | awk -F ',| ' '/release/{print $6}')
pip install setuptools wheel ninja packaging
python -c "import torch; assert not torch.cuda.is_available()"
unset TORCH_CUDA_ARCH_LIST
BUILD_EXT=1 pip wheel --no-deps --no-build-isolation --wheel-dir dist .

# check if it is CUDA 10.2
# download cub
if [ "$CUDA_VERSION" = "10.2" ]; then
wget https://github.com/NVIDIA/cub/archive/refs/tags/1.8.0.zip
unzip 1.8.0.zip
cp -r cub-1.8.0/cub/ colossalai/kernel/cuda_native/csrc/kernels/include/
fi

- name: Build
- name: Check wheel contains Ampere and Hopper kernels
run: |
BUILD_EXT=1 pip install -v -e .
python - <<'PY'
import re
import subprocess
import tempfile
import zipfile
from pathlib import Path

cuda_modules = {
"layernorm_cuda", "moe_cuda", "fused_optim_cuda", "inference_ops_cuda",
"scaled_masked_softmax_cuda", "scaled_upper_triangle_masked_softmax_cuda",
}
wheels = list(Path("dist").glob("*.whl"))
assert len(wheels) == 1, wheels
with zipfile.ZipFile(wheels[0]) as wheel, tempfile.TemporaryDirectory() as temp:
modules = {
Path(name).name.split(".")[0]: name for name in wheel.namelist()
if name.startswith("colossalai/_C/") and name.endswith(".so")
}
assert set(modules) == cuda_modules | {"cpu_adam_x86"}, modules
for name in sorted(cuda_modules):
binary = Path(temp) / f"{name}.so"
binary.write_bytes(wheel.read(modules[name]))
output = subprocess.check_output(["cuobjdump", "--list-elf", str(binary)], text=True)
arches = set(re.findall(r"\bsm_\d+[af]?\b", output))
assert {"sm_80", "sm_90"} <= arches, f"{name}: {sorted(arches)}"
print(f"{name}: {sorted(arches)}")
PY
2 changes: 1 addition & 1 deletion .github/workflows/doc_check_on_pr.yml
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ jobs:

- uses: actions/setup-python@v2
with:
python-version: "3.9"
python-version: "3.10"

# we use the versions in the main branch as the guide for versions to display
# checkout will give your merged branch
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/release_nightly_on_schedule.yml
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ jobs:

- uses: actions/setup-python@v2
with:
python-version: '3.9'
python-version: '3.10'

- run: |
python .github/workflows/scripts/update_setup_for_nightly.py
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/release_pypi_after_merge.yml
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ jobs:

- uses: actions/setup-python@v2
with:
python-version: '3.9'
python-version: '3.10'

- run: python setup.py sdist build

Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/release_test_pypi_before_merge.yml
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ jobs:

- uses: actions/setup-python@v2
with:
python-version: '3.9'
python-version: '3.10'

- name: add timestamp to the version
id: prep-version
Expand Down
8 changes: 4 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -468,7 +468,7 @@ Please visit our [documentation](https://www.colossalai.org/) and [examples](htt
## Installation

Requirements:
- PyTorch >= 2.2
- 2.2 <= PyTorch <= 2.5.1
- Python >= 3.10 and < 3.13
- CUDA >= 11.0
- [NVIDIA GPU Compute Capability](https://developer.nvidia.com/cuda-gpus) >= 7.0 (V100/RTX20 and higher)
Expand All @@ -489,7 +489,7 @@ pip install colossalai
However, if you want to build the PyTorch extensions during installation, you can set `BUILD_EXT=1`.

```bash
BUILD_EXT=1 pip install colossalai
BUILD_EXT=1 pip install --no-binary=colossalai --no-build-isolation colossalai
```

**Otherwise, CUDA kernels will be built during runtime when you actually need them.**
Expand Down Expand Up @@ -517,7 +517,7 @@ By default, we do not compile CUDA/C++ kernels. ColossalAI will build them durin
If you want to install and enable CUDA kernel fusion (compulsory installation when using fused optimizer):

```shell
BUILD_EXT=1 pip install .
BUILD_EXT=1 pip install --no-build-isolation .
```

For Users with CUDA 10.2, you can still build ColossalAI from source. However, you need to manually download the cub library and copy it to the corresponding directory.
Expand All @@ -533,7 +533,7 @@ unzip 1.8.0.zip
cp -r cub-1.8.0/cub/ colossalai/kernel/cuda_native/csrc/kernels/include/

# install
BUILD_EXT=1 pip install .
BUILD_EXT=1 pip install --no-build-isolation .
```

<p align="right">(<a href="#top">back to top</a>)</p>
Expand Down
7 changes: 3 additions & 4 deletions colossalai/booster/plugin/hybrid_parallel_plugin.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,14 +29,15 @@
from colossalai.interface.model import PeftUnwrapMixin
from colossalai.interface.optimizer import DistributedOptim
from colossalai.logging import get_dist_logger
from colossalai.nn.optimizer import DistGaloreAwamW, cast_to_distributed
from colossalai.nn.optimizer import DistGaloreAdamW, cast_to_distributed
from colossalai.pipeline.schedule import InterleavedSchedule, OneForwardOneBackwardSchedule, ZeroBubbleVPipeScheduler
from colossalai.pipeline.stage_manager import PipelineStageManager
from colossalai.quantization import BnbQuantizationConfig, quantize_model
from colossalai.quantization.fp8_hook import FP8Hook
from colossalai.shardformer import GradientCheckpointConfig, ShardConfig, ShardFormer
from colossalai.shardformer.layer.utils import SeqParallelUtils, is_share_sp_tp
from colossalai.shardformer.policies.base_policy import Policy
from colossalai.shardformer.shard.shard_config import SUPPORT_SP_MODE
from colossalai.tensor.colo_parameter import ColoParameter
from colossalai.tensor.d_tensor.api import is_distributed_tensor
from colossalai.tensor.param_op_hook import ColoParamOpHookManager
Expand All @@ -45,8 +46,6 @@

from .pp_plugin_base import PipelinePluginBase

SUPPORT_SP_MODE = ["split_gather", "ring", "all_to_all", "ring_attn"]

PRECISION_TORCH_TYPE = {"fp16": torch.float16, "fp32": torch.float32, "bf16": torch.bfloat16}


Expand Down Expand Up @@ -1298,7 +1297,7 @@ def configure(

# Replace with distributed implementation if exists
optimizer = cast_to_distributed(optimizer)
if isinstance(optimizer, DistGaloreAwamW) and zero_stage > 0 and self.dp_size > 0:
if isinstance(optimizer, DistGaloreAdamW) and zero_stage > 0 and self.dp_size > 0:
self.logger.warning(
"Galore is only supported for Tensor Parallel and vanilla Data Parallel yet. Disabling ZeRO.",
ranks=[0],
Expand Down
4 changes: 2 additions & 2 deletions colossalai/booster/plugin/low_level_zero_plugin.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@
from colossalai.interface import AMPModelMixin, ModelWrapper, OptimizerWrapper
from colossalai.interface.optimizer import DistributedOptim
from colossalai.logging import get_dist_logger
from colossalai.nn.optimizer import DistGaloreAwamW, cast_to_distributed
from colossalai.nn.optimizer import DistGaloreAdamW, cast_to_distributed
from colossalai.quantization import BnbQuantizationConfig, quantize_model
from colossalai.quantization.fp8_hook import FP8Hook
from colossalai.tensor.colo_parameter import ColoParameter
Expand Down Expand Up @@ -595,7 +595,7 @@ def configure(
# Replace with the distributed implementation if exists
optimizer = cast_to_distributed(optimizer)

if isinstance(optimizer, DistGaloreAwamW) and zero_stage > 0 and dp_size > 0:
if isinstance(optimizer, DistGaloreAdamW) and zero_stage > 0 and dp_size > 0:
self.logger.warning(
"Galore is only supported for Tensor Parallel and vanilla Data Parallel yet. Disabling ZeRO.",
ranks=[0],
Expand Down
3 changes: 1 addition & 2 deletions colossalai/booster/plugin/moe_hybrid_parallel_plugin.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,6 @@

from colossalai.booster.plugin.hybrid_parallel_plugin import (
PRECISION_TORCH_TYPE,
SUPPORT_SP_MODE,
HybridParallelAMPOptimizer,
HybridParallelModule,
HybridParallelNaiveOptimizer,
Expand All @@ -32,7 +31,7 @@
from colossalai.pipeline.stage_manager import PipelineStageManager
from colossalai.shardformer.policies.base_policy import Policy
from colossalai.shardformer.shard.grad_ckpt_config import GradientCheckpointConfig
from colossalai.shardformer.shard.shard_config import ShardConfig
from colossalai.shardformer.shard.shard_config import SUPPORT_SP_MODE, ShardConfig
from colossalai.tensor.moe_tensor.api import is_moe_tensor


Expand Down
2 changes: 1 addition & 1 deletion colossalai/inference/modeling/policy/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,5 +18,5 @@
"GlideLlamaModelPolicy",
"StableDiffusion3InferPolicy",
"PixArtAlphaInferPolicy",
"model_polic_map",
"model_policy_map",
]
5 changes: 3 additions & 2 deletions colossalai/kernel/kernel_loader.py
Original file line number Diff line number Diff line change
Expand Up @@ -68,12 +68,13 @@ def load(self, ext_name: str = None):
usable_exts = []
for ext in exts:
if ext.is_available():
# make sure the machine is compatible during kernel loading
ext.assert_compatible()
usable_exts.append(ext)

assert len(usable_exts) != 0, f"No usable kernel found for {self.__class__.__name__} on the current machine."

for ext in usable_exts:
ext.assert_compatible()

if len(usable_exts) > 1:
# if more than one usable kernel is found, we will try to load the kernel with the highest priority
usable_exts = sorted(usable_exts, key=lambda ext: ext.priority, reverse=True)
Expand Down
30 changes: 30 additions & 0 deletions colossalai/moe/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
"""MoE operations, loaded on demand to keep package imports lightweight."""

from importlib import import_module

__all__ = [
"AllGather",
"AllToAll",
"DPGradScalerIn",
"DPGradScalerOut",
"EPGradScalerIn",
"EPGradScalerOut",
"HierarchicalAllToAll",
"MoeCombine",
"MoeDispatch",
"ReduceScatter",
"all_to_all_uneven",
"moe_cumsum",
]


def __getattr__(name):
if name in __all__:
operation = getattr(import_module("._operation", __name__), name)
globals()[name] = operation
return operation
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")


def __dir__():
return sorted(set(globals()) | set(__all__))
5 changes: 3 additions & 2 deletions colossalai/nn/optimizer/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
from .cpu_adam import CPUAdam
from .distributed_adafactor import DistributedAdaFactor
from .distributed_came import DistributedCAME
from .distributed_galore import DistGaloreAwamW
from .distributed_galore import DistGaloreAdamW, DistGaloreAwamW
from .distributed_lamb import DistributedLamb
from .fused_adam import FusedAdam
from .fused_lamb import FusedLAMB
Expand All @@ -27,6 +27,7 @@
"CPUAdam",
"HybridAdam",
"DistributedLamb",
"DistGaloreAdamW",
"DistGaloreAwamW",
"GaLoreAdamW",
"GaLoreAdafactor",
Expand All @@ -38,7 +39,7 @@
]

optim2DistOptim = {
GaLoreAdamW8bit: DistGaloreAwamW,
GaLoreAdamW8bit: DistGaloreAdamW,
Lamb: DistributedLamb,
CAME: DistributedCAME,
Adafactor: DistributedAdaFactor,
Expand Down
8 changes: 6 additions & 2 deletions colossalai/nn/optimizer/distributed_galore.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,11 +14,11 @@

from .galore import GaLoreProjector, make_low_rank_buffer

__all__ = ["DistributedGalore"]
__all__ = ["DistGaloreAdamW", "DistGaloreAwamW"]
# Mark sharded dimension


class DistGaloreAwamW(DistributedOptim, Optimizer2State):
class DistGaloreAdamW(DistributedOptim, Optimizer2State):
r"""Implements Galore, a optimizer-agonistic gradient compression technique on 8-bit AdamW.
It largely compresses gradient via low-rank projection and is claimed to be insensitive to hyperparams like lr.
Supports Tensor Parallel and ZeRO stage 1 and 2 via booster and plugin.
Expand Down Expand Up @@ -280,3 +280,7 @@ def __del__(self):
for p in group["params"]:
if hasattr(p, "saved_data"):
del p.saved_data


# Keep the original misspelling importable for existing users and serialized objects.
DistGaloreAwamW = DistGaloreAdamW
8 changes: 4 additions & 4 deletions docs/source/en/get_started/installation.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Setup

Requirements:
- PyTorch >= 2.1
- 2.2 <= PyTorch <= 2.5.1
- Python >= 3.10 and < 3.13
- CUDA >= 11.0
- [NVIDIA GPU Compute Capability](https://developer.nvidia.com/cuda-gpus) >= 7.0 (V100/RTX20 and higher)
Expand All @@ -23,7 +23,7 @@ pip install colossalai
If you want to build PyTorch extensions during installation, you can use the command below. Otherwise, the PyTorch extensions will be built during runtime.

```shell
BUILD_EXT=1 pip install colossalai
BUILD_EXT=1 pip install --no-binary=colossalai --no-build-isolation colossalai
```


Expand All @@ -39,7 +39,7 @@ cd ColossalAI
pip install -r requirements/requirements.txt

# install colossalai
BUILD_EXT=1 pip install .
BUILD_EXT=1 pip install --no-build-isolation .
```

If you don't want to install and enable CUDA kernel fusion (compulsory installation when using fused optimizer), just don't specify the `BUILD_EXT`:
Expand All @@ -61,7 +61,7 @@ unzip 1.8.0.zip
cp -r cub-1.8.0/cub/ colossalai/kernel/cuda_native/csrc/kernels/include/

# install
BUILD_EXT=1 pip install .
BUILD_EXT=1 pip install --no-build-isolation .
```

<!-- doc-test-command: echo "installation.md does not need test" -->
8 changes: 4 additions & 4 deletions docs/source/zh-Hans/get_started/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

环境要求:

- PyTorch >= 2.1
- 2.2 <= PyTorch <= 2.5.1
- Python >= 3.10 且 < 3.13
- CUDA >= 11.0
- [NVIDIA GPU Compute Capability](https://developer.nvidia.com/cuda-gpus) >= 7.0 (V100/RTX20 and higher)
Expand All @@ -23,7 +23,7 @@ pip install colossalai
如果你想同时安装PyTorch扩展的话,可以添加`BUILD_EXT=1`。如果不添加的话,PyTorch扩展会在运行时自动安装。

```shell
BUILD_EXT=1 pip install colossalai
BUILD_EXT=1 pip install --no-binary=colossalai --no-build-isolation colossalai
```

## 从源安装
Expand All @@ -38,7 +38,7 @@ cd ColossalAI
pip install -r requirements/requirements.txt

# install colossalai
BUILD_EXT=1 pip install .
BUILD_EXT=1 pip install --no-build-isolation .
```

如果您不想安装和启用 CUDA 内核融合(使用融合优化器时强制安装),您可以不添加`BUILD_EXT=1`:
Expand All @@ -60,7 +60,7 @@ unzip 1.8.0.zip
cp -r cub-1.8.0/cub/ colossalai/kernel/cuda_native/csrc/kernels/include/

# install
BUILD_EXT=1 pip install .
BUILD_EXT=1 pip install --no-build-isolation .
```

<!-- doc-test-command: echo "installation.md does not need test" -->
2 changes: 1 addition & 1 deletion examples/images/dreambooth/requirements.txt
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
diffusers>==0.5.0
diffusers>=0.5.0
accelerate
torchvision
transformers>=4.21.0
Expand Down
Loading
Loading