Hi, and thanks for open-sourcing the MiniMax-H3 work.
I'd like to reproduce the training adapter itself — https://modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter — not the downstream LoRAs.
What I've already found in the repo
examples/minimax_h3/model_training/train.py and the two-stage sft:data_process → sft:train workflow.
- The
.sh scripts under examples/minimax_h3/model_training/lora/ (e.g. MiniMax-H3-Pruned-FL2VA.sh, the Ref2VA variants). These train downstream LoRAs and reference the adapter only as an optional --preset_lora_path / --preset_lora_model "dit" to fuse in during training.
- The base weights (
Comfy-Org/MiniMax-H3, MiniMax/MiniMax-H3) and the MiniMax-H3-Self-Generated-Dataset.
What I can't find
A script or config that produces the adapter itself (the rank-64 DeCFG LoRA, FL2VA and Ref2VA variants) from the CFG-distilled base + the self-generated dataset. The example scripts consume the adapter rather than create it.
Could you share the following for the adapter's own training?
- The exact training script / command (equivalent
.sh) used to produce model_for_comfy_dit.safetensors and model_ref2va_for_comfy_dit.safetensors.
- The DeCFG / "differential training" details — what the objective is and how it differs from standard SFT LoRA training (loss formulation, any CFG/DeCFG-specific handling, timestep sampling/shift).
- Hyperparameters: LoRA rank (64?) and target modules, learning rate, batch size, number of steps/epochs,
dataset_repeat, resolution / num_frames, audio_loss_weight, optimizer/schedule, and seed if fixed.
- Which subset/split of MiniMax-H3-Self-Generated-Dataset was used, and whether the released dataset is the complete training set or a sample.
- Base checkpoint used as the starting point (the pruned bf16 DiT, or another), plus any pre/post-processing of the adapter weights (e.g. the ComfyUI qkv layout conversion).
Happy to open a PR to add a reproduction script/README under examples/minimax_h3/ if that's easier on your side. Thanks!
Hi, and thanks for open-sourcing the MiniMax-H3 work.
I'd like to reproduce the training adapter itself — https://modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter — not the downstream LoRAs.
What I've already found in the repo
examples/minimax_h3/model_training/train.pyand the two-stagesft:data_process→sft:trainworkflow..shscripts underexamples/minimax_h3/model_training/lora/(e.g.MiniMax-H3-Pruned-FL2VA.sh, the Ref2VA variants). These train downstream LoRAs and reference the adapter only as an optional--preset_lora_path/--preset_lora_model "dit"to fuse in during training.Comfy-Org/MiniMax-H3,MiniMax/MiniMax-H3) and the MiniMax-H3-Self-Generated-Dataset.What I can't find
A script or config that produces the adapter itself (the rank-64 DeCFG LoRA, FL2VA and Ref2VA variants) from the CFG-distilled base + the self-generated dataset. The example scripts consume the adapter rather than create it.
Could you share the following for the adapter's own training?
.sh) used to producemodel_for_comfy_dit.safetensorsandmodel_ref2va_for_comfy_dit.safetensors.dataset_repeat, resolution /num_frames,audio_loss_weight, optimizer/schedule, and seed if fixed.Happy to open a PR to add a reproduction script/README under
examples/minimax_h3/if that's easier on your side. Thanks!