Skip to content

New merge method API - #706

Merged
cg123 merged 36 commits into
mainfrom
mmapi-9-26
Sep 12, 2026
Merged

cg123 merged 36 commits into
mainfrom
mmapi-9-26

Conversation

@cg123

@cg123 cg123 commented Sep 12, 2026 •

Copy link
Copy Markdown
Collaborator

Note

High Risk
Large refactor of merge execution, dtype promotion, and vocabulary metadata touches every weight merge path; misaligned vocab or truncation options can silently produce invalid checkpoints.

Overview
Introduces a unified merge-method execution model centered on TensorGroup / TensorBatch, with public merge_tensors and merge_state_dicts for in-memory merges (batched floating weights, strict buffer handling, dtype / out_dtype, and BatchOptions).

Built-in and custom methods are refactored around MergeMethodSpec, @merge_method (group vs batch kernels), register() / get(), and shared parameter binding (per-input maps, PerGroupValues, typed YAML resolution via evaluate_setting). Graph-facing code is adapted to method.spec metadata and optional-tensor policies instead of per-method Task classes where migrated (e.g. Arcee Fusion).

Vocabulary handling replaces WeightInfo.is_embed with vocabulary_axis across architecture JSON and auto-detection (infer_vocabulary_axes); several arch defs drop legacy tensor-space/head-split metadata. Docs expand tokenizer alignment, computation precision, and --unsafe-truncate-embeddings.

ConfigReader no longer resolves parameters per tensor name inline—it exposes parameter_sources for the planner to resolve with filters. Minor: float64 dtype in config, mergekit index version 0.1.5.

Reviewed by Cursor Bugbot for commit 34972dd. Bugbot is set up for automated code reviews on this repo. Configure here.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 53757a3. Configure here.

Comment thread mergekit/plan.py
parameters=global_params,
input_parameters=tensor_params,
base_model=base_model,
output_weight=weight,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

YAML dtype recasts non-floating buffers

Medium Severity

The YAML planner never passes dtype or out_dtype into ExecuteMergeMethodTask, so input casting happens in GatherTensors and output casting in SaveTensor. Both apply to every tensor, including integer or boolean buffers. That defeats copy_non_floating_buffer and can numerically merge or recast BatchNorm-style counters, contrary to the new API contract that non-floating buffers stay unchanged. The raw-PyTorch path already avoids this by casting only floating tensors at load and applying out_dtype inside the merge task after buffers are copied.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 53757a3. Configure here.

Comment thread mergekit/scripts/merge_raw_pytorch.py
@cg123
cg123 merged commit 9eeb539 into main Sep 12, 2026
12 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Sep 12, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant