Repository navigation
New merge method API - #706
Conversation
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 53757a3. Configure here.
| parameters=global_params, | ||
| input_parameters=tensor_params, | ||
| base_model=base_model, | ||
| output_weight=weight, |
There was a problem hiding this comment.
YAML dtype recasts non-floating buffers
Medium Severity
The YAML planner never passes dtype or out_dtype into ExecuteMergeMethodTask, so input casting happens in GatherTensors and output casting in SaveTensor. Both apply to every tensor, including integer or boolean buffers. That defeats copy_non_floating_buffer and can numerically merge or recast BatchNorm-style counters, contrary to the new API contract that non-floating buffers stay unchanged. The raw-PyTorch path already avoids this by casting only floating tensors at load and applying out_dtype inside the merge task after buffers are copied.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit 53757a3. Configure here.


Note
High Risk
Large refactor of merge execution, dtype promotion, and vocabulary metadata touches every weight merge path; misaligned vocab or truncation options can silently produce invalid checkpoints.
Overview
Introduces a unified merge-method execution model centered on
TensorGroup/TensorBatch, with publicmerge_tensorsandmerge_state_dictsfor in-memory merges (batched floating weights, strict buffer handling,dtype/out_dtype, andBatchOptions).Built-in and custom methods are refactored around
MergeMethodSpec,@merge_method(group vs batch kernels),register()/get(), and shared parameter binding (per-input maps,PerGroupValues, typed YAML resolution viaevaluate_setting). Graph-facing code is adapted tomethod.specmetadata and optional-tensor policies instead of per-methodTaskclasses where migrated (e.g. Arcee Fusion).Vocabulary handling replaces
WeightInfo.is_embedwithvocabulary_axisacross architecture JSON and auto-detection (infer_vocabulary_axes); several arch defs drop legacy tensor-space/head-split metadata. Docs expand tokenizer alignment, computation precision, and--unsafe-truncate-embeddings.ConfigReader no longer resolves parameters per tensor name inline—it exposes
parameter_sourcesfor the planner to resolve with filters. Minor:float64dtype in config, mergekit index version 0.1.5.Reviewed by Cursor Bugbot for commit 34972dd. Bugbot is set up for automated code reviews on this repo. Configure here.