Skip to content

architecture: canonical parameter-name mapping (first increment of #708) - #709

Open
mgillr wants to merge 1 commit into
arcee-ai:mainfrom
mgillr:canonical-role-keys
Open

mgillr wants to merge 1 commit into
arcee-ai:mainfrom
mgillr:canonical-role-keys

Conversation

@mgillr

@mgillr mgillr commented Sep 21, 2026 •

Copy link
Copy Markdown

First increment of #708: the canonical role-key space and the Llama-family detector.

mergekit/architecture/canonical.py maps native parameter names into one schema shared across architecture families, so parameters can be addressed uniformly before any merge method runs:

model.layers.0.self_attn.q_proj.weight  ->  layer_0.attn_q.weight
model.layers.11.mlp.up_proj.weight      ->  layer_11.ffn_up.weight
model.embed_tokens.weight               ->  embed_tokens.weight
model.norm.weight                       ->  final_norm.weight

Scope, deliberately narrow for a first PR:

  • Name mapping only. Values pass through by reference; no tensor ops, no new dependencies, no model downloads in tests.
  • The Llama-family detector (Llama, Qwen, Mistral, Gemma, TinyLlama layouts) with detect, canonicalize, and an exact decanonicalize inverse — round-trip tested on key-set and value identity over the mappable subset.
  • Remaining families (gpt2, bert, neox, opt, t5, phi) and the shape bridge (pad/truncate + optional per-key alignment) follow in the same pattern per Cross-architecture merging: proposal — canonicalization as a preprocessing transform #708, with this module as the settled home.

The detector pattern is exercised end-to-end by the cross-family merge pipeline behind the published Optitransfer converge collectives — 9 models across 4 architecture families merged into one set of weights (28.5 min on one A100, streaming).


Note

Low Risk
Additive, dependency-free name mapping with tests; it does not hook into existing merge pipelines in this diff.

Overview
Introduces mergekit/architecture/canonical.py, the first slice of cross-architecture merging (#708): a shared canonical role-key namespace (e.g. layer_0.attn_q.weight, embed_tokens.weight) so checkpoints from different families can be addressed the same way before merge logic runs.

LlamaFamilyDetector is included for Llama-style layouts (Llama, Qwen, Mistral, Gemma, etc.): detect (rejects OPT model.decoder.* and non-Llama keys), canonicalize (name-only remap; tensors unchanged; drops unmapped keys like rotary buffers), and decanonicalize for a round-trip back to native names. Helpers layer_key, global_key, and is_canonical validate the schema. No tensor ops or new runtime dependencies.

tests/test_canonical.py covers detection boundaries, role mapping, dropped buffers, round-trip identity on the mappable subset, and a fallback import path when the full package isn’t available. Other families and shape bridging are explicitly out of scope for this PR.

Reviewed by Cursor Bugbot for commit cd49106. Bugbot is set up for automated code reviews on this repo. Configure here.

…cee-ai#708)

A shared role-key space so parameters from different architecture
families can be addressed uniformly before any merge method runs:

  model.layers.0.self_attn.q_proj.weight -> layer_0.attn_q.weight
  model.layers.11.mlp.up_proj.weight     -> layer_11.ffn_up.weight

Dependency-free name mapping only: values pass through by reference,
no tensor ops, no new requirements. Ships the Llama-family detector
(Llama, Qwen, Mistral, Gemma, TinyLlama layouts) with detection,
canonicalization, and an exact decanonicalize inverse (round-trip
tested). Remaining families (gpt2, bert, neox, opt, t5, phi) and the
shape bridge follow in the same pattern per arcee-ai#708.

The detector set is exercised end-to-end by the cross-family merge
pipeline behind the published Optitransfer converge collectives
(9 models across 4 architecture families in one set of weights).

Signed-off-by: Ryan Gillespie <mgillr@users.noreply.github.com>
@github-actions

github-actions Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@mgillr

mgillr commented Sep 21, 2026

Copy link
Copy Markdown
Author

To make the integration discussion concrete, here is the wiring sketch for the two remaining increments — happy to adjust to whatever home the maintainers prefer.

P1b (next PR): the remaining pure-name detectors — gpt2, bert, neox, opt — identical shape to LlamaFamilyDetector. (phi and t5 involve fused-QKV splits and encoder/decoder prefixes respectively; slightly larger but same pattern.)

P2 (the shape bridge): the piece that makes --allow_crimes unnecessary. Sketch of the hook in parameter resolution, when canonicalize: true is set on a model entry:

  1. each model's tensors are canonicalized at load (detector auto-selected, as in P1);
  2. donors whose canonical shapes differ from the anchor's (e.g. hidden 4096 vs 3584) pass through the bridge: zero-pad/truncate to anchor geometry, optionally followed by per-key alignment (SVD/Procrustes) — measured cost ~30 min pad-only, 2-3 h with alignment, per 7-8B model on one A100;
  3. existing merge methods then run on the bridged groups unchanged — no method signature changes.

Config surface:

models:
  - model: Qwen/Qwen2.5-7B-Instruct
  - model: allenai/Llama-3.1-Tulu-3-8B-SFT
    canonicalize: true
    bridge: pad          # or: procrustes
merge_method: ties       # any existing method

Open question for maintainers: should canonicalize be per-model (as sketched, since one side may already be canonical) or a merge-level flag? I'd lean per-model with merge-level default-false.

@mgillr

mgillr commented Sep 21, 2026

Copy link
Copy Markdown
Author

I have read the CLA Document and I hereby sign the CLA

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant