Repository navigation
Conversation
…cee-ai#708) A shared role-key space so parameters from different architecture families can be addressed uniformly before any merge method runs: model.layers.0.self_attn.q_proj.weight -> layer_0.attn_q.weight model.layers.11.mlp.up_proj.weight -> layer_11.ffn_up.weight Dependency-free name mapping only: values pass through by reference, no tensor ops, no new requirements. Ships the Llama-family detector (Llama, Qwen, Mistral, Gemma, TinyLlama layouts) with detection, canonicalization, and an exact decanonicalize inverse (round-trip tested). Remaining families (gpt2, bert, neox, opt, t5, phi) and the shape bridge follow in the same pattern per arcee-ai#708. The detector set is exercised end-to-end by the cross-family merge pipeline behind the published Optitransfer converge collectives (9 models across 4 architecture families in one set of weights). Signed-off-by: Ryan Gillespie <mgillr@users.noreply.github.com>
|
All contributors have signed the CLA ✍️ ✅ |
|
To make the integration discussion concrete, here is the wiring sketch for the two remaining increments — happy to adjust to whatever home the maintainers prefer. P1b (next PR): the remaining pure-name detectors — P2 (the shape bridge): the piece that makes
Config surface: models:
- model: Qwen/Qwen2.5-7B-Instruct
- model: allenai/Llama-3.1-Tulu-3-8B-SFT
canonicalize: true
bridge: pad # or: procrustes
merge_method: ties # any existing methodOpen question for maintainers: should |
|
I have read the CLA Document and I hereby sign the CLA |
First increment of #708: the canonical role-key space and the Llama-family detector.
mergekit/architecture/canonical.pymaps native parameter names into one schema shared across architecture families, so parameters can be addressed uniformly before any merge method runs:Scope, deliberately narrow for a first PR:
detect,canonicalize, and an exactdecanonicalizeinverse — round-trip tested on key-set and value identity over the mappable subset.The detector pattern is exercised end-to-end by the cross-family merge pipeline behind the published Optitransfer converge collectives — 9 models across 4 architecture families merged into one set of weights (28.5 min on one A100, streaming).
Note
Low Risk
Additive, dependency-free name mapping with tests; it does not hook into existing merge pipelines in this diff.
Overview
Introduces
mergekit/architecture/canonical.py, the first slice of cross-architecture merging (#708): a shared canonical role-key namespace (e.g.layer_0.attn_q.weight,embed_tokens.weight) so checkpoints from different families can be addressed the same way before merge logic runs.LlamaFamilyDetectoris included for Llama-style layouts (Llama, Qwen, Mistral, Gemma, etc.):detect(rejects OPTmodel.decoder.*and non-Llama keys),canonicalize(name-only remap; tensors unchanged; drops unmapped keys like rotary buffers), anddecanonicalizefor a round-trip back to native names. Helperslayer_key,global_key, andis_canonicalvalidate the schema. No tensor ops or new runtime dependencies.tests/test_canonical.pycovers detection boundaries, role mapping, dropped buffers, round-trip identity on the mappable subset, and a fallback import path when the full package isn’t available. Other families and shape bridging are explicitly out of scope for this PR.Reviewed by Cursor Bugbot for commit cd49106. Bugbot is set up for automated code reviews on this repo. Configure here.