Skip to content

architecture: bert and opt detectors (second increment of #708) - #710

Open
mgillr wants to merge 4 commits into
arcee-ai:mainfrom
mgillr:p1b-detectors
Open

mgillr wants to merge 4 commits into
arcee-ai:mainfrom
mgillr:p1b-detectors

Conversation

@mgillr

@mgillr mgillr commented Sep 21, 2026 •

Copy link
Copy Markdown

Stacks on #709 (contains its commits; review the second commit for the delta).

Adds two more pure-name detectors in the same pattern:

  • BertFamilyDetector — BERT and RoBERTa encoder layouts; post-norm LayerNorms map onto the post_attn_norm / post_ffn_norm canonical slots (schema extended accordingly, plus embed_positions and the norm/head bias globals).
  • OPTDetector — the OPT decoder layout (model.decoder.layers.N.*), distinguished from the Llama layout that architecture: canonical parameter-name mapping (first increment of #708) #709's detector already rejects.

Tests: per-family mapping with value identity, all outputs pass is_canonical, and cross-detector exclusivity (a Llama-layout dict matches neither new detector). gpt2/neox/phi remain for the third increment — they need fused-QKV splits, which are tensor ops and belong with the shape-bridge increment rather than this dependency-free module.


Note

Low Risk
Additive, dependency-free name remapping with unit tests only; no changes to existing merge execution paths.

Overview
Extends the cross-architecture canonical name space (#708) with two more pure-name detectors alongside the existing Llama-family mapper: BertFamilyDetector (BERT/RoBERTa encoder keys, including post-norm LayerNorms mapped to post_attn_norm / post_ffn_norm) and OPTDetector (model.decoder.layers.*, kept distinct from Llama via model.decoder. detection).

The shared schema gains post_attn_norm and post_ffn_norm roles plus globals for embed_positions, final_norm.bias, and head_lm.bias so OPT and encoder layouts can round-trip into the same layer_N.role.weight keys without tensor ops.

Tests cover BERT/OPT detect + canonicalize (value identity, is_canonical), a regression that OPT per-layer final_layer_norm.bias maps to pre_ffn_norm.bias, and cross-detector exclusivity on Llama-layout state dicts.

Reviewed by Cursor Bugbot for commit f3b02aa. Bugbot is set up for automated code reviews on this repo. Configure here.

…cee-ai#708)

A shared role-key space so parameters from different architecture
families can be addressed uniformly before any merge method runs:

  model.layers.0.self_attn.q_proj.weight -> layer_0.attn_q.weight
  model.layers.11.mlp.up_proj.weight     -> layer_11.ffn_up.weight

Dependency-free name mapping only: values pass through by reference,
no tensor ops, no new requirements. Ships the Llama-family detector
(Llama, Qwen, Mistral, Gemma, TinyLlama layouts) with detection,
canonicalization, and an exact decanonicalize inverse (round-trip
tested). Remaining families (gpt2, bert, neox, opt, t5, phi) and the
shape bridge follow in the same pattern per arcee-ai#708.

The detector set is exercised end-to-end by the cross-family merge
pipeline behind the published Optitransfer converge collectives
(9 models across 4 architecture families in one set of weights).

Signed-off-by: Ryan Gillespie <mgillr@users.noreply.github.com>
BertFamilyDetector covers BERT and RoBERTa encoder layouts
(post-norm LayerNorms mapped onto the post_attn_norm / post_ffn_norm
canonical slots); OPTDetector covers the OPT decoder layout. Same
pattern as the Llama-family detector in arcee-ai#709: pure name mapping,
values by reference, no new dependencies. Cross-detector exclusivity
tested (a Llama-layout dict matches neither).

Signed-off-by: Ryan Gillespie <mgillr@users.noreply.github.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit ae93097. Configure here.

Comment thread mergekit/architecture/canonical.py
Comment thread tests/test_canonical.py
…t import branches

Signed-off-by: Ryan Gillespie <mgillr@users.noreply.github.com>
…r in CI)

Signed-off-by: Ryan Gillespie <mgillr@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant