Skip to content

[GKD] Preserve EOS supervision for off-policy distillation - #10285

Merged
hjh0119 merged 3 commits into
modelscope:mainfrom
taking-lying-flat:fix/gkd-off-policy-eos
Oct 2, 2026
Merged

hjh0119 merged 3 commits into
modelscope:mainfrom
taking-lying-flat:fix/gkd-off-policy-eos

Conversation

@taking-lying-flat

@taking-lying-flat taking-lying-flat commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

PR type

  • Bug Fix
  • New Feature
  • Document Updates
  • More Models or Datasets Support

PR information

Fixes #10282.

Off-policy GKD (--lmbda 0) inherits the rollout default add_eos=False, so dataset responses without a literal EOS lose EOS supervision in both distillation and the auxiliary SFT loss. Restoring only student labels leaves remote teacher prompt logprobs without the EOS position, which is then filtered out of the KL loss.

Use the normal training EOS policy (None) for dataset GKDSamples and preserve it in the OPSD teacher view. For remote scoring, add an explicit 'auto' value to RolloutInferRequest.add_eos and forward it from both HF/Megatron and Ray GKD requests. The template handles this value during actual encoding, using its existing stop-word, tokenizer-separator and token-overlap rules. The existing None, True and False behavior remains unchanged; generated samples still set False, including length-truncated rollouts.

The final diff changes five production files: 10 insertions and 4 deletions. The base template changes only two condition lines; no extra rendering or helper for inferring EOS is needed. Template registrations and model-specific templates are unchanged. No test files are included.

Both trainer and teacher must use the updated implementation: older teacher servers do not support the new request mode.

Experiment results

  • With the real local Qwen3 tokenizer, the original dataset sample had 15 supervised tokens without EOS; ordinary training had 17, including EOS and its trailing newline. Patched GKD matches ordinary training, and the JSON-round-tripped teacher request has identical input IDs. Synthetic top-64 teacher prompt logprobs retain EOS in the real GKD loss and produce finite gradients in the expected direction, with left/right padding and packing.
  • 414 local checks passed: 23 EOS/loss cases, 323 template differential cases and 68 explicit-auto request cases. Coverage includes existing stops, GLM observation stops, Gemma newline preprocessing, explicit overrides, tokenizer separators, empty suffixes, stopped/truncated rollouts, OPSD, tool schemas, JSON/dacite round trips, request isolation, Jinja and Ray request construction with replica calls stubbed. Request construction was also verified not to call the template renderer again.
  • Template comparisons cover all 248 registered metadata configurations through the shared serializer and text encoding/labels for 12 representative template types. These use one local Qwen tokenizer to isolate template behavior; TeleChat-specific token attributes use fixed placeholders. This is not full-model or multimodal-processor validation.
  • 115 existing tests passed, plus 107 subtests, covering GKD loss, teacher-logprob assembly, routing, teacher advantages, OPSD data handling, exact rollout token I/O and model templates. Five cases were excluded: four OPSD cases use a _FakeTemplate without _get_response_prefix and were separately confirmed to fail identically on unmodified main; one case requires an unavailable Qwen2-VL processor.
  • All applicable changed-file pre-commit hooks and git diff --check passed. The published patch was compared byte-for-byte with this validated candidate.

Validation used CPU. Teacher logprobs were synthetic; live vLLM scoring, distributed Ray/Megatron execution and the reported full 8B training run were not exercised.

@hjh0119
hjh0119 merged commit c2bcc23 into modelscope:main Oct 2, 2026
3 checks passed
@wodeai192

Copy link
Copy Markdown

Thanks for fixing this problem i found hope this SIWFT framework better

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GKD 的离线蒸馏没有办法学习 结束符号的 logits

3 participants