Does quantization compound abliteration's damage to structured-output adherence? #5005
behrnt-slatgng
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Disclosure first: I build Grunz, a hosted chat + coding agent running abliterated open-weight models. I'm a consumer of inference endpoints rather than an operator — I don't run lmdeploy myself — so this is a question from downstream rather than a report from inside the stack. I'd rather ask than keep guessing.
The observation I can stand behind: abliteration damages instruction-following and output-format adherence well before it damages knowledge. The model still knows the material; it stops reliably respecting the chat template, stop sequences and JSON schema contracts. Perplexity and MMLU-style numbers barely move while structured-output failure rate climbs, which means the standard metrics report the model as fine while the application says otherwise.
The thing I can't answer from where I sit: how that interacts with quantization.
My working suspicion is that the two degradations are not merely additive — that an abliterated checkpoint is disproportionately sensitive to precision loss specifically on the format-adherence axis, rather than on quality broadly. The intuition is that abliteration works by projecting a direction out of the weights, so the structure has already been deliberately flattened before quantization rounds anything away. But that is reasoning from a mechanism I only half understand, applied to a symptom I can see from outside, and it could easily be wrong.
Two questions:
Does anyone here have data on AWQ (or KV cache quantization) applied to modified checkpoints versus their unmodified counterparts, measured on format compliance rather than on quality benchmarks? If the compounding effect is real it should be measurable, and if it isn't I'd like to stop repeating it.
Is there existing tooling or a recommended eval pattern for tracking structured-output compliance as its own metric across quant levels? The obvious approach is a schema-validation pass over generated output with a compliance rate reported per configuration, which feels like something that should already exist rather than something each person reinvents badly.
Happy to be told the framing is wrong — that would itself be useful.
All reactions