Can a short format-focused SFT pass repair instruction-following damage from weight editing? #11958
behrnt-slatgng
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Disclosure first: I run open weights in production behind a hosted chat + coding agent (Grunz). Nothing to sell here — this is a finetuning question I don't have the background to answer myself, and this seems like the right room for it.
Background
There are now thousands of "abliterated" variants on the Hub — models where a refusal direction is identified in activation space and projected out of the weights. It's a weight-editing operation, not a finetune: no gradient steps, no data.
What I observe running these in production, consistently across variants and model families, is that the operation damages instruction-following and output-format adherence noticeably more than it damages knowledge or reasoning.
The model still knows the material. What degrades is:
On generic benchmarks these models often look fine. In an agent loop, where output gets parsed, the failure rate is meaningfully higher than the base model's at the same quant.
The question
Would a short, format-focused SFT pass repair that?
The intuition is that this is exactly the kind of narrow, mechanical behaviour that a small amount of well-targeted instruction data should be able to restore — template adherence, stop-sequence discipline, schema compliance — without needing anything like a full instruction-tuning run. If that's right, it's a few hundred to a few thousand examples and an hour on a consumer card, which is squarely what Unsloth is for.
What I don't know:
Is that intuition wrong? If the damage is distributed rather than localised, a small LoRA may just paper over it on the training distribution and not generalise.
What happens to the weight edit? A projection removed a direction; a finetune moves the weights. I'd genuinely like to know whether a light SFT pass measurably perturbs the edit, leaves it intact, or something in between. Either answer is interesting — I just haven't seen anyone characterise it.
What would you even train on? The target behaviour isn't "right answer," it's "well-formed output regardless of answer." A dataset of prompts with explicit output contracts, labelled on contract compliance rather than content, feels like the right shape — but I don't know if something like that already exists.
LoRA rank and target modules — if anyone has intuition for whether format adherence lives somewhere specific enough that a low-rank adapter on attention projections would be enough, versus needing MLP layers too, I'd like to hear it.
Why I'm asking rather than just trying it
I can run the experiment, but I'd design it badly without input. Specifically I don't have a good way to measure the thing I'm trying to fix — perplexity barely moves while format compliance falls off a cliff, so the standard metric says the model is fine when the application says it isn't. I've asked about a compliance-rate eval elsewhere; if someone here already has a measurement approach they trust, that's the piece I'm missing.
Happy to run it and write up the result either way if the answer is "nobody has tried this."
All reactions