fix: harden checkpoint blending contracts

This commit is contained in:
Joseph Magly
2026-08-21 12:12:14 -04:00
parent 17f1d5eb59
commit b746a87a00
6 changed files with 755 additions and 128 deletions
+68 -41
View File
@@ -1,26 +1,27 @@
# Complementary Abliteration Blending
**Date:** 2026-08-20
**Status:** Empirically Validated — V2 Shipped
**Authors:** OBLITERATUS Contributors
**Status:** Contributor-reported preliminary results; implementation independently tested
**Authors:** OBLITERATUS contributors and maintainers
---
## Abstract
We present **complementary abliteration blending**, a novel technique that combines two
abliteration methods with different failure modes via weight-space interpolation. The result
is the first abliterated model to exceed stock capability on MMLU (+1.1pp, n=570, lm-eval-harness)
while maintaining 0% refusal rate across 842 harmful prompts.
This document describes **complementary abliteration blending**, which combines two compatible
abliterated checkpoints via weight-space interpolation. PR #127 did not include the raw benchmark
outputs, prompt-level refusal evidence, model revisions, environment capture, or artifact hashes
needed to independently reproduce its numerical results. The tables below therefore preserve the
contributor's preliminary report as context; they are not independently verified project claims.
---
## 1. Motivation
All prior abliteration techniques face a fundamental tradeoff: deeper refusal removal causes
greater capability loss. Single-direction methods (Arditi et al., huihui-ai) preserve capability
but leave residual refusals. Multi-direction methods (OBLITERATUS V1, Gabliteration) achieve
complete refusal removal but at -6pp MMLU cost.
The contributor report frames abliteration as a tradeoff between refusal removal and measured
capability. Single-direction work such as [Arditi et al.][arditi] motivates the approach, while the
reported OBLITERATUS comparison observed lower MMLU for a more aggressive surgery. Raw evidence
for the specific refusal and MMLU figures was not included in PR #127.
We hypothesized that different direction-finding methods damage different parts of the model's
capability geometry, and that blending their outputs could cancel these damages.
@@ -39,9 +40,9 @@ obliteratus obliterate $BASE --method aggressive --n-directions 3 \
--min-layer-fraction 0.45
```
**Properties:** Greedy variance capture via SVD finds refusal directions that overlap with
capability-encoding subspaces. Deep refusal removal (0% refuse, 100% usable output) but
measurable capability damage (-2pp MMLU on 15-subject spot check).
**Contributor-reported properties:** deep refusal removal with measurable capability damage on a
small spot check. The proposed explanation—SVD directions overlapping capability subspaces—is a
hypothesis that requires activation and controlled model-comparison evidence.
### 2.2 Surgery B: LEACE
@@ -53,10 +54,10 @@ obliteratus obliterate $BASE --method aggressive --direction-method leace \
--refinement-passes 3 --min-layer-fraction 0.40
```
**Properties:** LEACE minimizes mutual information between the concept (refusal) and the
representation, preserving maximum non-refusal information by construction. Excellent capability
retention (+0.7pp MMLU) but weaker output quality (50% usable) because refusal removal is
less complete in the generation pathway.
LEACE is a closed-form linear concept-erasure method designed to remove linearly available concept
information while minimizing distortion ([Belrose et al.][leace]). The contributor reported
capability retention and weaker output quality for this surgery; the asserted mechanism in the
generation pathway was not evidenced in this PR.
### 2.3 Weight-Space LERP Blend
@@ -71,7 +72,11 @@ The blend ratio `alpha = 0.60` was found by binary search over {0.30, 0.50, 0.55
---
## 3. Results
## 3. Contributor-Reported Results (Not Independently Reproduced)
No machine-readable benchmark or refusal-evaluation artifacts accompany these tables. Counts,
uncertainty, prompt selection, exact model revisions, and evaluator configuration must be supplied
before the results can support a release or research conclusion.
### 3.1 Blend Ratio Search (15-subject MMLU, 150 questions)
@@ -86,9 +91,10 @@ The blend ratio `alpha = 0.60` was found by binary search over {0.30, 0.50, 0.55
| 70% | 0% | 80% | 86.7% | +0.7pp |
| 100% (pure LEACE)| 0% | 50% | 86.7% | +0.7pp |
The 60% blend is the unique optimum: maximum MMLU with 100% usable output.
The contributor selected the 60% blend from this small search. The evidence is insufficient to
establish a unique optimum or distinguish the apparent differences from sampling noise.
### 3.2 Full Validation (57-subject MMLU, 570 questions)
### 3.2 Larger Contributor Spot Check (57-subject MMLU, 570 questions)
| Model | MMLU | Stderr | vs Stock |
|----------------------|---------|--------|----------|
@@ -108,8 +114,8 @@ Capability gains span both safety-adjacent and neutral reasoning topics:
| Business Ethics | 80% | 100% | +20pp | Sensitive |
| Professional Law | 80% | 100% | +20pp | Sensitive |
Gains on neutral topics (math, logic) suggest real capability improvement, not just
reduced hedging on sensitive questions.
The reported neutral-topic gains motivate a controlled follow-up; five questions per subject are
not sufficient to establish capability improvement or rule out sampling variation.
### 3.4 Real-World Practical Tasks
@@ -119,7 +125,8 @@ reduced hedging on sensitive questions.
| Advanced (8 tasks) | 7/8 | 7/8 | ReAct agents, async refactor, K8s debug, |
| | | | security review, system design |
V2 matches stock on every practical capability while being fully uncensored.
The contributor reported comparable outcomes on this small practical-task set. The tasks and raw
outputs were not included, so maintainers have not independently verified that comparison.
---
@@ -127,38 +134,35 @@ V2 matches stock on every practical capability while being fully uncensored.
### 4.1 Complementary Error Cancellation
SVD and LEACE make **different mistakes in different parts of weight space**:
One hypothesis is that SVD and LEACE make different errors in weight space:
- **SVD** greedily captures maximum variance directions. Some captured variance encodes
capability, not just refusal. This damages specific weight regions.
- **LEACE** minimizes mutual information, preserving capability by construction. But it
leaves refusal residue in the generation pathway (attention heads, output projections)
that SVD would have removed.
- **LEACE** minimizes a linear erasure objective. Whether it leaves specific residue in attention
heads or output projections must be measured rather than inferred from output behavior.
Weight-space interpolation averages these complementary errors:
Weight-space interpolation may average complementary errors:
- Where SVD damaged capability, LEACE's intact weights dilute the damage
- Where LEACE left refusal residue, SVD's clean weights dilute the residue
### 4.2 Theoretical Connection to Model Merging
This technique is analogous to model merging (TIES, DARE, Model Soups) but applied within
the abliteration domain. The key insight is that the "task vectors" (weight deltas from stock)
created by different abliteration methods are approximately orthogonal in the dimensions that
matter — refusal removal is shared, but capability damage is method-specific.
This technique is related to model-merging work such as [Model Soups][model-soups] and
[TIES-Merging][ties], but is applied within the abliteration domain. PR #127 does not measure
task-vector orthogonality or error anti-correlation, so those remain testable explanations rather
than established properties.
### 4.3 Capacity Hypothesis
The +1.1pp MMLU improvement over stock raises the possibility that refusal training
consumes representational capacity that abliteration frees. Formal validation experiments
are provided in `obliteratus/capacity_hypothesis.py`:
The contributor-reported MMLU difference raises several competing hypotheses. A separate,
provenance-gated experiment framework is tracked in [issue #132][capacity-issue]:
1. **Activation Rank Analysis** — Does effective dimensionality increase after abliteration?
2. **Topic Cluster Analysis** — Do gains cluster on sensitive topics (hedging) or spread broadly (capacity)?
3. **Blend Control** — Does blending two identical SVD surgeries also gain MMLU? (Tests regularization hypothesis)
4. **Learning Absorption** — Does the abliterated model learn new information faster? (Tests freed capacity directly)
Preliminary topic cluster analysis shows gains on both sensitive (law, ethics) and neutral
(math, logic) topics, partially supporting the capacity hypothesis.
The small per-subject report does not distinguish these hypotheses.
---
@@ -177,22 +181,31 @@ obliteratus obliterate $BASE --method aggressive --direction-method leace \
# Step 3: Blend
obliteratus blend --model-a surgery_a --model-b surgery_b --alpha 0.6 \
--output blended_model
--config-source a --output blended_model
# Step 4: Validate
lm_eval --model hf --model_args pretrained=blended_model --tasks mmlu
```
The command requires matching `source_model` values in both checkpoints'
`abliteration_metadata.json`, identical tensor keys, compatible shapes/dtypes, floating-point
weights, and sharded safetensors indexes. For legacy checkpoints whose common lineage was verified
out of band, `--allow-unverified-lineage` is an explicit escape hatch. Output is staged and
validated before atomically replacing any existing destination.
---
## 6. Limitations
- Validated only on Qwen3.8-27B; generalization to other architectures is untested
- Contributor measurements cover only Qwen3.8-27B; independent validation is pending
- MMLU is a multiple-choice benchmark; gains may not transfer to all downstream tasks
- The 60/40 blend ratio may be model-specific
- `repetition_penalty=1.15` is still required for clean generation
- System prompts still reintroduce refusals
- Full 842-corpus validation in progress at time of writing
- The numerical results lack committed raw evidence and independent reproduction
- Single-file and quantized/integer checkpoints are not supported by the current blender
- Atomic promotion temporarily requires space for the complete staged output and any prior output
---
@@ -201,5 +214,19 @@ lm_eval --model hf --model_args pretrained=blended_model --tasks mmlu
- **Cross-architecture validation** on Llama, Gemma, Mistral
- **SLERP blending** instead of LERP (spherical interpolation may better preserve weight norms)
- **Three-way blends** with additional direction methods (diff_means, SOM)
- **Post-blend recovery** via QLoRA fine-tuning on capability data
- **Formal capacity hypothesis validation** using the provided experiment framework
- **Post-blend recovery**, tracked separately in [issue #133][recovery-issue]
- **Formal capacity-hypothesis validation**, tracked in [issue #132][capacity-issue]
## References
- [Arditi et al., *Refusal in Language Models Is Mediated by a Single Direction*][arditi]
- [Belrose et al., *LEACE: Perfect Linear Concept Erasure in Closed Form*][leace]
- [Wortsman et al., *Model Soups*][model-soups]
- [Yadav et al., *TIES-Merging*][ties]
[arditi]: https://arxiv.org/abs/2406.11717
[leace]: https://arxiv.org/abs/2306.03819
[model-soups]: https://proceedings.mlr.press/v162/wortsman22a.html
[ties]: https://arxiv.org/abs/2306.01708
[capacity-issue]: https://github.com/elder-plinius/OBLITERATUS/issues/132
[recovery-issue]: https://github.com/elder-plinius/OBLITERATUS/issues/133