mirror of
https://github.com/elder-plinius/OBLITERATUS.git
synced 2026-08-30 06:30:37 +02:00
fix: harden checkpoint blending contracts
This commit is contained in:
@@ -1,26 +1,27 @@
|
||||
# Complementary Abliteration Blending
|
||||
|
||||
**Date:** 2026-08-20
|
||||
**Status:** Empirically Validated — V2 Shipped
|
||||
**Authors:** OBLITERATUS Contributors
|
||||
**Status:** Contributor-reported preliminary results; implementation independently tested
|
||||
**Authors:** OBLITERATUS contributors and maintainers
|
||||
|
||||
---
|
||||
|
||||
## Abstract
|
||||
|
||||
We present **complementary abliteration blending**, a novel technique that combines two
|
||||
abliteration methods with different failure modes via weight-space interpolation. The result
|
||||
is the first abliterated model to exceed stock capability on MMLU (+1.1pp, n=570, lm-eval-harness)
|
||||
while maintaining 0% refusal rate across 842 harmful prompts.
|
||||
This document describes **complementary abliteration blending**, which combines two compatible
|
||||
abliterated checkpoints via weight-space interpolation. PR #127 did not include the raw benchmark
|
||||
outputs, prompt-level refusal evidence, model revisions, environment capture, or artifact hashes
|
||||
needed to independently reproduce its numerical results. The tables below therefore preserve the
|
||||
contributor's preliminary report as context; they are not independently verified project claims.
|
||||
|
||||
---
|
||||
|
||||
## 1. Motivation
|
||||
|
||||
All prior abliteration techniques face a fundamental tradeoff: deeper refusal removal causes
|
||||
greater capability loss. Single-direction methods (Arditi et al., huihui-ai) preserve capability
|
||||
but leave residual refusals. Multi-direction methods (OBLITERATUS V1, Gabliteration) achieve
|
||||
complete refusal removal but at -6pp MMLU cost.
|
||||
The contributor report frames abliteration as a tradeoff between refusal removal and measured
|
||||
capability. Single-direction work such as [Arditi et al.][arditi] motivates the approach, while the
|
||||
reported OBLITERATUS comparison observed lower MMLU for a more aggressive surgery. Raw evidence
|
||||
for the specific refusal and MMLU figures was not included in PR #127.
|
||||
|
||||
We hypothesized that different direction-finding methods damage different parts of the model's
|
||||
capability geometry, and that blending their outputs could cancel these damages.
|
||||
@@ -39,9 +40,9 @@ obliteratus obliterate $BASE --method aggressive --n-directions 3 \
|
||||
--min-layer-fraction 0.45
|
||||
```
|
||||
|
||||
**Properties:** Greedy variance capture via SVD finds refusal directions that overlap with
|
||||
capability-encoding subspaces. Deep refusal removal (0% refuse, 100% usable output) but
|
||||
measurable capability damage (-2pp MMLU on 15-subject spot check).
|
||||
**Contributor-reported properties:** deep refusal removal with measurable capability damage on a
|
||||
small spot check. The proposed explanation—SVD directions overlapping capability subspaces—is a
|
||||
hypothesis that requires activation and controlled model-comparison evidence.
|
||||
|
||||
### 2.2 Surgery B: LEACE
|
||||
|
||||
@@ -53,10 +54,10 @@ obliteratus obliterate $BASE --method aggressive --direction-method leace \
|
||||
--refinement-passes 3 --min-layer-fraction 0.40
|
||||
```
|
||||
|
||||
**Properties:** LEACE minimizes mutual information between the concept (refusal) and the
|
||||
representation, preserving maximum non-refusal information by construction. Excellent capability
|
||||
retention (+0.7pp MMLU) but weaker output quality (50% usable) because refusal removal is
|
||||
less complete in the generation pathway.
|
||||
LEACE is a closed-form linear concept-erasure method designed to remove linearly available concept
|
||||
information while minimizing distortion ([Belrose et al.][leace]). The contributor reported
|
||||
capability retention and weaker output quality for this surgery; the asserted mechanism in the
|
||||
generation pathway was not evidenced in this PR.
|
||||
|
||||
### 2.3 Weight-Space LERP Blend
|
||||
|
||||
@@ -71,7 +72,11 @@ The blend ratio `alpha = 0.60` was found by binary search over {0.30, 0.50, 0.55
|
||||
|
||||
---
|
||||
|
||||
## 3. Results
|
||||
## 3. Contributor-Reported Results (Not Independently Reproduced)
|
||||
|
||||
No machine-readable benchmark or refusal-evaluation artifacts accompany these tables. Counts,
|
||||
uncertainty, prompt selection, exact model revisions, and evaluator configuration must be supplied
|
||||
before the results can support a release or research conclusion.
|
||||
|
||||
### 3.1 Blend Ratio Search (15-subject MMLU, 150 questions)
|
||||
|
||||
@@ -86,9 +91,10 @@ The blend ratio `alpha = 0.60` was found by binary search over {0.30, 0.50, 0.55
|
||||
| 70% | 0% | 80% | 86.7% | +0.7pp |
|
||||
| 100% (pure LEACE)| 0% | 50% | 86.7% | +0.7pp |
|
||||
|
||||
The 60% blend is the unique optimum: maximum MMLU with 100% usable output.
|
||||
The contributor selected the 60% blend from this small search. The evidence is insufficient to
|
||||
establish a unique optimum or distinguish the apparent differences from sampling noise.
|
||||
|
||||
### 3.2 Full Validation (57-subject MMLU, 570 questions)
|
||||
### 3.2 Larger Contributor Spot Check (57-subject MMLU, 570 questions)
|
||||
|
||||
| Model | MMLU | Stderr | vs Stock |
|
||||
|----------------------|---------|--------|----------|
|
||||
@@ -108,8 +114,8 @@ Capability gains span both safety-adjacent and neutral reasoning topics:
|
||||
| Business Ethics | 80% | 100% | +20pp | Sensitive |
|
||||
| Professional Law | 80% | 100% | +20pp | Sensitive |
|
||||
|
||||
Gains on neutral topics (math, logic) suggest real capability improvement, not just
|
||||
reduced hedging on sensitive questions.
|
||||
The reported neutral-topic gains motivate a controlled follow-up; five questions per subject are
|
||||
not sufficient to establish capability improvement or rule out sampling variation.
|
||||
|
||||
### 3.4 Real-World Practical Tasks
|
||||
|
||||
@@ -119,7 +125,8 @@ reduced hedging on sensitive questions.
|
||||
| Advanced (8 tasks) | 7/8 | 7/8 | ReAct agents, async refactor, K8s debug, |
|
||||
| | | | security review, system design |
|
||||
|
||||
V2 matches stock on every practical capability while being fully uncensored.
|
||||
The contributor reported comparable outcomes on this small practical-task set. The tasks and raw
|
||||
outputs were not included, so maintainers have not independently verified that comparison.
|
||||
|
||||
---
|
||||
|
||||
@@ -127,38 +134,35 @@ V2 matches stock on every practical capability while being fully uncensored.
|
||||
|
||||
### 4.1 Complementary Error Cancellation
|
||||
|
||||
SVD and LEACE make **different mistakes in different parts of weight space**:
|
||||
One hypothesis is that SVD and LEACE make different errors in weight space:
|
||||
|
||||
- **SVD** greedily captures maximum variance directions. Some captured variance encodes
|
||||
capability, not just refusal. This damages specific weight regions.
|
||||
- **LEACE** minimizes mutual information, preserving capability by construction. But it
|
||||
leaves refusal residue in the generation pathway (attention heads, output projections)
|
||||
that SVD would have removed.
|
||||
- **LEACE** minimizes a linear erasure objective. Whether it leaves specific residue in attention
|
||||
heads or output projections must be measured rather than inferred from output behavior.
|
||||
|
||||
Weight-space interpolation averages these complementary errors:
|
||||
Weight-space interpolation may average complementary errors:
|
||||
- Where SVD damaged capability, LEACE's intact weights dilute the damage
|
||||
- Where LEACE left refusal residue, SVD's clean weights dilute the residue
|
||||
|
||||
### 4.2 Theoretical Connection to Model Merging
|
||||
|
||||
This technique is analogous to model merging (TIES, DARE, Model Soups) but applied within
|
||||
the abliteration domain. The key insight is that the "task vectors" (weight deltas from stock)
|
||||
created by different abliteration methods are approximately orthogonal in the dimensions that
|
||||
matter — refusal removal is shared, but capability damage is method-specific.
|
||||
This technique is related to model-merging work such as [Model Soups][model-soups] and
|
||||
[TIES-Merging][ties], but is applied within the abliteration domain. PR #127 does not measure
|
||||
task-vector orthogonality or error anti-correlation, so those remain testable explanations rather
|
||||
than established properties.
|
||||
|
||||
### 4.3 Capacity Hypothesis
|
||||
|
||||
The +1.1pp MMLU improvement over stock raises the possibility that refusal training
|
||||
consumes representational capacity that abliteration frees. Formal validation experiments
|
||||
are provided in `obliteratus/capacity_hypothesis.py`:
|
||||
The contributor-reported MMLU difference raises several competing hypotheses. A separate,
|
||||
provenance-gated experiment framework is tracked in [issue #132][capacity-issue]:
|
||||
|
||||
1. **Activation Rank Analysis** — Does effective dimensionality increase after abliteration?
|
||||
2. **Topic Cluster Analysis** — Do gains cluster on sensitive topics (hedging) or spread broadly (capacity)?
|
||||
3. **Blend Control** — Does blending two identical SVD surgeries also gain MMLU? (Tests regularization hypothesis)
|
||||
4. **Learning Absorption** — Does the abliterated model learn new information faster? (Tests freed capacity directly)
|
||||
|
||||
Preliminary topic cluster analysis shows gains on both sensitive (law, ethics) and neutral
|
||||
(math, logic) topics, partially supporting the capacity hypothesis.
|
||||
The small per-subject report does not distinguish these hypotheses.
|
||||
|
||||
---
|
||||
|
||||
@@ -177,22 +181,31 @@ obliteratus obliterate $BASE --method aggressive --direction-method leace \
|
||||
|
||||
# Step 3: Blend
|
||||
obliteratus blend --model-a surgery_a --model-b surgery_b --alpha 0.6 \
|
||||
--output blended_model
|
||||
--config-source a --output blended_model
|
||||
|
||||
# Step 4: Validate
|
||||
lm_eval --model hf --model_args pretrained=blended_model --tasks mmlu
|
||||
```
|
||||
|
||||
The command requires matching `source_model` values in both checkpoints'
|
||||
`abliteration_metadata.json`, identical tensor keys, compatible shapes/dtypes, floating-point
|
||||
weights, and sharded safetensors indexes. For legacy checkpoints whose common lineage was verified
|
||||
out of band, `--allow-unverified-lineage` is an explicit escape hatch. Output is staged and
|
||||
validated before atomically replacing any existing destination.
|
||||
|
||||
---
|
||||
|
||||
## 6. Limitations
|
||||
|
||||
- Validated only on Qwen3.8-27B; generalization to other architectures is untested
|
||||
- Contributor measurements cover only Qwen3.8-27B; independent validation is pending
|
||||
- MMLU is a multiple-choice benchmark; gains may not transfer to all downstream tasks
|
||||
- The 60/40 blend ratio may be model-specific
|
||||
- `repetition_penalty=1.15` is still required for clean generation
|
||||
- System prompts still reintroduce refusals
|
||||
- Full 842-corpus validation in progress at time of writing
|
||||
- The numerical results lack committed raw evidence and independent reproduction
|
||||
- Single-file and quantized/integer checkpoints are not supported by the current blender
|
||||
- Atomic promotion temporarily requires space for the complete staged output and any prior output
|
||||
|
||||
---
|
||||
|
||||
@@ -201,5 +214,19 @@ lm_eval --model hf --model_args pretrained=blended_model --tasks mmlu
|
||||
- **Cross-architecture validation** on Llama, Gemma, Mistral
|
||||
- **SLERP blending** instead of LERP (spherical interpolation may better preserve weight norms)
|
||||
- **Three-way blends** with additional direction methods (diff_means, SOM)
|
||||
- **Post-blend recovery** via QLoRA fine-tuning on capability data
|
||||
- **Formal capacity hypothesis validation** using the provided experiment framework
|
||||
- **Post-blend recovery**, tracked separately in [issue #133][recovery-issue]
|
||||
- **Formal capacity-hypothesis validation**, tracked in [issue #132][capacity-issue]
|
||||
|
||||
## References
|
||||
|
||||
- [Arditi et al., *Refusal in Language Models Is Mediated by a Single Direction*][arditi]
|
||||
- [Belrose et al., *LEACE: Perfect Linear Concept Erasure in Closed Form*][leace]
|
||||
- [Wortsman et al., *Model Soups*][model-soups]
|
||||
- [Yadav et al., *TIES-Merging*][ties]
|
||||
|
||||
[arditi]: https://arxiv.org/abs/2406.11717
|
||||
[leace]: https://arxiv.org/abs/2306.03819
|
||||
[model-soups]: https://proceedings.mlr.press/v162/wortsman22a.html
|
||||
[ties]: https://arxiv.org/abs/2306.01708
|
||||
[capacity-issue]: https://github.com/elder-plinius/OBLITERATUS/issues/132
|
||||
[recovery-issue]: https://github.com/elder-plinius/OBLITERATUS/issues/133
|
||||
|
||||
Reference in New Issue
Block a user