mirror of
https://github.com/elder-plinius/OBLITERATUS.git
synced 2026-08-30 06:30:37 +02:00
fix: harden checkpoint blending contracts
This commit is contained in:
@@ -1,26 +1,27 @@
|
||||
# Complementary Abliteration Blending
|
||||
|
||||
**Date:** 2026-08-20
|
||||
**Status:** Empirically Validated — V2 Shipped
|
||||
**Authors:** OBLITERATUS Contributors
|
||||
**Status:** Contributor-reported preliminary results; implementation independently tested
|
||||
**Authors:** OBLITERATUS contributors and maintainers
|
||||
|
||||
---
|
||||
|
||||
## Abstract
|
||||
|
||||
We present **complementary abliteration blending**, a novel technique that combines two
|
||||
abliteration methods with different failure modes via weight-space interpolation. The result
|
||||
is the first abliterated model to exceed stock capability on MMLU (+1.1pp, n=570, lm-eval-harness)
|
||||
while maintaining 0% refusal rate across 842 harmful prompts.
|
||||
This document describes **complementary abliteration blending**, which combines two compatible
|
||||
abliterated checkpoints via weight-space interpolation. PR #127 did not include the raw benchmark
|
||||
outputs, prompt-level refusal evidence, model revisions, environment capture, or artifact hashes
|
||||
needed to independently reproduce its numerical results. The tables below therefore preserve the
|
||||
contributor's preliminary report as context; they are not independently verified project claims.
|
||||
|
||||
---
|
||||
|
||||
## 1. Motivation
|
||||
|
||||
All prior abliteration techniques face a fundamental tradeoff: deeper refusal removal causes
|
||||
greater capability loss. Single-direction methods (Arditi et al., huihui-ai) preserve capability
|
||||
but leave residual refusals. Multi-direction methods (OBLITERATUS V1, Gabliteration) achieve
|
||||
complete refusal removal but at -6pp MMLU cost.
|
||||
The contributor report frames abliteration as a tradeoff between refusal removal and measured
|
||||
capability. Single-direction work such as [Arditi et al.][arditi] motivates the approach, while the
|
||||
reported OBLITERATUS comparison observed lower MMLU for a more aggressive surgery. Raw evidence
|
||||
for the specific refusal and MMLU figures was not included in PR #127.
|
||||
|
||||
We hypothesized that different direction-finding methods damage different parts of the model's
|
||||
capability geometry, and that blending their outputs could cancel these damages.
|
||||
@@ -39,9 +40,9 @@ obliteratus obliterate $BASE --method aggressive --n-directions 3 \
|
||||
--min-layer-fraction 0.45
|
||||
```
|
||||
|
||||
**Properties:** Greedy variance capture via SVD finds refusal directions that overlap with
|
||||
capability-encoding subspaces. Deep refusal removal (0% refuse, 100% usable output) but
|
||||
measurable capability damage (-2pp MMLU on 15-subject spot check).
|
||||
**Contributor-reported properties:** deep refusal removal with measurable capability damage on a
|
||||
small spot check. The proposed explanation—SVD directions overlapping capability subspaces—is a
|
||||
hypothesis that requires activation and controlled model-comparison evidence.
|
||||
|
||||
### 2.2 Surgery B: LEACE
|
||||
|
||||
@@ -53,10 +54,10 @@ obliteratus obliterate $BASE --method aggressive --direction-method leace \
|
||||
--refinement-passes 3 --min-layer-fraction 0.40
|
||||
```
|
||||
|
||||
**Properties:** LEACE minimizes mutual information between the concept (refusal) and the
|
||||
representation, preserving maximum non-refusal information by construction. Excellent capability
|
||||
retention (+0.7pp MMLU) but weaker output quality (50% usable) because refusal removal is
|
||||
less complete in the generation pathway.
|
||||
LEACE is a closed-form linear concept-erasure method designed to remove linearly available concept
|
||||
information while minimizing distortion ([Belrose et al.][leace]). The contributor reported
|
||||
capability retention and weaker output quality for this surgery; the asserted mechanism in the
|
||||
generation pathway was not evidenced in this PR.
|
||||
|
||||
### 2.3 Weight-Space LERP Blend
|
||||
|
||||
@@ -71,7 +72,11 @@ The blend ratio `alpha = 0.60` was found by binary search over {0.30, 0.50, 0.55
|
||||
|
||||
---
|
||||
|
||||
## 3. Results
|
||||
## 3. Contributor-Reported Results (Not Independently Reproduced)
|
||||
|
||||
No machine-readable benchmark or refusal-evaluation artifacts accompany these tables. Counts,
|
||||
uncertainty, prompt selection, exact model revisions, and evaluator configuration must be supplied
|
||||
before the results can support a release or research conclusion.
|
||||
|
||||
### 3.1 Blend Ratio Search (15-subject MMLU, 150 questions)
|
||||
|
||||
@@ -86,9 +91,10 @@ The blend ratio `alpha = 0.60` was found by binary search over {0.30, 0.50, 0.55
|
||||
| 70% | 0% | 80% | 86.7% | +0.7pp |
|
||||
| 100% (pure LEACE)| 0% | 50% | 86.7% | +0.7pp |
|
||||
|
||||
The 60% blend is the unique optimum: maximum MMLU with 100% usable output.
|
||||
The contributor selected the 60% blend from this small search. The evidence is insufficient to
|
||||
establish a unique optimum or distinguish the apparent differences from sampling noise.
|
||||
|
||||
### 3.2 Full Validation (57-subject MMLU, 570 questions)
|
||||
### 3.2 Larger Contributor Spot Check (57-subject MMLU, 570 questions)
|
||||
|
||||
| Model | MMLU | Stderr | vs Stock |
|
||||
|----------------------|---------|--------|----------|
|
||||
@@ -108,8 +114,8 @@ Capability gains span both safety-adjacent and neutral reasoning topics:
|
||||
| Business Ethics | 80% | 100% | +20pp | Sensitive |
|
||||
| Professional Law | 80% | 100% | +20pp | Sensitive |
|
||||
|
||||
Gains on neutral topics (math, logic) suggest real capability improvement, not just
|
||||
reduced hedging on sensitive questions.
|
||||
The reported neutral-topic gains motivate a controlled follow-up; five questions per subject are
|
||||
not sufficient to establish capability improvement or rule out sampling variation.
|
||||
|
||||
### 3.4 Real-World Practical Tasks
|
||||
|
||||
@@ -119,7 +125,8 @@ reduced hedging on sensitive questions.
|
||||
| Advanced (8 tasks) | 7/8 | 7/8 | ReAct agents, async refactor, K8s debug, |
|
||||
| | | | security review, system design |
|
||||
|
||||
V2 matches stock on every practical capability while being fully uncensored.
|
||||
The contributor reported comparable outcomes on this small practical-task set. The tasks and raw
|
||||
outputs were not included, so maintainers have not independently verified that comparison.
|
||||
|
||||
---
|
||||
|
||||
@@ -127,38 +134,35 @@ V2 matches stock on every practical capability while being fully uncensored.
|
||||
|
||||
### 4.1 Complementary Error Cancellation
|
||||
|
||||
SVD and LEACE make **different mistakes in different parts of weight space**:
|
||||
One hypothesis is that SVD and LEACE make different errors in weight space:
|
||||
|
||||
- **SVD** greedily captures maximum variance directions. Some captured variance encodes
|
||||
capability, not just refusal. This damages specific weight regions.
|
||||
- **LEACE** minimizes mutual information, preserving capability by construction. But it
|
||||
leaves refusal residue in the generation pathway (attention heads, output projections)
|
||||
that SVD would have removed.
|
||||
- **LEACE** minimizes a linear erasure objective. Whether it leaves specific residue in attention
|
||||
heads or output projections must be measured rather than inferred from output behavior.
|
||||
|
||||
Weight-space interpolation averages these complementary errors:
|
||||
Weight-space interpolation may average complementary errors:
|
||||
- Where SVD damaged capability, LEACE's intact weights dilute the damage
|
||||
- Where LEACE left refusal residue, SVD's clean weights dilute the residue
|
||||
|
||||
### 4.2 Theoretical Connection to Model Merging
|
||||
|
||||
This technique is analogous to model merging (TIES, DARE, Model Soups) but applied within
|
||||
the abliteration domain. The key insight is that the "task vectors" (weight deltas from stock)
|
||||
created by different abliteration methods are approximately orthogonal in the dimensions that
|
||||
matter — refusal removal is shared, but capability damage is method-specific.
|
||||
This technique is related to model-merging work such as [Model Soups][model-soups] and
|
||||
[TIES-Merging][ties], but is applied within the abliteration domain. PR #127 does not measure
|
||||
task-vector orthogonality or error anti-correlation, so those remain testable explanations rather
|
||||
than established properties.
|
||||
|
||||
### 4.3 Capacity Hypothesis
|
||||
|
||||
The +1.1pp MMLU improvement over stock raises the possibility that refusal training
|
||||
consumes representational capacity that abliteration frees. Formal validation experiments
|
||||
are provided in `obliteratus/capacity_hypothesis.py`:
|
||||
The contributor-reported MMLU difference raises several competing hypotheses. A separate,
|
||||
provenance-gated experiment framework is tracked in [issue #132][capacity-issue]:
|
||||
|
||||
1. **Activation Rank Analysis** — Does effective dimensionality increase after abliteration?
|
||||
2. **Topic Cluster Analysis** — Do gains cluster on sensitive topics (hedging) or spread broadly (capacity)?
|
||||
3. **Blend Control** — Does blending two identical SVD surgeries also gain MMLU? (Tests regularization hypothesis)
|
||||
4. **Learning Absorption** — Does the abliterated model learn new information faster? (Tests freed capacity directly)
|
||||
|
||||
Preliminary topic cluster analysis shows gains on both sensitive (law, ethics) and neutral
|
||||
(math, logic) topics, partially supporting the capacity hypothesis.
|
||||
The small per-subject report does not distinguish these hypotheses.
|
||||
|
||||
---
|
||||
|
||||
@@ -177,22 +181,31 @@ obliteratus obliterate $BASE --method aggressive --direction-method leace \
|
||||
|
||||
# Step 3: Blend
|
||||
obliteratus blend --model-a surgery_a --model-b surgery_b --alpha 0.6 \
|
||||
--output blended_model
|
||||
--config-source a --output blended_model
|
||||
|
||||
# Step 4: Validate
|
||||
lm_eval --model hf --model_args pretrained=blended_model --tasks mmlu
|
||||
```
|
||||
|
||||
The command requires matching `source_model` values in both checkpoints'
|
||||
`abliteration_metadata.json`, identical tensor keys, compatible shapes/dtypes, floating-point
|
||||
weights, and sharded safetensors indexes. For legacy checkpoints whose common lineage was verified
|
||||
out of band, `--allow-unverified-lineage` is an explicit escape hatch. Output is staged and
|
||||
validated before atomically replacing any existing destination.
|
||||
|
||||
---
|
||||
|
||||
## 6. Limitations
|
||||
|
||||
- Validated only on Qwen3.8-27B; generalization to other architectures is untested
|
||||
- Contributor measurements cover only Qwen3.8-27B; independent validation is pending
|
||||
- MMLU is a multiple-choice benchmark; gains may not transfer to all downstream tasks
|
||||
- The 60/40 blend ratio may be model-specific
|
||||
- `repetition_penalty=1.15` is still required for clean generation
|
||||
- System prompts still reintroduce refusals
|
||||
- Full 842-corpus validation in progress at time of writing
|
||||
- The numerical results lack committed raw evidence and independent reproduction
|
||||
- Single-file and quantized/integer checkpoints are not supported by the current blender
|
||||
- Atomic promotion temporarily requires space for the complete staged output and any prior output
|
||||
|
||||
---
|
||||
|
||||
@@ -201,5 +214,19 @@ lm_eval --model hf --model_args pretrained=blended_model --tasks mmlu
|
||||
- **Cross-architecture validation** on Llama, Gemma, Mistral
|
||||
- **SLERP blending** instead of LERP (spherical interpolation may better preserve weight norms)
|
||||
- **Three-way blends** with additional direction methods (diff_means, SOM)
|
||||
- **Post-blend recovery** via QLoRA fine-tuning on capability data
|
||||
- **Formal capacity hypothesis validation** using the provided experiment framework
|
||||
- **Post-blend recovery**, tracked separately in [issue #133][recovery-issue]
|
||||
- **Formal capacity-hypothesis validation**, tracked in [issue #132][capacity-issue]
|
||||
|
||||
## References
|
||||
|
||||
- [Arditi et al., *Refusal in Language Models Is Mediated by a Single Direction*][arditi]
|
||||
- [Belrose et al., *LEACE: Perfect Linear Concept Erasure in Closed Form*][leace]
|
||||
- [Wortsman et al., *Model Soups*][model-soups]
|
||||
- [Yadav et al., *TIES-Merging*][ties]
|
||||
|
||||
[arditi]: https://arxiv.org/abs/2406.11717
|
||||
[leace]: https://arxiv.org/abs/2306.03819
|
||||
[model-soups]: https://proceedings.mlr.press/v162/wortsman22a.html
|
||||
[ties]: https://arxiv.org/abs/2306.01708
|
||||
[capacity-issue]: https://github.com/elder-plinius/OBLITERATUS/issues/132
|
||||
[recovery-issue]: https://github.com/elder-plinius/OBLITERATUS/issues/133
|
||||
|
||||
@@ -2,11 +2,18 @@
|
||||
|
||||
**OBLITERATUS Project — August 2026**
|
||||
|
||||
> **Evidence status:** This summary preserves contributor-reported preliminary measurements from
|
||||
> PR #127. The PR did not include raw benchmark outputs, prompt-level refusal evidence, exact model
|
||||
> revisions, environment capture, or artifact hashes. Maintainers independently verified the blend
|
||||
> implementation and its CPU contracts, but not the numerical research results below.
|
||||
|
||||
---
|
||||
|
||||
## The Problem
|
||||
|
||||
All prior abliteration techniques face a fundamental tradeoff: deeper refusal removal causes greater capability loss. This tradeoff appeared to be intrinsic to the geometry of refusal-trained models — removing refusal directions inevitably damages overlapping capability directions.
|
||||
The contributor framed existing abliteration approaches as trading deeper refusal removal for
|
||||
greater capability loss. Whether that pattern generalizes or follows from refusal-model geometry
|
||||
has not been established by the evidence included with PR #127.
|
||||
|
||||
| Approach | Refusal Rate | MMLU Delta | Source |
|
||||
|---|---|---|---|
|
||||
@@ -15,17 +22,23 @@ All prior abliteration techniques face a fundamental tradeoff: deeper refusal re
|
||||
| huihui-ai (1-dir, skip layers) | Low | ~0pp | Community |
|
||||
| OBLITERATUS V1 (5-dir SVD) | 0.0% | -6.0pp | This work |
|
||||
|
||||
Complete refusal removal (0%) seemed to require accepting significant capability loss. V1 proved 0% was achievable but at -6pp MMLU — a cost that users noticed and complained about.
|
||||
The contributor reported that a V1 configuration reached 0% refusal on its sampled prompts with a
|
||||
-6pp MMLU difference. Those figures are retained as an unverified observation, not proof of a
|
||||
general tradeoff.
|
||||
|
||||
## The Insight
|
||||
|
||||
Different direction-finding algorithms damage different regions of weight space.
|
||||
|
||||
**SVD (Singular Value Decomposition):** Extracts directions by maximizing captured variance. This is greedy — it grabs high-variance components that encode both refusal AND capability. Deep refusal removal, but collateral capability damage concentrated in high-variance weight regions.
|
||||
**SVD (Singular Value Decomposition):** Extracts high-variance directions. The contributor proposes
|
||||
that some directions encode both refusal and capability; this mechanism was not measured in PR #127.
|
||||
|
||||
**LEACE (Linear Erasure of Concept Embeddings):** Finds directions by minimizing mutual information between the concept (refusal) and the representation. Mathematically constrained to preserve maximum non-refusal information. Excellent capability retention, but conservative — leaves refusal residue in the generation pathway (attention projections, output heads).
|
||||
**LEACE (Linear Erasure of Concept Embeddings):** Provides closed-form linear concept erasure while
|
||||
minimizing distortion. The reported capability retention and generation-pathway residue remain
|
||||
contributor observations requiring artifact-backed reproduction.
|
||||
|
||||
These methods make **complementary errors.** SVD damages regions LEACE preserves. LEACE leaves residue in regions SVD cleans.
|
||||
The working hypothesis is that these methods make complementary errors. The PR does not measure
|
||||
their error correlation or task-vector geometry.
|
||||
|
||||
## The Method
|
||||
|
||||
@@ -35,11 +48,14 @@ Run both surgeries independently on the same base model, then interpolate in wei
|
||||
blended_weight = α × LEACE_weight + (1 - α) × SVD_weight
|
||||
```
|
||||
|
||||
We binary-searched α over {0.30, 0.50, 0.55, 0.60, 0.65, 0.70} using 15-subject MMLU and a 10-prompt usability check as the objective. The optimal ratio for Qwen3.8-27B was α = 0.60.
|
||||
The contributor searched α over {0.30, 0.50, 0.55, 0.60, 0.65, 0.70} using 15-subject MMLU and a
|
||||
10-prompt usability check, then selected α = 0.60 for Qwen3.8-27B. The small, incomplete search
|
||||
does not establish a unique optimum.
|
||||
|
||||
**Why interpolation works:** Where SVD damaged capability, LEACE's intact weights dilute the damage. Where LEACE left refusal residue, SVD's clean weights dilute the residue. The blend point exists because these error distributions are approximately complementary — not identical, not orthogonal, but anti-correlated enough that averaging produces a model better than either parent.
|
||||
**Proposed explanation:** interpolation may dilute method-specific damage. This must be tested
|
||||
against same-method and unrelated-checkpoint blend controls before it is treated as causal.
|
||||
|
||||
## Results
|
||||
## Contributor-Reported Results (Not Independently Reproduced)
|
||||
|
||||
### Headline
|
||||
|
||||
@@ -49,9 +65,13 @@ We binary-searched α over {0.30, 0.50, 0.55, 0.60, 0.65, 0.70} using 15-subject
|
||||
| V1 (aggressive/SVD) | 81.4% (n=285) | 0.0% | 80% |
|
||||
| **V2 (60/40 blend)** | **86.3% (n=570)** | **0.0%** | **100%** |
|
||||
|
||||
V2 achieves +1.1pp MMLU above stock while maintaining complete refusal removal. This is the first reported instance of an abliterated model exceeding stock capability on a standard benchmark.
|
||||
The contributor reported +1.1pp MMLU above stock while maintaining complete refusal removal. The
|
||||
repository does not claim priority or a confirmed capability improvement without reproducible raw
|
||||
evidence and appropriate statistical comparison.
|
||||
|
||||
**Important caveat:** MMLU was run with `--limit 10` (570 questions, 10 per subject). This is above typical spot-check sample sizes but below the full 14,042-question MMLU benchmark. Full-scale validation is in progress. The +1.1pp result should be interpreted as "strong preliminary evidence of capability retention or improvement" rather than a definitive measurement.
|
||||
**Important caveat:** the contributor reports MMLU with `--limit 10` (570 questions, 10 per
|
||||
subject), below the full benchmark. Without raw outputs and a paired statistical analysis, the
|
||||
+1.1pp difference is descriptive only and may reflect sampling variation.
|
||||
|
||||
### Per-Subject Analysis
|
||||
|
||||
@@ -67,7 +87,8 @@ Gains span both safety-adjacent and neutral reasoning topics (5 questions per su
|
||||
| High School Chemistry | 100% | 80% | -20pp | Neutral (regression) |
|
||||
| College Computer Science | 80% | 60% | -20pp | Neutral (regression) |
|
||||
|
||||
The presence of gains on neutral topics (math, logic) suggests the improvement is not solely attributable to reduced hedging on sensitive questions. However, per-subject samples are too small for statistical significance.
|
||||
The reported neutral-topic differences motivate a reduced-hedging control, but the per-subject
|
||||
samples are too small to distinguish a capability effect from sampling variation.
|
||||
|
||||
### Practical Capability
|
||||
|
||||
@@ -81,7 +102,8 @@ The presence of gains on neutral topics (math, logic) suggests the improvement i
|
||||
| Structured output (JSON schema) | ✓ | ✓ | — |
|
||||
| System design | ✓ | ✓ | — |
|
||||
|
||||
V2 matches stock on every practical task tested while maintaining 0% refusal.
|
||||
The contributor reported comparable outcomes on this small practical-task set. The underlying
|
||||
tasks and outputs were not included for independent review.
|
||||
|
||||
## What We Don't Know Yet
|
||||
|
||||
@@ -94,7 +116,8 @@ V2 matches stock on every practical task tested while maintaining 0% refusal.
|
||||
- **Reduced hedging:** Stock model hedges on questions adjacent to sensitive topics; abliteration removes the hedging. Gains on law/ethics support this.
|
||||
- **Blend regularization:** Weight averaging of any two diverse models acts as implicit regularization (analogous to model soups/ensembling). The improvement may not be specific to abliteration.
|
||||
|
||||
3. **Is the 60/40 ratio model-specific?** We only tested on Qwen3.8-27B. The optimal ratio likely varies by architecture, model size, and alignment training method.
|
||||
3. **Is the 60/40 ratio model-specific?** The contributor reported testing only Qwen3.8-27B.
|
||||
Useful ratios may vary by architecture, model size, and alignment training method.
|
||||
|
||||
4. **Does this generalize beyond SVD + LEACE?** Other direction-finding methods (diff_means, SOM, nuclear/SAE) may offer additional complementary error profiles for three-way or N-way blends.
|
||||
|
||||
@@ -111,13 +134,15 @@ V2 matches stock on every practical task tested while maintaining 0% refusal.
|
||||
|
||||
## Experimental Framework for Future Validation
|
||||
|
||||
We built (but have not yet run) four experiments to distinguish between the competing hypotheses:
|
||||
Four proposed experiments are tracked in [issue #132](https://github.com/elder-plinius/OBLITERATUS/issues/132):
|
||||
|
||||
### Experiment 1: Activation Rank Analysis (Tests "freed capacity")
|
||||
Run diverse prompts through stock and abliterated models. Capture hidden states at each layer. Compute effective rank via SVD. If abliteration frees capacity, the effective dimensionality of activations should increase.
|
||||
|
||||
### Experiment 2: Topic Cluster Analysis (Tests "reduced hedging")
|
||||
Compare per-subject MMLU gains between sensitive topics (ethics, law, medicine) and neutral topics (physics, math). If gains cluster exclusively on sensitive topics, the improvement is hedging reduction, not capability gain. Preliminary results show mixed distribution — both types gain.
|
||||
Compare per-subject MMLU differences between prespecified sensitive and neutral groups. The test
|
||||
must define its statistical decision rule before examining results; the small table above is not
|
||||
such a test.
|
||||
|
||||
### Experiment 3: Blend Control (Tests "blend regularization")
|
||||
Blend two identical SVD surgeries (same method, different random seeds) at 60/40. If this blend also gains MMLU, the improvement comes from weight averaging itself, not from the SVD/LEACE complementarity. This is the critical control experiment.
|
||||
@@ -125,23 +150,31 @@ Blend two identical SVD surgeries (same method, different random seeds) at 60/40
|
||||
### Experiment 4: Learning Absorption (Tests "freed capacity" directly)
|
||||
QLoRA fine-tune both stock and abliterated models on identical small datasets. Compare loss curves. If the abliterated model learns faster (lower loss at same step count), it has more absorptive capacity — direct evidence for freed representational space. Requires GPU infrastructure (A100+, not feasible on MPS).
|
||||
|
||||
Code for all four experiments: `obliteratus/capacity_hypothesis.py`
|
||||
Their implementation is intentionally deferred until the hypotheses, controls, provenance, CPU
|
||||
contracts, and conditional GPU/network gates are specified.
|
||||
|
||||
## Future Directions
|
||||
|
||||
### Near-term (validated technique, ready to explore)
|
||||
### Near-term (implementation available; research validation pending)
|
||||
|
||||
1. **Cross-architecture replication.** Run the identical pipeline on Llama-3.1-70B, Gemma-2-27B, Mistral-Large. If the technique generalizes, it becomes a universal abliteration upgrade. The recipe is model-agnostic — only the blend ratio needs tuning per model.
|
||||
1. **Cross-architecture replication.** Run the identical pipeline on additional model families.
|
||||
This is needed to determine whether the technique generalizes and which parts of the recipe
|
||||
require model-specific tuning.
|
||||
|
||||
2. **Full-scale benchmarking.** Complete MMLU (14k), MMLU-Pro, HumanEval, GSM8K, ARC-Challenge on V2 to establish definitive capability numbers. Publish a proper eval table that the community can cite.
|
||||
2. **Full-scale benchmarking.** Complete MMLU (14k), MMLU-Pro, HumanEval, GSM8K, and
|
||||
ARC-Challenge with pinned inputs, raw results, and uncertainty estimates.
|
||||
|
||||
3. **N-way blending.** Blend three or more surgeries using different direction methods (SVD, LEACE, diff_means, SOM). If each adds complementary error cancellation, the optimal blend of N methods should outperform any pair.
|
||||
3. **N-way blending.** Test three or more surgeries using different direction methods (SVD,
|
||||
LEACE, diff_means, SOM) against prespecified pairwise and same-method controls.
|
||||
|
||||
4. **Blend ratio as a function of model properties.** Study how the optimal α relates to model size, architecture, alignment training intensity, and number of refusal directions. Build a predictor so users don't need to binary-search.
|
||||
4. **Blend ratio as a function of model properties.** Study how selected α values relate to model
|
||||
size, architecture, alignment training intensity, and number of refusal directions.
|
||||
|
||||
### Medium-term (theoretical, needs investigation)
|
||||
|
||||
5. **Post-blend capability recovery.** QLoRA fine-tune the blended model on a curated capability dataset (MMLU train split, code exercises, reasoning chains). If the "freed capacity" hypothesis holds, the abliterated model should absorb new capability faster than stock. We built the dataset (4,874 refusal-free examples) and the training code (`obliteratus/recover.py`) but MPS was insufficient for 27B QLoRA — needs A100+.
|
||||
5. **Post-blend capability recovery.** A provenance-safe dataset and QLoRA pipeline are proposed in
|
||||
[issue #133](https://github.com/elder-plinius/OBLITERATUS/issues/133). No corpus or recovery
|
||||
trainer is shipped by this change.
|
||||
|
||||
6. **SLERP and task-arithmetic blending.** Replace LERP with spherical interpolation (preserves weight norms) or task-arithmetic approaches (TIES-Merging, DARE) that handle parameter conflicts more intelligently. LERP is the simplest possible blend — there is likely headroom from more sophisticated interpolation.
|
||||
|
||||
@@ -151,7 +184,10 @@ Code for all four experiments: `obliteratus/capacity_hypothesis.py`
|
||||
|
||||
### Long-term (speculative, high-impact if true)
|
||||
|
||||
9. **Capacity hypothesis validation and exploitation.** If Experiment 1 confirms that abliteration increases effective activation rank, this has implications beyond abliteration — it suggests that safety training in general consumes representational capacity that could be allocated to capability. This would mean: (a) safety-capability tradeoffs are not fundamental but artifacts of training methodology, and (b) better alignment techniques could achieve safety without capacity cost.
|
||||
9. **Capacity-hypothesis validation.** Increased activation rank alone would not demonstrate freed
|
||||
representational capacity; the proposed work must control for prompt sampling, layer selection,
|
||||
model identity, numerical thresholds, and alternative explanations before drawing implications
|
||||
about safety training.
|
||||
|
||||
10. **Generalized complementary merging.** The principle — "combine models that fail in different ways" — may extend beyond abliteration to any model merging scenario. Fine-tunes optimized for different objectives (code, math, reasoning) could be blended using the same complementary error cancellation principle, with direction-specific merge ratios instead of uniform interpolation.
|
||||
|
||||
@@ -174,7 +210,7 @@ obliteratus obliterate $BASE --method aggressive --direction-method leace \
|
||||
|
||||
# Step 3: Blend
|
||||
obliteratus blend --model-a surgery_svd --model-b surgery_leace \
|
||||
--alpha 0.6 --output blended
|
||||
--alpha 0.6 --config-source a --output blended
|
||||
|
||||
# Step 4: Validate
|
||||
lm_eval --model hf --model_args pretrained=blended --tasks mmlu --device auto
|
||||
@@ -182,6 +218,11 @@ lm_eval --model hf --model_args pretrained=blended --tasks mmlu --device auto
|
||||
|
||||
All code is open source: [github.com/elder-plinius/OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS)
|
||||
|
||||
Primary background: [Arditi et al.](https://arxiv.org/abs/2406.11717),
|
||||
[LEACE](https://arxiv.org/abs/2306.03819),
|
||||
[Model Soups](https://proceedings.mlr.press/v162/wortsman22a.html), and
|
||||
[TIES-Merging](https://arxiv.org/abs/2306.01708).
|
||||
|
||||
---
|
||||
|
||||
*OBLITERATUS Contributors, August 2026*
|
||||
|
||||
Reference in New Issue
Block a user