mirror of
https://github.com/elder-plinius/OBLITERATUS.git
synced 2026-08-18 00:47:23 +02:00
Four bugs prevented bitsandbytes 4-bit quantized models from completing ablation studies on GPUs with 16GB VRAM: 1. runner.py: quantization parameter was never passed from StudyConfig to load_model(), so the loader had no idea quantization was enabled. 2. loader.py (max_memory): GPU memory budget was calculated against the unquantized model size, causing accelerate to offload layers to meta device even though the quantized model fits comfortably. Now divides estimate by 4 (4-bit) or 2 (8-bit) before deciding. 3. evaluator.py: empty strings in wikitext dataset caused zero-length tensors that crashed the forward pass with a reshape error. Now filters empty/whitespace-only texts and skips empty batches. 4. loader.py (snapshot/restore): snapshot skip decision used unquantized size estimate, and restore used strict=True which rejects bitsandbytes metadata keys (.absmax, .quant_map, .quant_state). Now uses quantized estimate and strict=False. Tested on RTX 5060 Ti (16GB) with Qwen2.5-Coder-7B-Instruct in 4-bit. Quick Scan (layer_removal + ffn_ablation) completes all 56 specs.