mirror of
https://github.com/wiltodelta/remove-ai-watermarks.git
synced 2026-08-09 23:50:40 +02:00
The sdxl-zimage profile crashed on every image containing a face. The remover gives it torch.float16, because SDXL ships fp16 weights and an fp16-safe VAE, and that dtype reached the inherited _load_zimage while _zimage_vram_config hardcodes bfloat16 for its offload, onload and computation dtypes. Z-Image was therefore built bf16 and handed fp16 latents, dying in the VAE with "Input type (c10::Half) and bias type (c10::BFloat16) should be the same". Zero-face inputs never enter _run_faces, so the profile passed every timing run it was given, and its tests avoid model downloads, so nothing exercised the loader. Every face-stage loader now reads _face_stage_dtype(), the computation dtype of the VRAM config it is paired with. SAM is routed through it too: it never crashed, since it casts its own inputs and leaves through .float(), but it read the same field and would have re-landed the bug for the next profile with a different global dtype. That field was never the global dtype on this profile anyway - _load_sdxl hardcodes fp16 for its own ControlNet, VAE and pipeline - so its only readers were face-stage code. This also fixes a second instance transitively: the persisted prompt-embedding cache restores payloads at the DiffSynth pipe's dtype, which was fp16 into a bf16 stack before this change. For qwen-zimage the whole change is a strict no-op. The remover already hands it bfloat16, the same value _face_stage_dtype() returns, so production is untouched; verified on an H100 against the deployed pin. The guard asserts the dtype the Z-Image and SAM loaders actually receive rather than comparing the accessor to the config it derives from, which would restate the implementation and pass for any consistently wrong value. Both assertions were mutation-tested against the pre-fix line. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>