mirror of
https://github.com/wiltodelta/remove-ai-watermarks.git
synced 2026-08-10 08:00:32 +02:00
Keep the qwen-zimage global stack resident on a card that can hold it
The mandatory Qwen stack was configured to offload to disk unconditionally. DiffSynth implements that by dropping the weights to the meta device and re-reading every parameter through its DiskMap on the next onload, and the pipeline moves between text encoder, transformer and VAE on every pass, so each generation paid a full model reload. That is the right trade on a consumer card, where it is what makes a 20B model runnable at all, and pure waste on a card that can simply hold the stack. Residency is now resolved from total VRAM, mirroring how the optional Z-Image face stack is already gated. Above the floor the config passes no "disk" value anywhere, which is what actually disables the behavior: DiffSynth latches disk_offload once from offload_dtype, so pointing every device at CUDA while leaving the sentinel would keep both the meta-drop and the re-read. The floor is set equal to the face floor rather than lower because that is the configuration measured with both stacks resident; a tighter gate is plausible but unvalidated. cpu_offload now forces both stacks to stream, so a caller asking for low VRAM no longer gets the larger stack pinned anyway. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
643557b7bf
commit
a64d81f446
@@ -293,11 +293,32 @@ Regression coverage:
|
||||
|
||||
CPU offload is enabled only when requested on CUDA. The standard Diffusers
|
||||
profiles call `enable_model_cpu_offload`. The `qwen-zimage` profile uses the
|
||||
same flag to force its face stack out of automatic device residency.
|
||||
same flag to force **both** its stacks out of automatic device residency.
|
||||
|
||||
Residency is otherwise chosen from the card's total VRAM, once per stack:
|
||||
`resolve_global_model_residency` gates the mandatory Qwen stack at
|
||||
`RESIDENT_GLOBAL_MODEL_MIN_VRAM_GIB` and `resolve_face_model_residency` gates the
|
||||
optional Z-Image stack at `RESIDENT_FACE_MODEL_MIN_VRAM_GIB`.
|
||||
|
||||
Below the global floor, `_qwen_vram_config` streams the stack from disk, which is
|
||||
what makes a 20B model runnable on a consumer card. At or above it, streaming is
|
||||
pure waste and the weights stay on the GPU. The difference is not marginal:
|
||||
DiffSynth offloads by dropping the weights to the meta device and re-reading every
|
||||
parameter through its `DiskMap` on the next onload, and the pipeline moves between
|
||||
text encoder, transformer and VAE on each pass. Measured on an H100 (80 GiB), a warm
|
||||
global pass took 37.3 s at 0.8 GiB resident with the streaming config, against 2.2 s
|
||||
at 28.7 GiB with the stack resident; both stacks resident peaked at 48.0 GiB. Faster
|
||||
storage cannot close that gap, because the cost is the reload itself rather than the
|
||||
read.
|
||||
|
||||
The resident config deliberately passes no `"disk"` value anywhere. DiffSynth latches
|
||||
`disk_offload` once, from `offload_dtype`, so leaving the sentinel in place while
|
||||
pointing every device at CUDA would keep the meta-drop and re-read.
|
||||
|
||||
Regression coverage:
|
||||
|
||||
- [`test_cpu_offload.py`](../tests/test_cpu_offload.py)
|
||||
- [`test_qwen_zimage_pipeline.py`](../tests/test_qwen_zimage_pipeline.py)
|
||||
|
||||
### Qwen plus Z-Image
|
||||
|
||||
|
||||
Reference in New Issue
Block a user