mirror of
https://github.com/wiltodelta/remove-ai-watermarks.git
synced 2026-08-09 23:50:40 +02:00
Both qwen-zimage stages prompt with module constants, and at CFG 1.0 DiffSynth's PipelineUnitRunner reuses the positive embedding for the negative side rather than encoding it, so exactly one embedding per stage is ever computed. Persist it and neither text encoder has to be loaded at all. Measured on an H100 volume: this drops 15.45 GiB (Qwen2.5-VL) and 7.49 GiB (Z-Image) of an 87.6 GiB per-request read, worth a median 11.76 s and 4.10 s of load time paired within five containers. A nine-face fixture returned sha256 c8567e11077de32a both with and without the cache, so the output is byte-identical and the provider-oracle clearance is untouched. The cache key carries the cache version, model id, pipeline output params and the exact prompt, so a model bump or a prompt edit recomputes instead of reading a stale embedding. The write is atomic because a torn file must never read back as a hit, and a miss after the text encoder was already dropped raises rather than calling a model that is not loaded. _model_cache_dir now prefers HF_HOME: on a scale-to-zero runner that is the only persistently mounted path, so anything below it is re-derived every request. The YuNet download follows the same root. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>