Merge pull request #83 from wiltodelta/chore/text-preservation-evaluation

Add verified text restoration
This commit is contained in:
Victor Kuznetsov
2026-08-15 12:30:27 -07:00
committed by GitHub
35 changed files with 2582 additions and 98 deletions
+13
View File
@@ -199,6 +199,19 @@ image came from an AI generator, add `--force`:
remove-ai-watermarks invisible image.png -o clean.png --force
```
Typography-heavy images can opt into the experimental verified-text post-pass.
It requires manually reviewed strings and line boxes; it never trusts OCR as ground
truth or runs automatically:
```bash
uv tool install --force "remove-ai-watermarks[text-restoration]"
remove-ai-watermarks invisible image.png -o clean.png \
--text-manifest verified-lines.json --force
```
See the [CLI guide](docs/cli.md#restore-operator-verified-text) for the manifest
schema, compatibility restrictions, and oracle caveats.
See the [installation guide](docs/installation.md) for Homebrew, uv, optional
features, and development setup.
+366 -3
View File
@@ -7,31 +7,394 @@ ground truth and instructions.
| Original | Provider | Content | Exercises |
| --- | --- | --- | --- |
| `ChatGPT Image May 31, 2026, 02_02_23 PM.png` | OpenAI | Light multilingual typography | Text preservation |
| `ChatGPT Image May 31, 2026, 02_03_55 PM.png` | OpenAI | Multilingual typography | Text preservation |
| `Gemini_Generated_Image_633uuy633uuy633u.png` | Google | Landscape with a Chinese sign | CJK text preservation |
| `Gemini_Generated_Image_y48j3cy48j3cy48j.png` | Google | Portrait grid | Face identity and skin texture |
## Text ground truth
`ground-truth.json` contains hand-verified OCR for the two text-bearing
`ground-truth.json` contains hand-verified OCR for the three text-bearing
originals. To regenerate an OCR seed:
```bash
uv run scripts/fidelity_metrics.py ocr \
"data/synthid/originals/ChatGPT Image May 31, 2026, 02_03_55 PM.png" \
"data/synthid/originals/ChatGPT Image May 31, 2026, 02_02_23 PM.png" \
data/synthid/originals/Gemini_Generated_Image_633uuy633uuy633u.png \
--langs en,ru,ch \
--out data/evaluations/fidelity/ground-truth.json
```
Verify and correct the generated text by hand before using it as ground truth.
`text-lines.json` contains the verified per-line strings and source-space boxes
used by the evaluation-only selective-restoration experiment. It is not an
automatic scene-text annotation set.
`scripts/infer_text_lines.py` can generate a draft from source pixels without
existing annotations. On the two posters it proposed 20 and 18 lines at the
default threshold, but exact-text precision was only 90.0% and 94.4% because
high-confidence OCR still dropped an English comma and replaced a Chinese comma
with ASCII. Its output therefore requires manual verification of every line;
`accepted` means crop-stable, not ground-truth-correct.
## Text-preservation benchmark
`text-preservation-2026-08-13.csv` records a fixed-seed comparison of the two
current profiles and a global Z-Image Turbo prototype on all three text
fixtures. Each candidate ran through the
complete `visible -> invisible -> metadata` route with its profile default
strength and adaptive-polish setting. The Z-Image prototype instead sweeps
0.08, 0.10, 0.15, 0.20, and 0.25 without polish because its provider-specific
removal floor is not known. The output hash identifies the exact bytes
measured; generated outputs remain outside the repository.
The character-weighted added CER is 0.262 for `qwen-zimage` and 0.256 for
`sdxl-zimage`. That 0.006 absolute difference is not a
stable ordering: SDXL wins the light poster, Qwen wins the dark poster, and the
Chinese sign is tied. A paired sign test on the two non-ties is 1-1 (`p=1.0`).
The measured sample therefore does not support a general text-preservation
winner. Both profiles substantially degrade the smallest multilingual poster
text and preserve the larger Chinese sign.
Z-Image Turbo is the clear fidelity lead. At strength 0.10 its
character-weighted CER is 0.093, against the unchanged sources' 0.124 OCR
floor, while whole-image LPIPS stays between 0.047 and 0.086. Visual inspection
still finds substitutions in the smallest Cyrillic and Chinese poster lines,
so a favorable OCR score does not mean pixel-exact text. Quality drops quickly
above 0.15; at 0.25 the weighted CER is 0.428.
Provider-oracle checks on 2026-08-13 bracket the OpenAI requirement at no more
than 0.10 for this sample. Both original OpenAI controls were detected by
`openai.com/verify`; the light poster was clean at 0.08, while the dark poster
was detected at 0.08 and clean at 0.10. The light 0.10 output was not separately
checked. Google fails the entire tested sweep. The original Gemini image was
detected through C2PA, and a pixel-identical copy with AI metadata stripped was
separately detected by Gemini's built-in SynthID verifier. Z-Image outputs at
0.08, 0.10, 0.15, 0.20, and 0.25 were all still detected. At 0.25 the weighted
CER has already risen to 0.428, worse than the raw weighted CER of both current
profiles (Qwen 0.387, SDXL 0.381). Increasing strength beyond the measured grid
would therefore no longer serve the text-preservation objective without a new
mechanism or hypothesis. Z-Image is not a viable global replacement on this
evidence: no clean Google operating point was found before it lost its fidelity
advantage.
Qwen-Image-2.0 was not added to the numeric comparison. Its weights are not
published, and its hosted edit API exposes an editing instruction and seed but
no low-strength denoise control. It can be evaluated as a separate hosted edit
strategy when credentials are available, but it is not a drop-in replacement
for the partial-regeneration mechanism measured here.
The OCR floor is the source image scored against the hand-verified text. Use
`added_cer = text_cer - ocr_floor` when interpreting pipeline damage, because the
unchanged poster sources already score 0.127 CER. `oracle_rechecked=true` marks
the exact Z-Image bytes checked above; the remaining rows were not rechecked.
The table does not certify other seeds, content classes, or strengths beyond
the recorded provider verdicts.
## Text-restoration prototype
`text-restoration-2026-08-13.csv` evaluates an OCR-driven post-pass on the exact
Qwen outputs above. The prototype recognizes English and Russian with macOS
Vision and CJK with PaddleOCR, derives glyph masks independently from the
source and Qwen output, removes both sets of glyphs with block-wise LaMa, and
draws the recognized strings with new system-font pixels. It never composites
source pixels back into the result.
On the two multilingual posters, character-weighted CER fell from 0.338 and
0.305 to 0.007 on both. OpenAI Verify reported no OpenAI signals for both Qwen
controls and both restored outputs in the same run. The improvement comes with
a substantial whole-image fidelity cost: LPIPS rose from 0.107 to 0.174 and
from 0.095 to 0.162, while PSNR fell by about 10.5-11.4 dB. Visual inspection
found one residual shadow in the smallest English line of the light poster;
the dark poster was clean but the substitute fonts visibly changed typography.
The Chinese sign did not improve: CER rose from 0.074 to 0.111 because the OCR
and renderer changed punctuation. The Gemini verifier returned detected for
both the restored output and its byte-identical Qwen control on a third work
account, although that Qwen hash had previously returned clean on another work
account. This run therefore does not isolate a restoration-stage regression;
its Google verdict is inconclusive until a source-positive, Qwen-negative
control can be reproduced in the same available account.
This prototype is not ready to ship. Its strongest result establishes that
fresh-glyph reconstruction can recover literal text without reintroducing an
OpenAI signal, but portable OCR, font/style reconstruction, a tighter mask, and
a reproducible Google oracle control remain prerequisites.
### Selective restoration follow-up
`selective-text-restoration-2026-08-13.csv` compares that full compositor with
a selective prototype on the same two posters. The prototype leaves a Qwen line
unchanged when padded source and output recognition agree, and applies the same
LaMa plus fresh-system-font reconstruction only to lines whose recognized text
changed. This reduced the edited area from 15.4% and 17.0% to 5.5% on both
posters.
Under one consistent Paddle `en+ru+ch` measurement route, selective restoration
reduced Qwen CER from 0.378 to 0.101 on the light poster and from 0.413 to 0.112
on the dark poster. Its image LPIPS was 0.120 and 0.103, substantially closer to
Qwen than the full compositor's 0.174 and 0.162. These CER values must not be
mixed with the preceding table's Vision/Paddle hybrid values: the comparison
file remeasures all three variants through Paddle so their relative result is
valid on one OCR route.
OpenAI Verify returned `No OpenAI signals detected` for both selective outputs,
then detected the original light poster as `Generated with OpenAI tools` in the
same Chrome sequence. Visual inspection found the dark output clean, but the
light output still retained a local shadow around one replaced fine-text line.
Selective restoration is therefore the strongest current direction, not a
production-ready default. The next implementation needs automatic line matching
and a tighter source-plus-candidate glyph mask before it can be proposed for the
pipeline.
A mask-only follow-up added two pixels of dilation around every selected glyph.
It visually removed the light poster's shadow and improved its CER from 0.101 to
0.081 with effectively unchanged LPIPS, but increased the edited area from 5.5%
to 7.0%. Applying the same rule to the dark poster worsened CER from 0.112 to
0.147. These expanded-mask rows were not resubmitted to the provider oracle.
The opposing result rules out one global dilation constant and makes per-line
residual detection the next fidelity gate.
An automatic follow-up removed the hand-selected line list. For each known
source line, it runs a recognition-only model on padded, upscaled source and
Qwen crops. It preserves a line only when both confidence scores are at least
0.75, the source recognition is within 0.25 normalized edit distance of the
verified line, and normalized source and candidate recognition match exactly.
This reproduced the manual dark-poster line selection and additionally kept one
correct large Russian line on the light poster.
The automatic prototype also replaced uniform dilation with a second measured
LaMa pass. After the first erase, it extracts contrast remaining at the original
glyph positions, dilates only that residual, and inpaints it before drawing fresh
text. Visual inspection found no remaining double glyphs on either output. The
effective edited fractions were 4.1% and 5.7%. Image LPIPS was 0.113 and 0.104,
against 0.120 and 0.103 for the manual selective version. Detector-based page CER
was 0.123 and 0.119; the light value is inflated by the detector splitting the
preserved Cyrillic line into Latin-like fragments even though the padded Russian
recognizer reads it exactly. This is another reason that line decisions use
recognition-only crops rather than the page detector's aggregate CER.
Both exact automatic output hashes returned `No OpenAI signals detected`, and
the original light poster returned `Generated with OpenAI tools` immediately
afterward in the same Chrome sequence. The tracked script reproduced the dark
hash exactly. Its light output differed from the initial prototype in only 386
pixels by at most one channel value because separate ONNX environments rounded
differently; the script's exact `900def5a...` hash was therefore submitted and
also returned clean before the positive source control.
The script and its required verified line annotations are tracked as
`scripts/selective_text_restoration.py` and `text-lines.json`. For example:
```bash
PADDLE_PDX_DISABLE_MODEL_SOURCE_CHECK=True \
uv run scripts/selective_text_restoration.py \
data/synthid/originals/Gemini_Generated_Image_633uuy633uuy633u.png \
path/to/qwen-candidate.png \
--restoration vae-glyphs \
--glyph-donor path/to/qwen-vae-reconstruction.png \
--glyph-feather 0.5 \
--output path/to/selective-output.png \
--manifest path/to/manifest.json
```
The restoration mode is deliberately explicit. `vae-glyphs` requires a
separately generated VAE reconstruction in `--glyph-donor`; it takes exact
thresholded glyph cores from that donor, uses a fresh silhouette beneath them,
and defaults to a narrow 0.5-pixel donor edge. The donor and the resulting exact bytes still require
full-pipeline oracle evaluation. `rerender` reproduces the tracked
fresh-system-font experiment. `source-glyphs` is only for compositing a
separately regenerated, oracle-evaluated layer; feeding it the watermarked
original would paste provenance-bearing pixels back into the result. The
`--keep-background --composite-mask boxes` combination is an aligned-layer
experiment and is not a production text restorer.
The `source-silhouette` follow-up keeps the source glyph geometry but discards
its pixel amplitudes: it thresholds each source line to a binary shape, samples
one foreground color, and synthesizes new antialiasing over the scrubbed image.
On a 24-image matrix spanning serif, Latin-diacritic, Cyrillic, CJK, tiny-UI,
and rotated text, the original outer-feather compositor preserved all 34
source-readable lines exactly under crop OCR, with median text-box SSIM 0.859.
An inner-antialias variant raised median text-box SSIM to 0.902 and left
whole-image SSIM unchanged at 0.672, but preserved 33/34 lines: OCR read one
middle dot as a colon. The higher-fidelity antialiasing is retained in the
evaluation script, with that punctuation miss recorded as a caveat. Its exact
outputs were not submitted to the provider oracle because the
public verifier reached its request limit; a clean verdict from the earlier
outer-feather bytes does not transfer to the new hashes.
Follow-up visual review on a typography-rich control rejected both silhouette
compositors despite their OCR and SSIM scores. They preserved literal content
and approximate glyph geometry, but changed stroke weight, color variation,
edge antialiasing, and small decorative details enough to be plainly different
from the source. OCR exactness and text-box SSIM are therefore screening metrics,
not acceptance gates for source-typography preservation. A candidate must also
retain source-like edge pixels and pass direct visual comparison at native size.
A narrower Qwen-VAE donor follow-up keeps the scrubbed background, then copies
only VAE-reconstructed pixels through the source-silhouette mask with a
0.5-pixel feather. Across the 48-case typography matrix (548 annotated lines),
text-box SSIM improved in 47/48 cases and glyph-edge error improved in 48/48.
Median text-box SSIM rose from 0.854 to 0.914, while median glyph-edge MAE fell
from 37.59 to 32.42. The median nonzero alpha area was 3.62%; on dense or large
text this is still substantially wider than the target control. These are
fidelity results only; the 48 matrix outputs were not submitted to the provider
oracle. On a separate untracked dense-typography control, the exact core-only
Qwen-VAE donor with a 1.43% source-silhouette area returned `No OpenAI signals
detected`. Its 0.5-pixel feathered sibling, which raises the nonzero donor area
to 2.75%, returned the same verdict and improved mean text-box SSIM from 0.957
to 0.960. Crop OCR recovered from 7/15 exact lines on the raw pass to 14/15,
matching the source OCR floor. Whole-image LPIPS was 0.082, but only 0.108% of
pixels were exactly unchanged and the detected face retained 0.670 of source
Laplacian variance. The tracked script reproduced the feathered file byte for
byte. These two exact-byte verdicts do not certify other images or the larger
matrix masks, and the global smoothing fails a strict unchanged-image criterion.
The opt-in production port was rechecked separately on 2026-08-15. Its current
LaMa runtime did not reproduce the earlier evaluation PNG byte for byte, but all
changed pixels were confined to the erased background outside the donor glyph
core. The exact production artifact returned `No OpenAI signals detected` in
3/3 OpenAI Verify runs, while the matched source control returned `Generated
with OpenAI tools` in 2/2 runs in the same Chrome session; expanded details
identified SynthID and no C2PA manifest on the control. The private control and
artifact hashes remain outside the public repository. This certifies only that
runtime, verified manifest, and output, not arbitrary text masks or images.
The Google result is negative. On the synthetic CJK sign case, two separate
work-account runs both detected SynthID in the resaved source control and in the
exact Qwen-VAE donor output. The candidate improved mean text-box SSIM from
0.791 to 0.835 and glyph-edge MAE from 35.48 to 22.49 with a 3.67% donor area,
but Gemini still detected it. The intermediate Qwen silhouette base and the
earlier original-sign Qwen baseline were also detected in the same account.
This isolates the current blocker upstream of text restoration: the global
Google pass did not reach the SynthID removal floor, so the OpenAI-clean
`vae-glyphs` configuration cannot be used as a Google operating point.
On the large Chinese-sign control, source and candidate recognition agreed on
all three lines. The script selected no changed lines, emitted a zero mask, and
copied the candidate byte for byte instead of re-encoding it. This corrects the
earlier unnecessary CJK rerender and proves the no-edit branch. It does not add
a Google-negative oracle result: the available Google account still cannot
reproduce a source-positive, Qwen-negative control sequence.
The earlier automatic rerender was the first variant in the experiment to pass
the recorded visual, fidelity, and OpenAI-oracle gates on its two poster
fixtures. The later typography-rich control shows that result does not
generalize to source-typography preservation. It remains evaluation-only: it
depends on verified source text and source line boxes, uses macOS system fonts,
and has not been validated on natural scene text, rotated text, or automatic
line-box discovery.
An opt-in `--detect-boxes` follow-up tested automatic geometry. Grouping Paddle
word detections by vertical overlap found exactly 20/20, 20/20, and 3/3 lines;
mean IoU with verified boxes was 0.857, 0.847, and 1.000. Reusing the annotation
crop padding was unstable and preserved only 4/20 dark-poster lines. Reducing
vertical recognition padding to 10% restored the exact 8/20 and 7/20 selection
decisions, but detector CER was 0.127 and 0.154 rather than 0.123 and 0.119. The
dark regression failed the fidelity gate, so these hashes were not submitted to
the provider oracle. The flag is retained only to reproduce that negative
evaluation and still requires verified strings and an exact line-count match.
### AnyText2 glyph-conditioned follow-up
`anytext2-restoration-2026-08-13.csv` tests the official Apache-2.0
AnyText2 checkpoint as a local text-editing pass over the exact Qwen Chinese-sign
output. The checkpoint hash was verified against ModelScope. Its own edit example
successfully replaced a masked blackboard line with the requested `DADDY`, so the
runtime and checkpoint were functional before the tracked fixture was measured.
AnyText2 failed the fidelity gate. The standard full-image detector scored the
default edit at CER 0.185 and the source-font mimic at 0.222, against 0.074 for
Qwen; image LPIPS moved from 0.289 to 0.338 and 0.345. A padded crop-recognition
cross-check removed the detector's punctuation misses: Qwen scored 0.000, the
default edit still scored 0.074 with two substituted characters, and the mimic
edit stayed at 0.222. Visual inspection agreed, with additional line-shape and
punctuation errors in the mimic output. No provider-oracle submission was made
because neither candidate preserved the requested content better than its
already-clean Qwen input.
This result covers the published SD1.5 checkpoint, three Chinese lines of no more
than 20 characters, one seed, and one fixture. The repository still describes
AnyText2XL as unreleased, and the inference wrapper truncates each requested line
after 20 characters, so the published model cannot directly cover the longer
English and Russian poster lines.
The padded cross-check also changes the interpretation of this fixture's earlier
restoration result. Its Qwen image already contains all three correct Chinese
lines; the reported 0.074 came from the detector cropping the final punctuation
from two lines. A deterministic rerender with horizontally padded recognition
and a mask-matched Hiragino Sans GB W6 font reached detector CER 0.000, but image
LPIPS worsened from 0.289 to 0.327 and the font was visibly heavier. The correct
policy on this large-text fixture is therefore to leave a Qwen line unchanged
when padded recognition confirms the source and candidate strings agree.
## Face-stage isolation
`face-stage-isolation-2026-08-13.csv` separates the two stages on the 18-face
Gemini portrait-grid fixture. All variants use the same visible-cleaned,
metadata-stripped input and seed 0. The prepared input remained positive in the
Gemini SynthID verifier, so the surrounding visible and metadata stages did not
create the clean verdicts.
| Route | Gemini oracle | ID cosine | Face LPIPS | Laplacian ratio | Image LPIPS |
| --- | --- | ---: | ---: | ---: | ---: |
| Prepared control | detected | - | - | - | - |
| Z-Image face stage only | detected | 0.927 | 0.028 | 0.612 | 0.013 |
| Qwen global stage only | clean | 0.589 | 0.217 | 0.331 | 0.288 |
| Full Qwen then Z-Image | clean | 0.852 | 0.050 | 0.492 | 0.252 |
This isolates the roles: the Qwen global pass removes the watermark, while the
masked Z-Image pass restores face identity and detail without reintroducing a
detectable whole-image signal. Z-Image face-only does not remove SynthID. The
result explains why the face stage succeeded while global Z-Image failed: the
earlier clean verdict belonged to the two-stage pipeline, not to Z-Image as a
standalone remover. This is one fixture, one seed, and one oracle pass per
variant; it does not establish a general reintroduction threshold for mask size.
## Upstream Synthid-Bypass v2 reproduction
`upstream-v2-reproduction-2026-08-13.csv` records a source-level audit and a
close reproduction of the public
[`Synthid-Bypass-v2.0.json`](https://github.com/cebeuq/Synthid-Bypass/blob/3007d0351596ae0a78b7074dae7ad179710b1e48/Synthid-Bypass-v2.0.json).
The audited workflow is SHA-256
`41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6` at
upstream commit `3007d0351596ae0a78b7074dae7ad179710b1e48`.
Tracing links backward from `SaveImage` confirms that the active global stage is
Qwen-Image-2512 Q4 with the Lightning four-step LoRA at 0.8, DiffSynth Canny at
1.0, AuraFlow shift 3, `dpmpp_2m` plus `sgm_uniform`, CFG 1, and four steps. Its
resolution helper uses adaptive level 6 over 0.08..0.15, which resolves to the
0.154 ceiling for the 2816x1536 portrait fixture. Z-Image Turbo is used only by
the face detailer: eight steps, CFG 1, `res_2s` plus `bong_tangent`, a 768 px
guide, 1024 px cap, and direct adaptive denoise
`clamp(0.10 * largest_face_ratio / 0.03, 0.05, 0.28)`. The active face regions
come from YOLOv8-face plus SAM. Although the README describes MediaPipe as part
of the strict path, the MediaPipe nodes do not reach the saved output. The
1.2-megapixel scaler is also bypassed (`mode=4`).
The published upstream pair 12 changed from SynthID-positive to clean in the
Gemini verifier and scored 0.975 face identity. On the project portrait fixture,
the prepared control was positive, while both the close global reproduction and
the close full reproduction were clean. The global stage scored 0.589 identity
and 0.217 face LPIPS; the upstream-strength face stage improved those to 0.783
and 0.083. The maintained profile's weaker face pass scored 0.852 and 0.050 on
the same control, so copying upstream's roughly doubled face strength would be a
quality regression on this fixture.
The close reproduction is intentionally labeled rather than presented as an
exact ComfyUI run. It uses full-precision DiffSynth bf16 weights instead of the
Q4 GGUF files, the maintained DiffSynth Lightning scheduler approximation
instead of the ComfyUI sampler pair, YuNet plus SAM instead of YOLOv8-face plus
SAM, and fixed seed 0 instead of randomized seeds. The two independent oracle
controls and the published upstream pair establish the direction of the result;
they do not make the local output byte-equivalent to upstream.
## Compare
```bash
uv run scripts/fidelity_metrics.py compare \
--original data/synthid/originals/Gemini_Generated_Image_y48j3cy48j3cy48j.png \
--variant controlnet=<out>.png \
--variant qwen=<out>.png \
--variant qwen-zimage=<qwen-out>.png \
--variant sdxl-zimage=<sdxl-out>.png \
--ocr-langs ""
```
@@ -0,0 +1,4 @@
date,source_commit,fixture,variant,model,model_revision,model_sha256,ocr_route,text_cer,img_lpips,ssim,psnr_db,output_sha256,provider_oracle,notes
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,qwen-zimage,,,,"Paddle ch",0.074,0.289,0.696,23.3,4b05489c7a79f993d35d32528de792baad5a3e33ee950afa147caf5d3e0621,not_run,Exact Qwen baseline reused from text-restoration benchmark
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,qwen-zimage-anytext2-default,AnyText2-SD1.5,b06c583a583818f3679665ef67b51363f107853c,7d5d593928d842500e0da101ab6a1ae28a8d21403441ca62c813c4f92529ace8,"Paddle ch",0.185,0.338,0.642,18.2,87729fe7e0482d2cead6d3b7c7dd386ecc75ab3d8658131c54e67528b8f9feeb,not_run,Fidelity gate failed; padded crop recognition CER 0.074 still shows two substituted characters; locally composited three-line edit; 20 DDIM steps; seed 0
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,qwen-zimage-anytext2-mimic,AnyText2-SD1.5,b06c583a583818f3679665ef67b51363f107853c,7d5d593928d842500e0da101ab6a1ae28a8d21403441ca62c813c4f92529ace8,"Paddle ch",0.222,0.345,0.644,17.9,f5198d8f2365b79880aba8ea6b0c3885142b9de21c9f7f4f6712ac4a0db8dd85,not_run,Fidelity gate failed; source glyph masks supplied as font hints; 20 DDIM steps; seed 0
1 date source_commit fixture variant model model_revision model_sha256 ocr_route text_cer img_lpips ssim psnr_db output_sha256 provider_oracle notes
2 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png qwen-zimage Paddle ch 0.074 0.289 0.696 23.3 4b05489c7a79f993d35d32528de792baad5a3e33ee950afa147caf5d3e0621 not_run Exact Qwen baseline reused from text-restoration benchmark
3 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png qwen-zimage-anytext2-default AnyText2-SD1.5 b06c583a583818f3679665ef67b51363f107853c 7d5d593928d842500e0da101ab6a1ae28a8d21403441ca62c813c4f92529ace8 Paddle ch 0.185 0.338 0.642 18.2 87729fe7e0482d2cead6d3b7c7dd386ecc75ab3d8658131c54e67528b8f9feeb not_run Fidelity gate failed; padded crop recognition CER 0.074 still shows two substituted characters; locally composited three-line edit; 20 DDIM steps; seed 0
4 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png qwen-zimage-anytext2-mimic AnyText2-SD1.5 b06c583a583818f3679665ef67b51363f107853c 7d5d593928d842500e0da101ab6a1ae28a8d21403441ca62c813c4f92529ace8 Paddle ch 0.222 0.345 0.644 17.9 f5198d8f2365b79880aba8ea6b0c3885142b9de21c9f7f4f6712ac4a0db8dd85 not_run Fidelity gate failed; source glyph masks supplied as font hints; 20 DDIM steps; seed 0
@@ -0,0 +1,5 @@
date,source_commit,fixture,route,seed,global_strength,face_strength,face_count,id_cos,face_lpips,lapvar_ratio,img_lpips,ssim,psnr_db,output_sha256,gemini_oracle
2026-08-13,2d03a00,Gemini_Generated_Image_y48j3cy48j3cy48j.png,visible-metadata-control,0,,,18,,,,,,,9f67264beb1da66a710885464a11f2ae85989094149f6145521d302cd2db940b,synthid_detected
2026-08-13,2d03a00,Gemini_Generated_Image_y48j3cy48j3cy48j.png,zimage-face-only,0,,0.06604512532552083,18,0.927,0.028,0.612,0.013,0.959,34.7,c388e9bbf737bb9b22a7cb2216ffb8026321ca632f9db29e1e456b7c14a8553f,synthid_detected
2026-08-13,2d03a00,Gemini_Generated_Image_y48j3cy48j3cy48j.png,qwen-global-only,0,0.153999996,,18,0.589,0.217,0.331,0.288,0.646,24.2,281a1d28d786d157cbbd13d00a32b215de4dfb3174a3c5e4e83c843e5402138f,clean
2026-08-13,2d03a00,Gemini_Generated_Image_y48j3cy48j3cy48j.png,qwen-zimage-full,0,0.153999996,0.06604512532552083,18,0.852,0.050,0.492,0.252,0.671,24.7,69de191bddbfbf53348c7638bdba661b6d45d257c00a8a2646b2558670c5a4c2,clean
1 date source_commit fixture route seed global_strength face_strength face_count id_cos face_lpips lapvar_ratio img_lpips ssim psnr_db output_sha256 gemini_oracle
2 2026-08-13 2d03a00 Gemini_Generated_Image_y48j3cy48j3cy48j.png visible-metadata-control 0 18 9f67264beb1da66a710885464a11f2ae85989094149f6145521d302cd2db940b synthid_detected
3 2026-08-13 2d03a00 Gemini_Generated_Image_y48j3cy48j3cy48j.png zimage-face-only 0 0.06604512532552083 18 0.927 0.028 0.612 0.013 0.959 34.7 c388e9bbf737bb9b22a7cb2216ffb8026321ca632f9db29e1e456b7c14a8553f synthid_detected
4 2026-08-13 2d03a00 Gemini_Generated_Image_y48j3cy48j3cy48j.png qwen-global-only 0 0.153999996 18 0.589 0.217 0.331 0.288 0.646 24.2 281a1d28d786d157cbbd13d00a32b215de4dfb3174a3c5e4e83c843e5402138f clean
5 2026-08-13 2d03a00 Gemini_Generated_Image_y48j3cy48j3cy48j.png qwen-zimage-full 0 0.153999996 0.06604512532552083 18 0.852 0.050 0.492 0.252 0.671 24.7 69de191bddbfbf53348c7638bdba661b6d45d257c00a8a2646b2558670c5a4c2 clean
@@ -1,4 +1,5 @@
{
"ChatGPT Image May 31, 2026, 02_02_23 PM.png": "This is a longer sample text in English.\nTypography can flow smoothly from large to medium size.\nSmaller lines help demonstrate hierarchy, rhythm, and clarity.\nEven the finest text should remain clean and readable.\n这是一段较长的中文示例文本。\n排版可以从大字号逐渐过渡到中字号。\n更小的文字能够展示层次、节奏与清晰度。\n即使是最小的一行,也应该保持清楚易读。\nЭто более длинный пример текста на русском языке.\nТипографика может плавно переходить от крупного размера к среднему.\nБолее мелкие строки показывают иерархию, ритм и ясность.\nДаже самый маленький текст должен оставаться чистым и читаемым.",
"ChatGPT Image May 31, 2026, 02_03_55 PM.png": "This is a longer sample text in English.\nTypography can flow smoothly from large to medium size.\nSmaller lines help demonstrate hierarchy, rhythm, and clarity.\nEven the finest text should remain clean and readable.\n这是一段较长的中文示例文本。\n排版可以从大字号逐渐过渡到中字号。\n更小的文字能够展示层次、节奏与清晰度。\n即使是最小的一行,也应该保持清楚易读。\nЭто более длинный пример текста на русском языке.\nТипографика может плавно переходить от крупного размера к среднему.\nБолее мелкие строки показывают иерархию, ритм и ясность.\nДаже самый маленький текст должен оставаться чистым и читаемым.",
"Gemini_Generated_Image_633uuy633uuy633u.png": "每天都是一个新的机会。\n用微笑开始它。\n世界也会向你微笑。"
}
@@ -0,0 +1,14 @@
date,source_commit,fixture,variant,ocr_route,text_cer,img_lpips,ssim,psnr_db,mask_fraction,output_sha256,provider_oracle,oracle_control
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage,"Paddle en+ru+ch",0.378,0.107,0.832,24.4,,bed68cf79a023cad582a11b41c4c2962bd90a218037ce931a2b3a8697e3bf4bd,no_openai_signals_detected,source-positive control established earlier
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage-text-restored,"Paddle en+ru+ch",0.048,0.174,0.703,13.9,0.154,3221b2694f39420d509dd37cdaf69581432f9a189e87397346de0c78762402c6,no_openai_signals_detected,qwen-zimage clean in same run
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage-text-selective,"Paddle en+ru+ch",0.101,0.120,0.796,19.5,0.054799,15b40b8cdf9e97558c66d4acdcb97d46769e64f04f143f8061746bc731fbbc21,no_openai_signals_detected,source detected in same browser sequence
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage-text-selective-expanded-mask,"Paddle en+ru+ch",0.081,0.119,0.796,19.5,0.070267,5a8a0bcbe5763d1b1af0d64149cec7e35b16c5751e995d3bacf17581d64177ec,not_checked,fidelity-only follow-up
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage-text-selective-automatic,"Paddle detector en+ru+ch; padded recognizer selection",0.123,0.113,0.808,21.2,0.041279,900def5a433c039aa474789eddc4dc4c63b8126ba1aff6829d2e8d600b9f9c71,no_openai_signals_detected,source detected immediately afterward in the same Chrome sequence
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage,"Paddle en+ru+ch",0.413,0.095,0.818,26.0,,83e777ad28bf80a8e5aa7d2e6ef2348337c667212e2c71075460352ad7423713,no_openai_signals_detected,source-positive control established earlier
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage-text-restored,"Paddle en+ru+ch",0.086,0.162,0.682,14.6,0.170,95812399424feb70fc56d3e518d1ae0a3a50a3c2049d5902b00686ede9d3022d,no_openai_signals_detected,qwen-zimage clean in same run
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage-text-selective,"Paddle en+ru+ch",0.112,0.103,0.780,20.8,0.054912,846f39db0ad4a48a734a5019f0b0ff5d110eecb3210223407e72529dfa4fcf54,no_openai_signals_detected,source detected in same browser sequence
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage-text-selective-expanded-mask,"Paddle en+ru+ch",0.147,0.104,0.779,20.8,0.072556,8ae76e1dedeec26bebb8d35e71fd92d77f142dbb0f3440b330cdf33b79558afe,not_checked,fidelity-only follow-up
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage-text-selective-automatic,"Paddle detector en+ru+ch; padded recognizer selection",0.119,0.104,0.779,20.8,0.057110,ac4201bd2e342b1a37737319438cedb2203f6bb58cd9113f825cf9b80d913755,no_openai_signals_detected,source detected in same browser sequence
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,qwen-zimage-text-selective-automatic-passthrough,"Paddle detector ch; horizontal-only padded recognizer selection",0.074,0.289,0.696,23.3,0.000000,4b05489c7a79f993d35d32528de792baad5a3e33ee950afa147caf5d3e0621a6,not_checked,all three source and candidate recognitions agreed exactly; output is byte-identical to candidate
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage-text-selective-detected-boxes,"Paddle ch detector boxes; 0.10-height recognizer padding",0.127,0.113,0.805,20.9,0.041280,70c52fed8b8a1d98f4cda8558ae08825313e510dfb372e448370d8dcf1da73b0,not_checked,fidelity-only automatic-box follow-up; same 8/20 preserve decisions
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage-text-selective-detected-boxes,"Paddle ch detector boxes; 0.10-height recognizer padding",0.154,0.104,0.778,20.8,0.057146,1fa01ce706042c003a61429d5d97fc6c928ec32278071a91926d588f1e760b00,not_checked,fidelity-only automatic-box follow-up; same 7/20 preserve decisions but worse CER
1 date source_commit fixture variant ocr_route text_cer img_lpips ssim psnr_db mask_fraction output_sha256 provider_oracle oracle_control
2 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage Paddle en+ru+ch 0.378 0.107 0.832 24.4 bed68cf79a023cad582a11b41c4c2962bd90a218037ce931a2b3a8697e3bf4bd no_openai_signals_detected source-positive control established earlier
3 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage-text-restored Paddle en+ru+ch 0.048 0.174 0.703 13.9 0.154 3221b2694f39420d509dd37cdaf69581432f9a189e87397346de0c78762402c6 no_openai_signals_detected qwen-zimage clean in same run
4 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage-text-selective Paddle en+ru+ch 0.101 0.120 0.796 19.5 0.054799 15b40b8cdf9e97558c66d4acdcb97d46769e64f04f143f8061746bc731fbbc21 no_openai_signals_detected source detected in same browser sequence
5 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage-text-selective-expanded-mask Paddle en+ru+ch 0.081 0.119 0.796 19.5 0.070267 5a8a0bcbe5763d1b1af0d64149cec7e35b16c5751e995d3bacf17581d64177ec not_checked fidelity-only follow-up
6 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage-text-selective-automatic Paddle detector en+ru+ch; padded recognizer selection 0.123 0.113 0.808 21.2 0.041279 900def5a433c039aa474789eddc4dc4c63b8126ba1aff6829d2e8d600b9f9c71 no_openai_signals_detected source detected immediately afterward in the same Chrome sequence
7 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage Paddle en+ru+ch 0.413 0.095 0.818 26.0 83e777ad28bf80a8e5aa7d2e6ef2348337c667212e2c71075460352ad7423713 no_openai_signals_detected source-positive control established earlier
8 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage-text-restored Paddle en+ru+ch 0.086 0.162 0.682 14.6 0.170 95812399424feb70fc56d3e518d1ae0a3a50a3c2049d5902b00686ede9d3022d no_openai_signals_detected qwen-zimage clean in same run
9 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage-text-selective Paddle en+ru+ch 0.112 0.103 0.780 20.8 0.054912 846f39db0ad4a48a734a5019f0b0ff5d110eecb3210223407e72529dfa4fcf54 no_openai_signals_detected source detected in same browser sequence
10 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage-text-selective-expanded-mask Paddle en+ru+ch 0.147 0.104 0.779 20.8 0.072556 8ae76e1dedeec26bebb8d35e71fd92d77f142dbb0f3440b330cdf33b79558afe not_checked fidelity-only follow-up
11 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage-text-selective-automatic Paddle detector en+ru+ch; padded recognizer selection 0.119 0.104 0.779 20.8 0.057110 ac4201bd2e342b1a37737319438cedb2203f6bb58cd9113f825cf9b80d913755 no_openai_signals_detected source detected in same browser sequence
12 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png qwen-zimage-text-selective-automatic-passthrough Paddle detector ch; horizontal-only padded recognizer selection 0.074 0.289 0.696 23.3 0.000000 4b05489c7a79f993d35d32528de792baad5a3e33ee950afa147caf5d3e0621a6 not_checked all three source and candidate recognitions agreed exactly; output is byte-identical to candidate
13 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage-text-selective-detected-boxes Paddle ch detector boxes; 0.10-height recognizer padding 0.127 0.113 0.805 20.9 0.041280 70c52fed8b8a1d98f4cda8558ae08825313e510dfb372e448370d8dcf1da73b0 not_checked fidelity-only automatic-box follow-up; same 8/20 preserve decisions
14 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage-text-selective-detected-boxes Paddle ch detector boxes; 0.10-height recognizer padding 0.154 0.104 0.778 20.8 0.057146 1fa01ce706042c003a61429d5d97fc6c928ec32278071a91926d588f1e760b00 not_checked fidelity-only automatic-box follow-up; same 7/20 preserve decisions but worse CER
+51
View File
@@ -0,0 +1,51 @@
{
"ChatGPT Image May 31, 2026, 02_02_23 PM.png": [
{"box": [156, 206, 709, 296], "text": "This is a longer", "script": "alphabetic"},
{"box": [157, 293, 939, 379], "text": "sample text in English.", "script": "alphabetic"},
{"box": [165, 393, 774, 434], "text": "Typography can flow smoothly from", "script": "alphabetic"},
{"box": [161, 434, 521, 477], "text": "large to medium size.", "script": "alphabetic"},
{"box": [165, 495, 566, 522], "text": "Smaller lines help demonstrate hierarchy,", "script": "alphabetic"},
{"box": [165, 524, 354, 550], "text": "rhythm, and clarity.", "script": "alphabetic"},
{"box": [165, 564, 448, 585], "text": "Even the finest text should remain", "script": "alphabetic"},
{"box": [165, 593, 330, 613], "text": "clean and readable.", "script": "alphabetic"},
{"box": [158, 673, 903, 740], "text": "这是一段较长的中文示例文本。", "script": "cjk"},
{"box": [162, 753, 697, 793], "text": "排版可以从大字号逐渐过渡到中字号。", "script": "cjk"},
{"box": [164, 807, 524, 829], "text": "更小的文字能够展示层次、节奏与清晰度。", "script": "cjk"},
{"box": [163, 845, 483, 868], "text": "即使是最小的一行,也应该保持清楚易读。", "script": "cjk"},
{"box": [162, 951, 780, 1009], "text": "Это более длинный пример", "script": "alphabetic"},
{"box": [161, 1019, 723, 1062], "text": "текста на русском языке.", "script": "alphabetic"},
{"box": [165, 1076, 639, 1104], "text": "Типографика может плавно переходить от", "script": "alphabetic"},
{"box": [165, 1111, 505, 1133], "text": "крупного размера к среднему.", "script": "alphabetic"},
{"box": [165, 1151, 513, 1172], "text": "Более мелкие строки показывают иерархию,", "script": "alphabetic"},
{"box": [167, 1178, 287, 1192], "text": "ритм и ясность.", "script": "alphabetic"},
{"box": [165, 1210, 417, 1227], "text": "Даже самый маленький текст должен", "script": "alphabetic"},
{"box": [165, 1235, 381, 1247], "text": "оставаться чистым и читаемым.", "script": "alphabetic"}
],
"ChatGPT Image May 31, 2026, 02_03_55 PM.png": [
{"box": [165, 204, 754, 299], "text": "This is a longer", "script": "alphabetic"},
{"box": [167, 295, 1006, 385], "text": "sample text in English.", "script": "alphabetic"},
{"box": [171, 401, 831, 446], "text": "Typography can flow smoothly from", "script": "alphabetic"},
{"box": [169, 446, 560, 491], "text": "large to medium size.", "script": "alphabetic"},
{"box": [173, 507, 605, 536], "text": "Smaller lines help demonstrate hierarchy,", "script": "alphabetic"},
{"box": [173, 538, 379, 567], "text": "rhythm, and clarity.", "script": "alphabetic"},
{"box": [175, 583, 481, 603], "text": "Even the finest text should remain", "script": "alphabetic"},
{"box": [173, 611, 352, 632], "text": "clean and readable.", "script": "alphabetic"},
{"box": [170, 693, 956, 757], "text": "这是一段较长的中文示例文本。", "script": "cjk"},
{"box": [174, 775, 740, 812], "text": "排版可以从大字号逐渐过渡到中字号。", "script": "cjk"},
{"box": [175, 829, 561, 852], "text": "更小的文字能够展示层次、节奏与清晰度。", "script": "cjk"},
{"box": [174, 869, 534, 892], "text": "即使是最小的一行,也应该保持清楚易读。", "script": "cjk"},
{"box": [169, 958, 824, 1022], "text": "Это более длинный пример", "script": "alphabetic"},
{"box": [171, 1029, 764, 1074], "text": "текста на русском языке.", "script": "alphabetic"},
{"box": [175, 1090, 678, 1119], "text": "Типографика может плавно переходить от", "script": "alphabetic"},
{"box": [175, 1125, 546, 1151], "text": "крупного размера к среднему.", "script": "alphabetic"},
{"box": [175, 1169, 552, 1192], "text": "Более мелкие строки показывают иерархию,", "script": "alphabetic"},
{"box": [175, 1196, 312, 1217], "text": "ритм и ясность.", "script": "alphabetic"},
{"box": [175, 1233, 479, 1253], "text": "Даже самый маленький текст должен", "script": "alphabetic"},
{"box": [177, 1259, 440, 1272], "text": "оставаться чистым и читаемым.", "script": "alphabetic"}
],
"Gemini_Generated_Image_633uuy633uuy633u.png": [
{"box": [1281, 650, 2357, 818], "text": "每天都是一个新的机会。", "script": "cjk"},
{"box": [1466, 788, 2167, 957], "text": "用微笑开始它。", "script": "cjk"},
{"box": [1367, 910, 2264, 1112], "text": "世界也会向你微笑。", "script": "cjk"}
]
}
@@ -0,0 +1,25 @@
date,source_commit,fixture,profile,route,seed,strength,adaptive_polish,text_cer,ocr_floor,added_cer,img_lpips,ssim,psnr_db,output_sha256,oracle_rechecked
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",original,reference,,,,0.127,0.127,0.000,0.000,1.000,inf,,not_applicable
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage,visible-invisible-metadata,0,0.11020972541176471,false,0.378,0.127,0.251,0.107,0.832,24.4,bed68cf79a023cad582a11b41c4c2962bd90a218037ce931a2b3a8697e3bf4bd,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",sdxl-zimage,visible-invisible-metadata,0,0.15,true,0.356,0.127,0.229,0.125,0.707,25.3,9bb932d6e66b52f34ff07c3fcdbdb9c33eb601ae2cd84bec6b8c7b5e30df9926,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",original,reference,,,,0.127,0.127,0.000,0.000,1.000,inf,,not_applicable
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage,visible-invisible-metadata,0,0.11020972541176471,false,0.413,0.127,0.286,0.095,0.818,26.0,83e777ad28bf80a8e5aa7d2e6ef2348337c667212e2c71075460352ad7423713,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",sdxl-zimage,visible-invisible-metadata,0,0.15,true,0.422,0.127,0.295,0.099,0.754,25.9,7605b1175b0c99f481ae6929acc488f34d98bc0f2571d42a1d2972b671d9b8cb,false
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,original,reference,,,,0.037,0.037,0.000,0.000,1.000,inf,,not_applicable
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,qwen-zimage,visible-invisible-metadata,0,0.153999996,false,0.074,0.037,0.037,0.289,0.696,23.3,4b05489c7a79f993d35d32528de792baad5a3e33ee950afa147caf5d3e0621a6,false
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,sdxl-zimage,visible-invisible-metadata,0,0.25,true,0.074,0.037,0.037,0.176,0.769,26.1,d9f89465c31bd2d38ffb111775b9bbea4577f40713fcbebf421e05d344ab27ad,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",zimage-global,visible-invisible-metadata,0,0.08,false,0.068,0.127,-0.059,0.042,0.874,31.4,534de6422a3a03cfde44377aa9bf824a8267db3b73511240f8f07fb8dcb95eff,true
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",zimage-global,visible-invisible-metadata,0,0.10,false,0.086,0.127,-0.041,0.047,0.862,30.8,4dba8211e8d84ebbfff72de00f54d4365d0d202394acd412d1419ae93d1e9f71,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",zimage-global,visible-invisible-metadata,0,0.15,false,0.177,0.127,0.050,0.065,0.838,29.1,fc45c5fff6d33f15d30e05dd02580fe0addd97098b735b2894015f9abdaa43a7,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",zimage-global,visible-invisible-metadata,0,0.20,false,0.248,0.127,0.121,0.093,0.813,27.3,1e55e11f0e289a752098370e9f1d53fed6cd4fab7ceff46fe86dd20c927f9ffb,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",zimage-global,visible-invisible-metadata,0,0.25,false,0.343,0.127,0.216,0.120,0.787,25.7,f5a2d2829e88a790cff5bedc50fca1f878a92cdb392d530afb49d0d95c23b884,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",zimage-global,visible-invisible-metadata,0,0.08,false,0.143,0.127,0.016,0.042,0.877,31.5,5cfa466e276ab37cf532f7458d6dd86ea4f0807e2cbdb6788b7c01ca6ee35a1a,true
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",zimage-global,visible-invisible-metadata,0,0.10,false,0.101,0.127,-0.026,0.047,0.864,30.9,ade869ec7c5b08098495d9b9d67ee900773f189f94bbb411b043f242f5c2fa37,true
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",zimage-global,visible-invisible-metadata,0,0.15,false,0.208,0.127,0.081,0.060,0.839,29.3,a11976963e3f2ca856ce66278e22064d51345d449fa6d672029039a79517d60a,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",zimage-global,visible-invisible-metadata,0,0.20,false,0.341,0.127,0.214,0.070,0.819,27.6,1ab1ae075565e9d39960a80cf5d61e5b0c50d7e23283493ea0277056b3e964ab,false
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",zimage-global,visible-invisible-metadata,0,0.25,false,0.538,0.127,0.411,0.084,0.799,25.6,5ff5a364d84f738de4c56ead7d4950a2f9a1aaf4bec16e97bda6a65915744a57,false
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,zimage-global,visible-invisible-metadata,0,0.08,false,0.074,0.037,0.037,0.068,0.865,30.8,81419eddc3e5150956859dd10c26dd841bb93d9c8b34c01f344e6d69bca6edbf,true
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,zimage-global,visible-invisible-metadata,0,0.10,false,0.074,0.037,0.037,0.086,0.844,29.8,4d1f32bd0dac0e05885ff943e7c10230b06de7d8547e6215f8e2941f732a1df5,true
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,zimage-global,visible-invisible-metadata,0,0.15,false,0.000,0.037,-0.037,0.135,0.793,27.6,526c7987b34896f9513f7439b63b8a87573b59d15930fa307f3bb924435f3cec,true
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,zimage-global,visible-invisible-metadata,0,0.20,false,0.000,0.037,-0.037,0.179,0.746,25.9,325364cf7a3b5d39ac0c44af82181753ca4d6fde0f1b522c896414191725e8fc,true
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,zimage-global,visible-invisible-metadata,0,0.25,false,0.000,0.037,-0.037,0.218,0.703,24.5,f01d64dc0cabd45d66a93c1be39b819b988f29cda1f0a7bc0d5a1954571d807e,true
1 date source_commit fixture profile route seed strength adaptive_polish text_cer ocr_floor added_cer img_lpips ssim psnr_db output_sha256 oracle_rechecked
2 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png original reference 0.127 0.127 0.000 0.000 1.000 inf not_applicable
3 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage visible-invisible-metadata 0 0.11020972541176471 false 0.378 0.127 0.251 0.107 0.832 24.4 bed68cf79a023cad582a11b41c4c2962bd90a218037ce931a2b3a8697e3bf4bd false
4 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png sdxl-zimage visible-invisible-metadata 0 0.15 true 0.356 0.127 0.229 0.125 0.707 25.3 9bb932d6e66b52f34ff07c3fcdbdb9c33eb601ae2cd84bec6b8c7b5e30df9926 false
5 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png original reference 0.127 0.127 0.000 0.000 1.000 inf not_applicable
6 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage visible-invisible-metadata 0 0.11020972541176471 false 0.413 0.127 0.286 0.095 0.818 26.0 83e777ad28bf80a8e5aa7d2e6ef2348337c667212e2c71075460352ad7423713 false
7 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png sdxl-zimage visible-invisible-metadata 0 0.15 true 0.422 0.127 0.295 0.099 0.754 25.9 7605b1175b0c99f481ae6929acc488f34d98bc0f2571d42a1d2972b671d9b8cb false
8 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png original reference 0.037 0.037 0.000 0.000 1.000 inf not_applicable
9 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png qwen-zimage visible-invisible-metadata 0 0.153999996 false 0.074 0.037 0.037 0.289 0.696 23.3 4b05489c7a79f993d35d32528de792baad5a3e33ee950afa147caf5d3e0621a6 false
10 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png sdxl-zimage visible-invisible-metadata 0 0.25 true 0.074 0.037 0.037 0.176 0.769 26.1 d9f89465c31bd2d38ffb111775b9bbea4577f40713fcbebf421e05d344ab27ad false
11 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png zimage-global visible-invisible-metadata 0 0.08 false 0.068 0.127 -0.059 0.042 0.874 31.4 534de6422a3a03cfde44377aa9bf824a8267db3b73511240f8f07fb8dcb95eff true
12 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png zimage-global visible-invisible-metadata 0 0.10 false 0.086 0.127 -0.041 0.047 0.862 30.8 4dba8211e8d84ebbfff72de00f54d4365d0d202394acd412d1419ae93d1e9f71 false
13 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png zimage-global visible-invisible-metadata 0 0.15 false 0.177 0.127 0.050 0.065 0.838 29.1 fc45c5fff6d33f15d30e05dd02580fe0addd97098b735b2894015f9abdaa43a7 false
14 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png zimage-global visible-invisible-metadata 0 0.20 false 0.248 0.127 0.121 0.093 0.813 27.3 1e55e11f0e289a752098370e9f1d53fed6cd4fab7ceff46fe86dd20c927f9ffb false
15 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png zimage-global visible-invisible-metadata 0 0.25 false 0.343 0.127 0.216 0.120 0.787 25.7 f5a2d2829e88a790cff5bedc50fca1f878a92cdb392d530afb49d0d95c23b884 false
16 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png zimage-global visible-invisible-metadata 0 0.08 false 0.143 0.127 0.016 0.042 0.877 31.5 5cfa466e276ab37cf532f7458d6dd86ea4f0807e2cbdb6788b7c01ca6ee35a1a true
17 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png zimage-global visible-invisible-metadata 0 0.10 false 0.101 0.127 -0.026 0.047 0.864 30.9 ade869ec7c5b08098495d9b9d67ee900773f189f94bbb411b043f242f5c2fa37 true
18 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png zimage-global visible-invisible-metadata 0 0.15 false 0.208 0.127 0.081 0.060 0.839 29.3 a11976963e3f2ca856ce66278e22064d51345d449fa6d672029039a79517d60a false
19 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png zimage-global visible-invisible-metadata 0 0.20 false 0.341 0.127 0.214 0.070 0.819 27.6 1ab1ae075565e9d39960a80cf5d61e5b0c50d7e23283493ea0277056b3e964ab false
20 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png zimage-global visible-invisible-metadata 0 0.25 false 0.538 0.127 0.411 0.084 0.799 25.6 5ff5a364d84f738de4c56ead7d4950a2f9a1aaf4bec16e97bda6a65915744a57 false
21 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png zimage-global visible-invisible-metadata 0 0.08 false 0.074 0.037 0.037 0.068 0.865 30.8 81419eddc3e5150956859dd10c26dd841bb93d9c8b34c01f344e6d69bca6edbf true
22 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png zimage-global visible-invisible-metadata 0 0.10 false 0.074 0.037 0.037 0.086 0.844 29.8 4d1f32bd0dac0e05885ff943e7c10230b06de7d8547e6215f8e2941f732a1df5 true
23 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png zimage-global visible-invisible-metadata 0 0.15 false 0.000 0.037 -0.037 0.135 0.793 27.6 526c7987b34896f9513f7439b63b8a87573b59d15930fa307f3bb924435f3cec true
24 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png zimage-global visible-invisible-metadata 0 0.20 false 0.000 0.037 -0.037 0.179 0.746 25.9 325364cf7a3b5d39ac0c44af82181753ca4d6fde0f1b522c896414191725e8fc true
25 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png zimage-global visible-invisible-metadata 0 0.25 false 0.000 0.037 -0.037 0.218 0.703 24.5 f01d64dc0cabd45d66a93c1be39b819b988f29cda1f0a7bc0d5a1954571d807e true
@@ -0,0 +1,7 @@
date,source_commit,fixture,variant,ocr_route,text_cer,img_lpips,ssim,psnr_db,mask_fraction,output_sha256,provider_oracle,oracle_control
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage,"Vision en+ru; Paddle ch",0.338,0.107,0.832,24.4,,bed68cf79a023cad582a11b41c4c2962bd90a218037ce931a2b3a8697e3bf4bd,no_openai_signals_detected,source-positive control established earlier
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_02_23 PM.png",qwen-zimage-text-restored,"Vision en+ru; Paddle ch",0.007,0.174,0.703,13.9,0.154,3221b2694f39420d509dd37cdaf69581432f9a189e87397346de0c78762402c6,no_openai_signals_detected,qwen-zimage clean in same run
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage,"Vision en+ru; Paddle ch",0.305,0.095,0.818,26.0,,83e777ad28bf80a8e5aa7d2e6ef2348337c667212e2c71075460352ad7423713,no_openai_signals_detected,source-positive control established earlier
2026-08-13,2d03a00,"ChatGPT Image May 31, 2026, 02_03_55 PM.png",qwen-zimage-text-restored,"Vision en+ru; Paddle ch",0.007,0.162,0.682,14.6,0.170,95812399424feb70fc56d3e518d1ae0a3a50a3c2049d5902b00686ede9d3022d,no_openai_signals_detected,qwen-zimage clean in same run
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,qwen-zimage,"Paddle ch",0.074,0.289,0.696,23.3,,4b05489c7a79f993d35d32528de792baad5a3e33ee950afa147caf5d3e0621a6,synthid_detected,inconclusive account: clean baseline not reproduced
2026-08-13,2d03a00,Gemini_Generated_Image_633uuy633uuy633u.png,qwen-zimage-text-restored,"Paddle ch",0.111,0.328,0.653,18.1,0.051,2ab7666a4161ff0e12c53d9412b23d013970c7613a8a55a1b1ae8a2744e3290d,synthid_detected,inconclusive account: qwen-zimage control also detected
1 date source_commit fixture variant ocr_route text_cer img_lpips ssim psnr_db mask_fraction output_sha256 provider_oracle oracle_control
2 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage Vision en+ru; Paddle ch 0.338 0.107 0.832 24.4 bed68cf79a023cad582a11b41c4c2962bd90a218037ce931a2b3a8697e3bf4bd no_openai_signals_detected source-positive control established earlier
3 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_02_23 PM.png qwen-zimage-text-restored Vision en+ru; Paddle ch 0.007 0.174 0.703 13.9 0.154 3221b2694f39420d509dd37cdaf69581432f9a189e87397346de0c78762402c6 no_openai_signals_detected qwen-zimage clean in same run
4 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage Vision en+ru; Paddle ch 0.305 0.095 0.818 26.0 83e777ad28bf80a8e5aa7d2e6ef2348337c667212e2c71075460352ad7423713 no_openai_signals_detected source-positive control established earlier
5 2026-08-13 2d03a00 ChatGPT Image May 31, 2026, 02_03_55 PM.png qwen-zimage-text-restored Vision en+ru; Paddle ch 0.007 0.162 0.682 14.6 0.170 95812399424feb70fc56d3e518d1ae0a3a50a3c2049d5902b00686ede9d3022d no_openai_signals_detected qwen-zimage clean in same run
6 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png qwen-zimage Paddle ch 0.074 0.289 0.696 23.3 4b05489c7a79f993d35d32528de792baad5a3e33ee950afa147caf5d3e0621a6 synthid_detected inconclusive account: clean baseline not reproduced
7 2026-08-13 2d03a00 Gemini_Generated_Image_633uuy633uuy633u.png qwen-zimage-text-restored Paddle ch 0.111 0.328 0.653 18.1 0.051 2ab7666a4161ff0e12c53d9412b23d013970c7613a8a55a1b1ae8a2744e3290d synthid_detected inconclusive account: qwen-zimage control also detected
@@ -0,0 +1,6 @@
date,upstream_commit,workflow_sha256,fixture,route,seed,global_strength,face_strength,face_count,id_cos,face_lpips,lapvar_ratio,img_lpips,ssim,psnr_db,output_sha256,gemini_oracle,implementation
2026-08-13,3007d0351596ae0a78b7074dae7ad179710b1e48,41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6,comparison/before/12.png,official-before,,,,10,,,,,,,056276464387c883f5fd89bd2f1415faeb425abb3153cd10efb25d1b9053bdc2,synthid_detected,upstream-published-bytes
2026-08-13,3007d0351596ae0a78b7074dae7ad179710b1e48,41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6,comparison/after/12.png,official-after,,,,10,0.975,0.093,0.354,0.172,0.820,23.8,a822d805666224adb3e360b3ee89e14f20a341f1e7c4d3a5d3d5251e9286aefe,clean,upstream-published-bytes
2026-08-13,3007d0351596ae0a78b7074dae7ad179710b1e48,41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6,Gemini_Generated_Image_y48j3cy48j3cy48j.png,visible-metadata-control,0,,,18,,,,,,,9f67264beb1da66a710885464a11f2ae85989094149f6145521d302cd2db940b,synthid_detected,project-runtime-control
2026-08-13,3007d0351596ae0a78b7074dae7ad179710b1e48,41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6,Gemini_Generated_Image_y48j3cy48j3cy48j.png,upstream-global-close-reproduction,0,0.153999996,,18,0.589,0.217,0.331,0.286,0.647,24.3,281a1d28d786d157cbbd13d00a32b215de4dfb3174a3c5e4e83c843e5402138f,clean,diffsynth-bf16-yunet-sam
2026-08-13,3007d0351596ae0a78b7074dae7ad179710b1e48,41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6,Gemini_Generated_Image_y48j3cy48j3cy48j.png,upstream-full-close-reproduction,0,0.153999996,0.13209025065104166,18,0.783,0.083,0.470,0.260,0.659,24.7,2d7a8f76f98b905f14d69e5f9e11979610c72c6be7b6ac298dec4659f17bd819,clean,diffsynth-bf16-yunet-sam
1 date upstream_commit workflow_sha256 fixture route seed global_strength face_strength face_count id_cos face_lpips lapvar_ratio img_lpips ssim psnr_db output_sha256 gemini_oracle implementation
2 2026-08-13 3007d0351596ae0a78b7074dae7ad179710b1e48 41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6 comparison/before/12.png official-before 10 056276464387c883f5fd89bd2f1415faeb425abb3153cd10efb25d1b9053bdc2 synthid_detected upstream-published-bytes
3 2026-08-13 3007d0351596ae0a78b7074dae7ad179710b1e48 41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6 comparison/after/12.png official-after 10 0.975 0.093 0.354 0.172 0.820 23.8 a822d805666224adb3e360b3ee89e14f20a341f1e7c4d3a5d3d5251e9286aefe clean upstream-published-bytes
4 2026-08-13 3007d0351596ae0a78b7074dae7ad179710b1e48 41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6 Gemini_Generated_Image_y48j3cy48j3cy48j.png visible-metadata-control 0 18 9f67264beb1da66a710885464a11f2ae85989094149f6145521d302cd2db940b synthid_detected project-runtime-control
5 2026-08-13 3007d0351596ae0a78b7074dae7ad179710b1e48 41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6 Gemini_Generated_Image_y48j3cy48j3cy48j.png upstream-global-close-reproduction 0 0.153999996 18 0.589 0.217 0.331 0.286 0.647 24.3 281a1d28d786d157cbbd13d00a32b215de4dfb3174a3c5e4e83c843e5402138f clean diffsynth-bf16-yunet-sam
6 2026-08-13 3007d0351596ae0a78b7074dae7ad179710b1e48 41911b3b8e01bf51450361dc8beccd84c9513f78ddb160430cdd9bebc294adf6 Gemini_Generated_Image_y48j3cy48j3cy48j.png upstream-full-close-reproduction 0 0.153999996 0.13209025065104166 18 0.783 0.083 0.470 0.260 0.659 24.7 2d7a8f76f98b905f14d69e5f9e11979610c72c6be7b6ac298dec4659f17bd819 clean diffsynth-bf16-yunet-sam
+31
View File
@@ -398,6 +398,37 @@ schedule, CFG 1.0 and CUDA, so every one of those flags existed only to be refus
several layers down. They are not parsed at all now, which fails at the point the
user can act on rather than after a model load.
### Restore operator-verified text
`--text-manifest` enables the experimental `vae-glyphs` post-pass. It reconstructs
the source with the Qwen VAE, blends 15% of that reconstruction into the normal
`qwen-zimage` result, erases the annotated candidate glyphs with LaMa, and composites
only the reconstructed glyph cores through source-derived silhouettes. It does not
run OCR or choose which strings are correct.
Install the combined extra and run only with a manually reviewed manifest:
```bash
uv tool install --force "remove-ai-watermarks[text-restoration]"
remove-ai-watermarks invisible image.png -o clean.png \
--pipeline qwen-zimage --text-manifest verified-lines.json --force
```
The manifest is a JSON object with `schema_version: 1`, `verified: true`, decoded
RGB dimensions, `source_pixel_sha256`, and a non-empty `lines` array. Each line has
an integer `[x1, y1, x2, y2]` box, exact `text`, a non-empty `script`, and an optional
angle from -30 to 30 degrees. Lines must be in top-to-bottom, left-to-right order.
The hash binds the annotations to decoded RGB geometry and pixels, so metadata-only
container changes remain valid while a resized or edited source fails closed. The
experimental helper
`remove_ai_watermarks._internal.text_restoration.source_pixel_sha256` computes it.
This mode is supported only by `qwen-zimage` at native untiled geometry with
`humanize=0`, `unsharp=0`, and adaptive polish disabled. `all` also accepts the flag,
but its manifest must match the pixels entering the invisible stage; if visible-mark
removal changes those pixels, the hash check rejects the run. One oracle verdict does
not certify another manifest, seed, model/runtime version, or output hash.
### Work with limited memory
Lower CUDA memory pressure:
+12
View File
@@ -87,6 +87,15 @@ removal, metadata stripping and every `identify` command still run anywhere.
Video SynthID regeneration is a separate VAE path and does still run on CPU or MPS;
it needs the `diffusion` extra, not this one.
The experimental verified-text post-pass additionally needs LaMa:
```bash
uv tool install --force "remove-ai-watermarks[text-restoration]"
```
That extra includes `qwen-zimage` and `lama`; it does not add OCR. Text strings and
line boxes must be reviewed before the run.
## Feature extras
Extras are composable. Install only the capabilities and file formats the
@@ -104,6 +113,7 @@ application actually uses:
| `migan` | MI-GAN ONNX fill backend | `visible`, ONNX Runtime | Model download, no Torch |
| `lama` | big-LaMa ONNX fill backend | `visible`, ONNX Runtime | Model download, no Torch |
| `qwen-zimage` | Invisible image-watermark removal, both CUDA-only profiles | `diffusion`, DiffSynth | Yes |
| `text-restoration` | Opt-in verified Qwen-VAE glyph restoration | `qwen-zimage`, `lama` | Yes |
| `all` | Every production feature available on the active Python | All compatible rows above | Yes |
| `dev` | Tests, linting, typing, and upstream parity checks | `video`, `detect`, upstream invisible-watermark | Yes, for parity tests |
@@ -118,6 +128,8 @@ flowchart LR
migan --> visible
lama --> visible
qwen["qwen-zimage"] --> diffusion
text["text-restoration"] --> qwen
text --> lama
heif
trustmark
```
+28 -2
View File
@@ -61,8 +61,25 @@ change even when the watermark is successfully disrupted.
Both are CUDA only and differ only in the global regeneration model: each
conditions that stage on a canny edge map, which preserves structure but not
identity or exact texture, and each then runs the same face stage.
`qwen-zimage` is the higher fidelity of the two. Both are large, slow, and may
still alter small text or difficult faces.
Existing face evaluations favor `qwen-zimage`, but there is no blanket fidelity
ordering across content types. A fixed-seed, three-scene text comparison at the
profile defaults found no stable winner: SDXL won one poster, Qwen won one, and
the Chinese sign tied. Both are large, slow, and may still alter small text or
difficult faces. The measurements and their OCR and oracle caveats are tracked
in [`data/evaluations/fidelity/`](../data/evaluations/fidelity/README.md).
A global Z-Image Turbo prototype preserved text substantially better at low
strength, but it has no useful cross-provider operating point and is not a
supported profile. Automatic text restorers also remain research-only:
fresh-font and silhouette variants visibly changed typography. The higher-fidelity
`vae-glyphs` route is available only as an experimental opt-in with verified strings
and line geometry. It builds its donor internally but still requires an independently
clean global anchor. Automatic OCR and line-box proposals are not reliable enough to
remove those requirements, and exact oracle results do not establish a general mask,
seed, runtime, or provider operating range. Qwen-Image-2.0 is hosted-only and exposes
no equivalent low-strength denoise control. Exact experiments, controls, and pass
rates are kept in
[`text-protection-research.md`](text-protection-research.md) and the
[`fidelity` evaluation record](../data/evaluations/fidelity/README.md).
### Removal cannot be verified locally for proprietary SynthID
@@ -175,6 +192,15 @@ certified at a fixed seed. The live resolver is
| `qwen-zimage` | CUDA only, large model stack, and limited broad certification across seeds and content. |
| `sdxl-zimage` | CUDA only. Its strength ladder is flat per vendor, not a resolution curve, because flat values are what was measured. |
Only manually verified `vae-glyphs` is an optional production stage, and it is
experimental rather than a default.
OCR plus LaMa recovered literal poster text but changed fonts and worsened whole-image
fidelity. Restricting it to OCR-mismatched lines improved the tradeoff but still
left a local shadow on one poster. The published AnyText2 SD1.5 checkpoint
substituted Chinese characters and increased CER on the sign fixture. AnyText2XL
is not published, and the released wrapper truncates individual text lines after
20 characters.
The `controlnet`, `sdxl`, `qwen` and `default` profiles were removed, not aliased
onward: a retired name is rejected at parse time rather than routed into a profile
the caller never chose. There is no `--model`, `--steps`, `--guidance-scale`,
+51
View File
@@ -949,6 +949,57 @@ orchestration, YuNet integration, SAM selection, masks, sizing helpers, and pixe
compositing are implemented for this runtime. Changing a calibrated model input
requires the same provider-oracle and identity evaluation as a model change.
#### Verified text restoration
[`_internal/text_restoration.py`](../src/remove_ai_watermarks/_internal/text_restoration.py)
implements the opt-in `vae-glyphs` stage. A versioned manifest carries manually
reviewed strings and source-space line boxes, plus a SHA-256 over decoded RGB width,
height, and pixels. Validation happens before model loading. The product never treats
OCR confidence as verification.
When enabled, `QwenZImagePipeline` reconstructs the source once through its already
loaded Qwen VAE, runs the ordinary global and face stages, blends 15% of the VAE
reconstruction into that clean result, and calls the shared restoration compositor.
The compositor derives binary source and candidate silhouettes, groups nearby lines,
uses LaMa for the initial and residual-glyph erase passes, paints fresh silhouette
edges, then copies the Qwen-VAE core with a 0.5-pixel feather. The evaluation script
imports these same mask and compositing helpers so the two implementations cannot
silently drift.
The stage is deliberately narrower than the engine: it rejects `sdxl-zimage`, tiles,
resolution caps, humanize, unsharp, and adaptive polish. Those combinations change
geometry or final pixels after the verified layer and have no measured oracle result.
It remains opt-in because annotations are manual and provider verdicts apply only to
the exact tested output hashes, not to the mechanism in general.
A matched stage-isolation check on the 18-face Gemini portrait grid confirms the
division of responsibility. The visible-cleaned, metadata-stripped control and the
Z-Image face-only output were both SynthID-positive; Qwen global-only and the full
Qwen-then-Z-Image output were both clean. The face stage raised identity cosine from
0.589 to 0.852 and reduced face LPIPS from 0.217 to 0.050 without reintroducing a
detectable whole-image signal. Thus Z-Image is a masked fidelity repair stage here,
not the watermark-removal stage. Exact hashes, metrics, strengths, and the one-fixture,
one-seed caveat are recorded in
[`data/evaluations/fidelity/face-stage-isolation-2026-08-13.csv`](../data/evaluations/fidelity/face-stage-isolation-2026-08-13.csv).
The public Synthid-Bypass v2 graph was subsequently audited at upstream commit
`3007d035`. Its saved-output path confirms the same division: Qwen-Image-2512
Lightning plus Canny is global, and Z-Image Turbo exists only inside the masked
face detailer. The connected face path is YOLOv8-face plus SAM; the MediaPipe nodes
described by the upstream README and the 1.2-megapixel normalization node do not
reach `SaveImage`. Upstream also applies its adaptive face strength directly,
whereas this implementation multiplies it by `FACE_DENOISE_SCALE = 0.5`.
A close reproduction on the same portrait fixture kept the control positive and
made both Qwen global-only and full Qwen-then-Z-Image outputs clean. Applying the
upstream-strength face pass raised identity from 0.589 to 0.783 and reduced face
LPIPS from 0.217 to 0.083, but remained worse than this profile's 0.852 and 0.050.
The published upstream pair 12 was also independently checked positive before and
clean after, with 0.975 identity. Exact workflow provenance, hashes, metrics,
oracle outcomes, and the DiffSynth/GGUF, scheduler, detector, and seed caveats are
recorded in
[`data/evaluations/fidelity/upstream-v2-reproduction-2026-08-13.csv`](../data/evaluations/fidelity/upstream-v2-reproduction-2026-08-13.csv).
### SDXL plus Z-Image
[`_internal/sdxl_zimage_pipeline.py`](../src/remove_ai_watermarks/_internal/sdxl_zimage_pipeline.py)
+17
View File
@@ -596,6 +596,23 @@ engine = InvisibleEngine(pipeline="sdxl-zimage")
The `qwen-zimage` extra is required for both profiles: each runs the same
DiffSynth Z-Image face stage.
The opt-in verified-text stage uses the same `text_manifest` argument as the CLI:
```python
engine.remove_watermark(
Path("watermarked.png"),
Path("clean.png"),
text_manifest=Path("verified-lines.json"),
)
```
Install `remove-ai-watermarks[text-restoration]`. The manifest schema and safety
constraints are documented in the CLI guide. The engine verifies its decoded RGB
hash before loading the diffusion models and rejects SDXL, tiling, downscaling, and
postprocessing combinations that were not evaluated. `InvisibleOptions` exposes the
same field for `remove_all`; after a visible-stage edit, the manifest must be built
against the staged pixels rather than the pristine source.
`remove_watermark` takes strength, seed, tiling, resolution, and postprocessing
controls. It takes no model id, step count or guidance scale, and neither does the
constructor: each profile pins its model stack, its per-stage schedule and CFG
+200
View File
@@ -143,3 +143,203 @@ Deep-research workflow run `wf_118b9a03-3eb` (2026-05-29). Findings adversariall
(2/3 refutes required to kill a claim). This note records research only; no code change is
implied until a prototype validates fidelity and the SynthID-scrub guarantee on the
restored output.
## 2026-08-13 OCR plus LaMa prototype
A lightweight version of the recommended two-stage architecture was tested on
the three tracked text fixtures. It used the already-clean Qwen result as the
base, recognized the source text, removed source and Qwen glyph contours with
LaMa, and rendered the recognized strings as fresh pixels. No source pixels
were pasted back.
The result validates only part of the hypothesis. Character-weighted CER on the
two multilingual posters fell from 0.338/0.305 to 0.007/0.007, and OpenAI Verify
remained clean for both matched Qwen and restored pairs. However, replacement
fonts changed the design, whole-image LPIPS worsened by 0.067 on each poster,
and the light poster retained a shadow in one small English line. On the Chinese
sign, punctuation errors moved CER from 0.074 to 0.111. The Google verifier run
was inconclusive because the available account detected both the restored image
and the byte-identical Qwen control that a previous account had called clean.
The experiment supports a proper glyph-conditioned restorer, not shipping the
plain OCR/system-font compositor. Exact hashes, metrics, masks, and provider
verdicts are tracked in
[`data/evaluations/fidelity/text-restoration-2026-08-13.csv`](../data/evaluations/fidelity/text-restoration-2026-08-13.csv).
## 2026-08-13 AnyText2 follow-up
The official AnyText2 SD1.5 checkpoint was then tested as the glyph-conditioned
restorer. Its ModelScope entry and code are Apache 2.0, and the downloaded
checkpoint matched the published SHA-256. An official blackboard edit completed
successfully before the fixture run, establishing that the runtime reproduced
the model rather than silently exercising a fallback.
On the Chinese sign, a three-line local edit over the exact Qwen output scored
CER 0.185 under the standard detector. Font mimic from the source glyph masks
scored 0.222. Whole-image LPIPS changed from 0.289 to 0.338 and 0.345 respectively.
A padded crop-recognition check showed why the comparison matters: Qwen contained
the full correct text and scored 0.000, while default AnyText2 still scored 0.074
with two substituted characters and mimic remained at 0.222. The provider oracle
was deliberately not called because these variants had already failed the
content-fidelity gate.
AnyText2 is therefore not the missing production restorer in its published form.
The negative result is narrower than the model family: only the SD1.5 checkpoint
is public, while AnyText2XL remains listed as future work. The current wrapper
also truncates each quoted text line to 20 characters, which prevents a direct
test of several tracked English and Russian poster lines. Exact model provenance,
output hashes, and metrics are recorded in
[`data/evaluations/fidelity/anytext2-restoration-2026-08-13.csv`](../data/evaluations/fidelity/anytext2-restoration-2026-08-13.csv).
The same cross-check exposes a measurement bug in the earlier lightweight
restoration result. Paddle detection made tight boxes that omitted the final
Chinese full stop from two lines, producing Qwen CER 0.074 although the pixels
and padded recognition were correct. Adding 16-32 pixels of horizontal context
restored every punctuation mark. A deterministic rerender using the full OCR
strings and the closest of nine mask-scored CJK fonts also reached detector CER
0.000, but its heavier Hiragino Sans GB W6 glyphs raised LPIPS from 0.289 to 0.327.
The actionable design is selective restoration: compare padded source and output
recognition first, and preserve the Qwen output when they already match.
That policy was tested on the two multilingual posters. A manual prototype kept
the lines whose padded source and Qwen recognition agreed and rerendered only the
changed lines. The edited mask fell from 15.4%/17.0% in the full compositor to
5.5%/5.5%. On a single Paddle `en+ru+ch` route, CER changed from 0.378/0.413 for
Qwen to 0.101/0.112, while LPIPS was only 0.120/0.103 rather than the full
compositor's 0.174/0.162. The different OCR route is intentional and all three
variants were remeasured through it; these CER values are not directly
comparable with the earlier Vision/Paddle hybrid results.
Both selective outputs returned `No OpenAI signals detected` in OpenAI Verify,
and the original light poster returned `Generated with OpenAI tools` immediately
afterward as the positive control. The dark poster was visually clean. The light
poster still showed one local old-text shadow, so mask construction remains the
blocking defect. Exact hashes, metrics, mask fractions, and oracle controls are
in
[`data/evaluations/fidelity/selective-text-restoration-2026-08-13.csv`](../data/evaluations/fidelity/selective-text-restoration-2026-08-13.csv).
Uniformly expanding every selected glyph mask by two pixels removed that light
poster shadow and moved CER from 0.101 to 0.081 without a meaningful LPIPS cost.
It also expanded the edited area from 5.5% to 7.0%, and the identical rule made
the dark poster worse at CER 0.147 rather than 0.112. Those follow-up bytes have
not been oracle-checked. The next prototype should detect residual source glyphs
per line and expand only the failing component, rather than introducing another
poster-wide tuning constant.
The manual selection was then replaced with per-line padded recognition. A line
is left untouched only when the source recognizer is sufficiently consistent
with the verified line and normalized source/candidate recognition agree
exactly. The automatic rule reproduced the manual selection on the dark poster
and correctly kept one additional large Russian line on the light poster. It
reduced that poster's effective edited area from 5.5% to 4.1%.
Mask cleanup became a measured second pass rather than a global dilation. After
the first LaMa erase, the prototype finds contrast that remains specifically at
the original glyph positions, dilates only that residual, and erases it before
fresh text is drawn. This removed the visible double glyphs from both posters.
The automatic outputs scored LPIPS 0.113/0.104 and detector CER 0.123/0.119, with
effective edited fractions of 4.1%/5.7%. The light aggregate CER is pessimistic:
the page detector fragments its untouched Cyrillic line into Latin-like pieces,
while the padded Russian recognizer returns the exact expected text.
OpenAI Verify returned `No OpenAI signals detected` for both new hashes and then
`Generated with OpenAI tools` for the original light poster in the same browser
sequence. Reproducing the prototype from the tracked script exposed harmless
ONNX rounding on the light image: 386 pixels differed by at most one channel
value from the initially checked hash. The exact tracked output hash
`900def5a...` was therefore submitted separately, returned clean, and was
followed immediately by another positive source result. The dark tracked output
was byte-identical to the initially checked hash.
The Chinese sign provides a no-edit control for the selector. Detector boxes for
its three large lines overlap vertically, so Latin-style vertical padding
incorrectly mixed neighboring lines. Horizontal-only CJK padding captured the
terminal punctuation without mixing neighbors and made source and candidate
recognition agree exactly on all three lines.
The script then produced a zero mask and copied the Qwen candidate byte for byte.
This prevents the heavier-font regression seen in the earlier CJK compositor,
but does not resolve the unavailable Google-negative oracle control.
The result clears the measured OpenAI gate without a manual line selection, but
is not a production design yet. It still assumes verified source strings and
line boxes, relies on macOS system fonts, and needs evaluation on uncurated
layouts, rotated text, false OCR matches, automatic line-box discovery, and a
reproducible Google-negative control.
The later source-silhouette experiment removed the font lookup but did not meet
the actual visual requirement. Native-size review found changed stroke weight,
color variation, edge antialiasing, and decorative details even where crop OCR
was exact. A narrower `vae-glyphs` experiment uses the silhouette only as a
support mask: exact pixels come from a Qwen-VAE reconstruction, while a fresh
flat-color silhouette supplies only the outer edge beneath that core. Across 48
typography cases (548 annotated lines), this improved glyph-edge MAE in 48/48
and text-box SSIM in 47/48. The median values moved from 37.59 to 32.42 and from
0.854 to 0.914 respectively. One separate untracked core-only output with a
1.43% donor area returned `No OpenAI signals detected`; the tracked script then
reproduced those exact bytes. A 0.5-pixel feathered sibling with 2.75% nonzero
donor area improved mean text-box SSIM from 0.957 to 0.960, with a 0.918 minimum
across 15 verified lines, while full-image RGB SSIM reached 0.864. Crop OCR
recovered from 7/15 exact lines on the raw pass to 14/15, equal to the source's
own recognizer floor. Whole-image LPIPS was 0.082, but only 0.108% of pixels
remained exactly equal to the source and the detected face's Laplacian-variance
ratio was 0.670, confirming residual global smoothing. A same-session
OpenAI Verify sequence detected the exact source control in 1/1 check, returned
clean for the raw global pass and clean-fidelity anchor in 1/1 each, and returned
clean for the feathered output in 3/3 repeats. The tracked script reproduced
that feathered file byte for byte. This proves two materially better exact-output
Pareto points, not a general operating range. The 48 matrix outputs and other
mask sizes remain oracle-unverified.
The first Google oracle sample did not transfer. For the synthetic CJK sign,
two separate work accounts detected SynthID in both the resaved positive source
control and the exact Qwen-VAE donor candidate. The candidate's text-box SSIM
improved from 0.791 to 0.835 and its glyph-edge MAE from 35.48 to 22.49, but its
3.67% donor layer did not create a clean operating point. The Qwen silhouette
base was already detected, as was the earlier original-sign Qwen baseline in the
same account. The measured cause is therefore insufficient global Google
regeneration, not a demonstrated watermark regression from the text donor.
Google needs a stronger oracle-clean global anchor before the narrow donor can
be evaluated there.
A targeted follow-up supplied that missing anchor for one CJK case. An explicit
Qwen global pass at strength 0.30 returned no reliable SynthID signal in Gemini,
while the matched source control was detected. Applying the 0.5-pixel
`vae-glyphs` layer to that clean anchor changed 1.17% of the image and raised
mean text-box SSIM from 0.765 to 0.964 (minimum 0.963), with mean text-box MAE
falling from 20.74 to 5.38. The post-donor output then returned clean in 3/3
Gemini checks across two work accounts. A separate 18-face fixture with four
tiny UI-text lines also had a clean strength-0.30 anchor; its donor layer changed
0.70% of the image, raised mean text-box SSIM from 0.617 to 0.976 (minimum
0.972), and returned clean in 3/3 checks across the same two accounts. One
quota-exceeded response is excluded from both pass rates. These are two exact
oracle-certified outputs, not a general operating range: other layouts, masks,
strengths, seeds, and donor areas remain unverified.
An automatic-box follow-up merged Paddle word detections by vertical overlap.
It found exactly 20/20 poster lines and 3/3 sign lines, with mean IoU 0.857,
0.847, and 1.000 against the verified boxes. That structural match was not
sufficient: reusing the annotation crop padding changed recognition decisions,
reduced preserved dark-poster lines from 7 to 4, and expanded its edited
fraction from 5.7% to 11.2%.
A recognition-only sweep isolated the crop sensitivity. Limiting vertical
padding around detector boxes to 8-12% of line height reproduced the verified-
box decision vectors on both posters; 10% was used for a full follow-up. It kept
8/20 and 7/20 lines with edited fractions of 4.1% and 5.7%, and visual inspection
found no double glyphs. Whole-image LPIPS remained 0.113/0.104, but detector CER
was 0.127/0.154 instead of 0.123/0.119. The dark regression failed the fidelity
gate, so these hashes were not submitted to the provider oracle. The opt-in flag
remains only to reproduce the negative result. Count, IoU, and matching decision
vectors are therefore insufficient gates for automatic boxes; the next design
needs box rectification or recognition stability under crop jitter before it
can replace verified geometry.
Removing verified strings was tested separately with an annotation-seed dry
run. It detects boxes, chooses `en`, `ru`, or `ch` from Unicode script, and
accepts a draft only when three crop paddings normalize identically and every
confidence is at least 0.85. One execution proposed 20 and 18 poster lines, but
exact-text precision against the held-out annotations was only 90.0% and 94.4%.
The stable errors were punctuation: one lost English comma and an ideographic
comma consistently replaced by ASCII. Thus crop agreement and confidence do
not verify literal text. The helper is retained only to seed annotations; every
proposal needs manual review before it can enter the restoration compositor.
+6 -1
View File
@@ -120,6 +120,11 @@ qwen-zimage = [
"diffsynth>=2.0.17,<3",
"torchvision>=0.20.0",
]
# Opt-in verified-text reconstruction over qwen-zimage. LaMa removes the changed
# candidate glyphs before exact Qwen-VAE cores are composited back.
text-restoration = [
"remove-ai-watermarks[qwen-zimage,lama]",
]
# Adobe TrustMark decoder -- the open, keyless watermark behind Adobe Durable
# Content Credentials (soft-binding alg ``com.adobe.trustmark.P``). Optional
# because it pulls torch and downloads model weights on first use. identify()
@@ -161,7 +166,7 @@ dev = [
]
# ``qwen-zimage`` already pulls ``diffusion``; naming both would suggest diffusion is
# independently sufficient for a removal, which it is not.
all = ["remove-ai-watermarks[video,heif,detect,trustmark,qwen-zimage,lama,migan]"]
all = ["remove-ai-watermarks[video,heif,detect,trustmark,text-restoration,migan]"]
[project.scripts]
remove-ai-watermarks = "remove_ai_watermarks.cli:main"
+34
View File
@@ -0,0 +1,34 @@
"""Pure text-normalization helpers shared by evaluation scripts."""
from __future__ import annotations
import unicodedata
def normalize_text(text: str) -> str:
"""Normalize text for layout-independent evaluation comparisons."""
return "".join(unicodedata.normalize("NFC", text).casefold().split())
def levenshtein_normalized(left: str, right: str) -> float:
"""Return normalized Levenshtein distance without changing either input."""
if not left and not right:
return 0.0
previous = list(range(len(right) + 1))
for left_index, left_character in enumerate(left, start=1):
current = [left_index]
for right_index, right_character in enumerate(right, start=1):
current.append(
min(
current[-1] + 1,
previous[right_index] + 1,
previous[right_index - 1] + (left_character != right_character),
)
)
previous = current
return previous[-1] / max(len(left), len(right))
def normalized_edit_distance(left: str, right: str) -> float:
"""Normalize text, then return Levenshtein distance over the result."""
return levenshtein_normalized(normalize_text(left), normalize_text(right))
+8 -7
View File
@@ -3,10 +3,9 @@
# dependencies = [
# "click",
# "numpy",
# "opencv-python-headless",
# "opencv-python-headless<5",
# "pillow",
# "scikit-image",
# "rapidfuzz",
# "torch",
# "lpips",
# "paddleocr",
@@ -29,7 +28,7 @@ metrics run only where faces are detected, text metrics only where text is.
Two subcommands:
ocr -- OCR images (PaddleOCR PP-OCRv6) into a JSON {basename: text} file.
ocr -- OCR images (PaddleOCR defaults) into a JSON {basename: text} file.
Run this on the ORIGINALS, hand-verify/correct the file, and it
becomes the ground truth for ``compare --ground-truth`` -- the clean
way to score text, since OCR-vs-OCR is doubly noisy (errors on both
@@ -54,6 +53,7 @@ Usage:
from __future__ import annotations
import json
import sys
import unicodedata
from dataclasses import dataclass, field
from pathlib import Path
@@ -65,7 +65,10 @@ import numpy as np
from _plain_console import Console, Table
console = Console()
ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(0, str(ROOT))
from scripts._text_eval import levenshtein_normalized # noqa: E402
# ── helpers ──────────────────────────────────────────────────────────
@@ -89,7 +92,7 @@ def _norm(text: str) -> str:
return "".join(unicodedata.normalize("NFC", text).split())
# ── text: PaddleOCR (PP-OCRv6) ───────────────────────────────────────
# ── text: PaddleOCR defaults ─────────────────────────────────────────
# Our lang codes -> PaddleOCR lang. The 'ch' model also reads Latin; 'ru' reads
# Cyrillic + Latin. Multiple langs in one image -> run each model, union detections.
@@ -165,9 +168,7 @@ def _ocr_lines(bgr: np.ndarray, langs: list[str], min_score: float = 0.5) -> lis
def _cer(ref: str, hyp: str) -> float:
from rapidfuzz.distance import Levenshtein
return Levenshtein.normalized_distance(_norm(ref), _norm(hyp))
return levenshtein_normalized(_norm(ref), _norm(hyp))
# ── face: detection + ArcFace + texture ──────────────────────────────
+128
View File
@@ -0,0 +1,128 @@
# /// script
# requires-python = ">=3.11"
# dependencies = [
# "click",
# "huggingface-hub>=0.20.0",
# "numpy",
# "onnxruntime>=1.24.0",
# "opencv-python-headless<5",
# "paddleocr>=3.3.3",
# "paddlepaddle",
# "pillow",
# ]
# ///
"""Infer stable source-text lines without modifying an image.
This evaluation-only dry run proposes line annotations for selective text
restoration. Every proposal still needs human verification: stable OCR can lose
punctuation with high confidence. It separately flags lines whose recognition
changes under crop jitter or whose minimum confidence is below the threshold.
"""
from __future__ import annotations
import importlib.util
import json
import logging
import os
import sys
import unicodedata
from pathlib import Path
from typing import Any
import click
import numpy as np
from PIL import Image
log = logging.getLogger(__name__)
ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(0, str(ROOT))
RESTORATION_SCRIPT = ROOT / "scripts/selective_text_restoration.py"
from scripts._text_eval import normalize_text # noqa: E402
def _load_restoration_module() -> Any:
spec = importlib.util.spec_from_file_location("selective_text_restoration_for_inference", RESTORATION_SCRIPT)
if spec is None or spec.loader is None:
raise RuntimeError(f"unable to load {RESTORATION_SCRIPT}")
module = importlib.util.module_from_spec(spec)
sys.modules[spec.name] = module
spec.loader.exec_module(module)
return module
def _has_script(text: str, script: str) -> bool:
return any(script in unicodedata.name(character, "") for character in text)
def choose_language(probes: dict[str, tuple[str, float]]) -> str:
if _has_script(probes["ch"][0], "CJK"):
return "ch"
if _has_script(probes["ru"][0], "CYRILLIC"):
return "ru"
return "en"
def stable_recognition(reads: list[tuple[str, float]], min_score: float = 0.85) -> str | None:
normalized = {normalize_text(text) for text, _score in reads}
if len(normalized) != 1 or min(score for _text, score in reads) < min_score:
return None
return reads[0][0]
@click.command()
@click.argument("source", type=click.Path(exists=True, dir_okay=False, path_type=Path))
@click.option("--out", required=True, type=click.Path(dir_okay=False, path_type=Path))
@click.option("--min-score", default=0.85, show_default=True, type=click.FloatRange(0.0, 1.0))
def main(source: Path, out: Path, min_score: float) -> None:
"""Write draft line text for SOURCE; manually verify every proposal."""
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
os.environ["PADDLE_PDX_DISABLE_MODEL_SOURCE_CHECK"] = "True"
from paddleocr import PaddleOCR, TextRecognition
restoration = _load_restoration_module()
source_rgb = np.asarray(Image.open(source).convert("RGB"))
detector = PaddleOCR(
lang="ch",
use_doc_orientation_classify=False,
use_doc_unwarping=False,
use_textline_orientation=False,
)
engines = {
"en": TextRecognition(model_name="en_PP-OCRv5_mobile_rec"),
"ru": TextRecognition(model_name="eslav_PP-OCRv5_mobile_rec"),
"ch": TextRecognition(model_name="PP-OCRv5_server_rec"),
}
boxes = restoration.detect_line_boxes(detector, source_rgb)
accepted = []
rejected = []
for box in boxes:
probes = {}
for language, engine in engines.items():
script = "cjk" if language == "ch" else "alphabetic"
line = restoration.TextLine(box, "", script)
probes[language] = restoration._recognize(engine, source_rgb, line, 0.1)
language = choose_language(probes)
script = "cjk" if language == "ch" else "alphabetic"
line = restoration.TextLine(box, "", script)
reads = [restoration._recognize(engines[language], source_rgb, line, ratio) for ratio in (0.08, 0.12, 0.2)]
text = stable_recognition(reads, min_score)
result = {
"box": box,
"script": script,
"language": language,
"reads": [{"text": value, "score": score} for value, score in reads],
}
if text is None:
rejected.append(result)
else:
accepted.append({"box": box, "text": text, "script": script, "min_score": min(score for _, score in reads)})
payload = {"source": source.name, "accepted": accepted, "rejected": rejected}
out.parent.mkdir(parents=True, exist_ok=True)
out.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
log.info("Accepted %s lines and rejected %s uncertain lines", len(accepted), len(rejected))
if __name__ == "__main__":
main()
+528
View File
@@ -0,0 +1,528 @@
# /// script
# requires-python = ">=3.11"
# dependencies = [
# "click",
# "huggingface-hub>=0.20.0",
# "numpy",
# "onnxruntime>=1.24.0",
# "opencv-python-headless<5",
# "paddleocr>=3.3.3",
# "paddlepaddle",
# "pillow",
# ]
# ///
"""Evaluation-only text restoration over a scrubbed image.
``vae-glyphs`` composites only thresholded glyph-core pixels from a separately
generated VAE reconstruction over a fresh silhouette edge. ``source-glyphs``
can preserve typeface, layout, color, and antialiasing when
SOURCE is itself a regenerated layer. It must not be treated as safe when SOURCE
is the watermarked original: a provider oracle detected SynthID after that exact
paste-back experiment. ``source-silhouette`` instead transfers only a
thresholded glyph shape, then synthesizes fresh flat-color pixels and
antialiasing. ``rerender`` retains the system-font negative control. This is
not a production stage and does not add PaddleOCR to the package graph.
"""
from __future__ import annotations
import json
import logging
import math
import os
import shutil
import sys
import unicodedata
from dataclasses import asdict
from pathlib import Path
from typing import Any
import click
import cv2
import numpy as np
from PIL import Image, ImageDraw, ImageFont
log = logging.getLogger(__name__)
ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(0, str(ROOT))
sys.path.insert(0, str(ROOT / "src"))
from remove_ai_watermarks import region_eraser # noqa: E402
from remove_ai_watermarks._internal.text_restoration import VerifiedTextLine as TextLine # noqa: E402
from remove_ai_watermarks._internal.text_restoration import ( # noqa: E402
composite_fresh_text_edges,
composite_reconstructed_glyphs,
group_text_lines,
residual_glyph_mask,
source_silhouette_mask,
)
from scripts._text_eval import normalize_text, normalized_edit_distance # noqa: E402
if ROOT not in Path(region_eraser.__file__).resolve().parents:
raise RuntimeError("selective_text_restoration imported outside the current worktree")
REGULAR_FONT = Path("/System/Library/Fonts/Supplemental/Arial.ttf")
BOLD_FONT = Path("/System/Library/Fonts/Supplemental/Arial Bold.ttf")
CJK_FONT = Path("/System/Library/Fonts/STHeiti Medium.ttc")
def should_preserve_line(
expected: str,
source_text: str,
source_score: float,
candidate_text: str,
candidate_score: float,
) -> bool:
if min(source_score, candidate_score) < 0.75:
return False
if normalized_edit_distance(expected, source_text) > 0.25:
return False
return normalize_text(source_text) == normalize_text(candidate_text)
def composite_source_glyphs(
source_rgb: np.ndarray,
background_rgb: np.ndarray,
glyph_mask: np.ndarray,
*,
feather: float = 0.7,
) -> np.ndarray:
"""Composite exact source pixels inside a glyph mask with an outer feather."""
if source_rgb.shape != background_rgb.shape or source_rgb.shape[:2] != glyph_mask.shape:
raise ValueError("source, background, and glyph mask dimensions must match")
blurred = cv2.GaussianBlur(glyph_mask, (0, 0), feather) if feather > 0 else glyph_mask
alpha = np.maximum(glyph_mask, blurred).astype(np.float32) / 255.0
alpha = alpha[..., None]
combined = source_rgb.astype(np.float32) * alpha + background_rgb.astype(np.float32) * (1.0 - alpha)
return np.clip(combined, 0, 255).astype(np.uint8)
def source_box_mask(
shape: tuple[int, int],
boxes: list[tuple[int, int, int, int]],
) -> np.ndarray:
"""Build a padded text-line mask for aligned regenerated layer compositing."""
height, width = shape
mask = np.zeros((height, width), dtype=np.uint8)
for x1, y1, x2, y2 in boxes:
pad = max(8, (y2 - y1) // 4)
x1, y1, x2, y2 = _clip_box((x1, y1, x2, y2), width, height, pad=pad)
mask[y1:y2, x1:x2] = 255
return mask
def _clip_box(box: tuple[int, int, int, int], width: int, height: int, pad: int = 0) -> tuple[int, int, int, int]:
x1, y1, x2, y2 = box
return max(0, x1 - pad), max(0, y1 - pad), min(width, x2 + pad), min(height, y2 + pad)
def foreground_mask(source_rgb: np.ndarray, box: tuple[int, int, int, int]) -> np.ndarray:
height, width = source_rgb.shape[:2]
line_height = box[3] - box[1]
x1, y1, x2, y2 = _clip_box(box, width, height, pad=max(6, int(line_height * 0.12)))
gray = cv2.cvtColor(source_rgb[y1:y2, x1:x2], cv2.COLOR_RGB2GRAY)
ring_pad = max(8, min(24, (y2 - y1) // 5))
rx1, ry1, rx2, ry2 = _clip_box((x1, y1, x2, y2), width, height, pad=ring_pad)
context = cv2.cvtColor(source_rgb[ry1:ry2, rx1:rx2], cv2.COLOR_RGB2GRAY)
ring = np.ones(context.shape, dtype=bool)
ring[y1 - ry1 : y2 - ry1, x1 - rx1 : x2 - rx1] = False
background_luma = float(np.median(context[ring])) if ring.any() else float(np.median(gray))
low, high = float(np.percentile(gray, 4)), float(np.percentile(gray, 96))
dark_contrast, light_contrast = background_luma - low, high - background_luma
contrast = max(light_contrast, dark_contrast)
threshold = max(24.0, min(72.0, contrast * 0.32))
if light_contrast > dark_contrast:
mask = (gray.astype(np.float32) >= background_luma + threshold).astype(np.uint8) * 255
else:
mask = (gray.astype(np.float32) <= background_luma - threshold).astype(np.uint8) * 255
mask = cv2.morphologyEx(mask, cv2.MORPH_OPEN, np.ones((2, 2), np.uint8))
dilation = 5 if line_height >= 48 else 3
mask = cv2.dilate(mask, cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (2 * dilation + 1,) * 2))
result = np.zeros((height, width), dtype=np.uint8)
result[y1:y2, x1:x2] = mask
return result
def _line_language(line: TextLine) -> str:
if line.script == "cjk":
return "ch"
return "ru" if any("CYRILLIC" in unicodedata.name(character, "") for character in line.text) else "en"
def _recognition_box(
line: TextLine,
width: int,
height: int,
vertical_pad_ratio: float | None = None,
) -> tuple[int, int, int, int]:
x1, y1, x2, y2 = line.box
line_height = y2 - y1
if line.script == "cjk":
left_pad = max(16, round(line_height * 0.2))
right_pad = max(16, round(line_height * 0.6))
return max(0, x1 - left_pad), y1, min(width, x2 + right_pad), y2
pad_x = max(16, line_height)
pad_y = max(8, line_height // 3) if vertical_pad_ratio is None else max(8, round(line_height * vertical_pad_ratio))
return max(0, x1 - pad_x), max(0, y1 - pad_y), min(width, x2 + pad_x), min(height, y2 + pad_y)
def _recognize(
engine: Any,
image: np.ndarray,
line: TextLine,
vertical_pad_ratio: float | None = None,
) -> tuple[str, float]:
height, width = image.shape[:2]
x1, y1, x2, y2 = _recognition_box(line, width, height, vertical_pad_ratio)
crop = image[y1:y2, x1:x2]
if crop.shape[0] < 64:
scale = 64 / crop.shape[0]
crop = cv2.resize(crop, None, fx=scale, fy=scale, interpolation=cv2.INTER_CUBIC)
result = next(iter(engine.predict(crop)))
return str(result.get("rec_text", "")), float(result.get("rec_score", 0.0))
def _sample_text_color(
source_rgb: np.ndarray,
mask: np.ndarray,
box: tuple[int, int, int, int],
) -> tuple[int, int, int]:
height, width = source_rgb.shape[:2]
x1, y1, x2, y2 = _clip_box(box, width, height, pad=2)
crop = source_rgb[y1:y2, x1:x2]
active = mask[y1:y2, x1:x2] > 0
pixels = crop[active]
luma = pixels.mean(axis=1)
background_luma = float(crop[[0, -1], :, :].reshape(-1, 3).mean(axis=1).mean())
if background_luma >= 128:
selected = pixels[luma <= np.percentile(luma, 20)]
else:
selected = pixels[luma >= np.percentile(luma, 80)]
return tuple(int(value) for value in np.median(selected, axis=0))
def _render_line(image: Image.Image, line: TextLine, color: tuple[int, int, int]) -> None:
font_path = CJK_FONT if line.script == "cjk" else (BOLD_FONT if line.box[3] - line.box[1] >= 55 else REGULAR_FONT)
target_width, target_height = line.box[2] - line.box[0], line.box[3] - line.box[1]
draw = ImageDraw.Draw(image)
low, high = 4, max(8, target_height * 2)
font = ImageFont.truetype(str(font_path), low)
while low <= high:
size = (low + high) // 2
candidate = ImageFont.truetype(str(font_path), size)
bounds = draw.textbbox((0, 0), line.text, font=candidate)
if bounds[2] - bounds[0] <= target_width * 1.03 and bounds[3] - bounds[1] <= target_height * 1.08:
font, low = candidate, size + 1
else:
high = size - 1
bounds = draw.textbbox((0, 0), line.text, font=font)
y = line.box[1] + math.floor((target_height - (bounds[3] - bounds[1])) / 2) - bounds[1]
draw.text((line.box[0], y), line.text, fill=color, font=font)
def _write_manifest(path: Path | None, payload: dict[str, Any]) -> None:
if path is None:
return
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
def _vertical_overlap_ratio(left: tuple[int, int, int, int], right: tuple[int, int, int, int]) -> float:
overlap = max(0, min(left[3], right[3]) - max(left[1], right[1]))
return overlap / max(1, min(left[3] - left[1], right[3] - right[1]))
def group_word_boxes(boxes: list[tuple[int, int, int, int]]) -> list[tuple[int, int, int, int]]:
groups: list[tuple[int, int, int, int]] = []
for box in sorted(boxes, key=lambda item: ((item[1] + item[3]) / 2, item[0])):
matches = []
for index, group in enumerate(groups):
if _vertical_overlap_ratio(box, group) < 0.45:
continue
horizontal_gap = max(0, max(box[0], group[0]) - min(box[2], group[2]))
line_height = min(box[3] - box[1], group[3] - group[1])
if horizontal_gap <= max(24, line_height * 3):
matches.append(index)
if not matches:
groups.append(box)
continue
index = max(matches, key=lambda item: _vertical_overlap_ratio(box, groups[item]))
x1, y1, x2, y2 = groups[index]
groups[index] = min(x1, box[0]), min(y1, box[1]), max(x2, box[2]), max(y2, box[3])
return sorted(groups, key=lambda item: ((item[1] + item[3]) / 2, item[0]))
def detect_line_boxes(
engine: Any,
source_rgb: np.ndarray,
expected_count: int | None = None,
) -> list[tuple[int, int, int, int]]:
boxes: list[tuple[int, int, int, int]] = []
for page in engine.predict(source_rgb):
detected = page.get("rec_boxes", None)
if detected is None or len(detected) == 0:
detected = page.get("rec_polys", [])
for score, raw_box in zip(page.get("rec_scores", []), detected, strict=False):
if float(score) < 0.5:
continue
points = np.asarray(raw_box, dtype=np.float32).reshape(-1)
if points.size == 4:
x1, y1, x2, y2 = points
else:
points = points.reshape(-1, 2)
x1, y1 = points.min(axis=0)
x2, y2 = points.max(axis=0)
boxes.append((round(float(x1)), round(float(y1)), round(float(x2)), round(float(y2))))
lines = group_word_boxes(boxes)
if expected_count is not None and len(lines) != expected_count:
raise click.ClickException(f"detected {len(lines)} source lines; expected exactly {expected_count}")
return lines
def _load_lines(path: Path, key: str) -> list[TextLine]:
payload = json.loads(path.read_text(encoding="utf-8"))
return [
TextLine(tuple(item["box"]), item["text"], item["script"], float(item.get("angle", 0.0)))
for item in payload[key]
]
@click.command()
@click.argument("source", type=click.Path(exists=True, dir_okay=False, path_type=Path))
@click.argument("candidate", type=click.Path(exists=True, dir_okay=False, path_type=Path))
@click.option("--output", required=True, type=click.Path(dir_okay=False, path_type=Path))
@click.option(
"--lines-json",
default=ROOT / "data/evaluations/fidelity/text-lines.json",
type=click.Path(exists=True, dir_okay=False, path_type=Path),
)
@click.option("--source-key", help="Key in lines JSON; defaults to source basename.")
@click.option(
"--detect-boxes",
is_flag=True,
help="Detect source boxes; fail unless their count matches verified lines.",
)
@click.option(
"--restoration",
type=click.Choice(("vae-glyphs", "source-glyphs", "source-silhouette", "rerender")),
required=True,
help="Choose VAE glyph cores, regenerated pixels, fresh source shapes, or the system-font control.",
)
@click.option(
"--glyph-donor",
type=click.Path(exists=True, dir_okay=False, path_type=Path),
help="VAE reconstruction used only by --restoration vae-glyphs.",
)
@click.option(
"--glyph-feather",
type=click.FloatRange(min=0.0),
default=0.5,
show_default=True,
help="Outer donor-edge feather used only by --restoration vae-glyphs.",
)
@click.option(
"--selection",
type=click.Choice(("all", "changed")),
default="all",
show_default=True,
help="Restore every verified line or only OCR-confirmed changes.",
)
@click.option(
"--erase-background/--keep-background",
default=True,
show_default=True,
help="Erase candidate glyphs before compositing, or directly blend an aligned regenerated glyph layer.",
)
@click.option(
"--composite-mask",
type=click.Choice(("glyphs", "boxes")),
default="glyphs",
show_default=True,
help="Composite isolated glyphs or complete aligned text-line boxes.",
)
@click.option("--manifest", type=click.Path(dir_okay=False, path_type=Path))
def main(
source: Path,
candidate: Path,
output: Path,
lines_json: Path,
source_key: str | None,
detect_boxes: bool,
restoration: str,
glyph_donor: Path | None,
glyph_feather: float,
selection: str,
erase_background: bool,
composite_mask: str,
manifest: Path | None,
) -> None:
"""Restore SOURCE text over the scrubbed CANDIDATE."""
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
if restoration == "rerender":
for font in (REGULAR_FONT, BOLD_FONT, CJK_FONT):
if not font.exists():
raise click.ClickException(f"required evaluation font is unavailable: {font}")
if restoration == "vae-glyphs" and glyph_donor is None:
raise click.ClickException("--glyph-donor is required for --restoration vae-glyphs")
if restoration != "vae-glyphs" and glyph_donor is not None:
raise click.ClickException("--glyph-donor is only valid with --restoration vae-glyphs")
os.environ["PADDLE_PDX_DISABLE_MODEL_SOURCE_CHECK"] = "True"
source_rgb = np.asarray(Image.open(source).convert("RGB"))
candidate_rgb = np.asarray(Image.open(candidate).convert("RGB"))
if candidate_rgb.shape != source_rgb.shape:
raise click.ClickException("source and candidate dimensions must match")
donor_rgb = np.asarray(Image.open(glyph_donor).convert("RGB")) if glyph_donor else None
if donor_rgb is not None and donor_rgb.shape != source_rgb.shape:
raise click.ClickException("source and glyph donor dimensions must match")
lines = _load_lines(lines_json, source_key or source.name)
annotation_boxes = [line.box for line in lines]
if detect_boxes:
from paddleocr import PaddleOCR
page_engine = PaddleOCR(
lang="ch",
use_doc_orientation_classify=False,
use_doc_unwarping=False,
use_textline_orientation=False,
)
boxes = detect_line_boxes(page_engine, source_rgb, len(lines))
lines = [TextLine(box, line.text, line.script) for line, box in zip(lines, boxes, strict=True)]
vertical_pad_ratio = 0.1 if detect_boxes else None
decisions = []
selected = list(lines)
if selection == "all":
decisions = [{"line": asdict(line), "selected": True, "reason": "all-lines"} for line in lines]
else:
from paddleocr import TextRecognition
engines = {
"en": TextRecognition(model_name="en_PP-OCRv5_mobile_rec"),
"ru": TextRecognition(model_name="eslav_PP-OCRv5_mobile_rec"),
"ch": TextRecognition(model_name="PP-OCRv5_server_rec"),
}
selected = []
for line in lines:
language = _line_language(line)
source_text, source_score = _recognize(engines[language], source_rgb, line, vertical_pad_ratio)
candidate_text, candidate_score = _recognize(engines[language], candidate_rgb, line, vertical_pad_ratio)
preserve = should_preserve_line(line.text, source_text, source_score, candidate_text, candidate_score)
decisions.append(
{
"line": asdict(line),
"source_text": source_text,
"source_score": source_score,
"candidate_text": candidate_text,
"candidate_score": candidate_score,
"preserve": preserve,
"selected": not preserve,
}
)
if not preserve:
selected.append(line)
output.parent.mkdir(parents=True, exist_ok=True)
mask_path = output.with_name(output.stem + "_mask.png")
manifest_common = {
"source": source.name,
"candidate": candidate.name,
"output": output.name,
"mask": mask_path.name,
"glyph_donor": glyph_donor.name if glyph_donor else None,
"glyph_feather": glyph_feather if restoration == "vae-glyphs" else None,
"restoration": restoration,
"selection": selection,
"erase_background": erase_background,
"composite_mask": composite_mask,
"box_source": "detector" if detect_boxes else "verified_annotations",
"annotation_boxes": annotation_boxes,
"decisions": decisions,
}
if not selected:
combined = np.zeros(source_rgb.shape[:2], dtype=np.uint8)
shutil.copyfile(candidate, output)
Image.fromarray(combined).save(mask_path)
payload = {
**manifest_common,
"mask_fraction": 0.0,
"source_glyph_fraction": 0.0,
"source_layer_fraction": 0.0,
}
_write_manifest(manifest, payload)
log.info("Copied %s unchanged because every line passed", output)
return
if restoration in {"source-silhouette", "vae-glyphs"}:
source_masks = [source_silhouette_mask(source_rgb, line.box, line.angle) for line in selected]
candidate_masks = [source_silhouette_mask(candidate_rgb, line.box, line.angle) for line in selected]
line_masks = []
for line, source_mask, candidate_mask in zip(selected, source_masks, candidate_masks, strict=True):
radius = 5 if line.box[3] - line.box[1] >= 48 else 3
kernel = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (2 * radius + 1,) * 2)
line_masks.append(cv2.dilate(np.maximum(source_mask, candidate_mask), kernel))
masks = line_masks
else:
source_masks = [foreground_mask(source_rgb, line.box) for line in selected]
candidate_masks = [foreground_mask(candidate_rgb, line.box) for line in selected]
masks = [np.maximum(left, right) for left, right in zip(source_masks, candidate_masks, strict=True)]
del candidate_masks
groups = group_text_lines(selected)
if erase_background:
background = cv2.cvtColor(candidate_rgb, cv2.COLOR_RGB2BGR)
for group in groups:
background = region_eraser.erase_lama(
background,
np.maximum.reduce([masks[index] for index in group]),
)
background_rgb = cv2.cvtColor(background, cv2.COLOR_BGR2RGB)
residual_masks = [
residual_glyph_mask(background_rgb, mask, line.box) for line, mask in zip(selected, masks, strict=True)
]
for group in groups:
residual = np.maximum.reduce([residual_masks[index] for index in group])
if np.any(residual):
background = region_eraser.erase_lama(background, residual)
background_rgb = cv2.cvtColor(background, cv2.COLOR_BGR2RGB)
else:
background_rgb = candidate_rgb
residual_masks = []
source_glyph_mask = np.maximum.reduce(source_masks)
source_layer_mask = source_glyph_mask
if restoration == "source-glyphs":
source_layer_mask = (
source_box_mask(source_rgb.shape[:2], [line.box for line in selected])
if composite_mask == "boxes"
else source_glyph_mask
)
restored = composite_source_glyphs(source_rgb, background_rgb, source_layer_mask, feather=3.0)
Image.fromarray(restored).save(output)
elif restoration in {"source-silhouette", "vae-glyphs"}:
restored = composite_fresh_text_edges(source_rgb, background_rgb, selected, source_masks)
if restoration == "vae-glyphs":
if donor_rgb is None:
raise RuntimeError("VAE glyph restoration requires a loaded donor")
restored = composite_reconstructed_glyphs(
donor_rgb,
restored,
source_layer_mask,
feather=glyph_feather,
)
Image.fromarray(restored).save(output)
else:
rendered = Image.fromarray(background_rgb)
for line, source_mask in zip(selected, source_masks, strict=True):
_render_line(rendered, line, _sample_text_color(source_rgb, source_mask, line.box))
rendered.save(output)
combined = np.maximum.reduce([*masks, *residual_masks])
if restoration == "source-glyphs" and not erase_background:
combined = source_layer_mask
Image.fromarray(combined).save(mask_path)
payload = {
**manifest_common,
"mask_fraction": float((combined > 0).mean()),
"source_glyph_fraction": float((source_glyph_mask > 0).mean()),
"source_layer_fraction": float((source_layer_mask > 0).mean()),
}
_write_manifest(manifest, payload)
log.info("Wrote %s with %.4f edited fraction", output, payload["mask_fraction"])
if __name__ == "__main__":
main()
@@ -30,6 +30,8 @@ from remove_ai_watermarks._internal.watermark_profiles import resolve_seed
if TYPE_CHECKING:
from collections.abc import Callable
from remove_ai_watermarks._internal.text_restoration import VerifiedTextManifest
log = logging.getLogger(__name__)
QWEN_IMAGE_2512_MODEL_ID = "Qwen/Qwen-Image-2512"
@@ -973,6 +975,30 @@ class QwenZImagePipeline:
result = result.resize(image.size, Image.Resampling.LANCZOS)
return result.convert("RGB")
def _qwen_vae_roundtrip(self, image: Image.Image) -> Image.Image:
"""Reconstruct source pixels through the already loaded Qwen VAE."""
import torch
pipe, _controlnet_input_cls = self._load_qwen()
source_width, source_height = image.size
pad_width = (-source_width) % 8
pad_height = (-source_height) % 8
padded = image.convert("RGB")
if pad_width or pad_height:
padded = Image.fromarray(
np.pad(
np.asarray(padded),
((0, pad_height), (0, pad_width), (0, 0)),
mode="edge",
)
)
pipe.load_models_to_device(["vae"])
tensor = pipe.preprocess_image(padded).to(device=self.device, dtype=self.torch_dtype)
with torch.inference_mode():
latents = pipe.vae.encode(tensor)
decoded = pipe.vae.decode(latents)
return pipe.vae_output_to_image(decoded).crop((0, 0, source_width, source_height)).convert("RGB")
@staticmethod
def _detail_size(
crop_size: tuple[int, int],
@@ -1044,10 +1070,15 @@ class QwenZImagePipeline:
tile: bool = False,
tile_size: int = 1024,
tile_overlap: int = 128,
text_manifest: VerifiedTextManifest | None = None,
) -> Image.Image:
"""Execute global regeneration and masked face repair."""
self._require_cuda()
seed = resolve_seed(seed)
donor = None
if text_manifest is not None:
self._progress("Reconstructing the verified text donor with the Qwen VAE...")
donor = self._qwen_vae_roundtrip(image)
global_strength = (
resolution_adaptive_denoise(image.width, image.height) if strength is None else float(strength)
)
@@ -1068,14 +1099,28 @@ class QwenZImagePipeline:
boxes = detect_faces(image)
if not boxes:
self._progress("No faces detected; keeping the Qwen global result.")
return global_result
masks = self._sam_masks(image, boxes)
face_strength = largest_face_denoise(boxes, image.size) * FACE_DENOISE_SCALE
return self._run_faces(
image,
global_result,
boxes,
masks,
strength=face_strength,
seed=seed,
result = global_result
else:
masks = self._sam_masks(image, boxes)
face_strength = largest_face_denoise(boxes, image.size) * FACE_DENOISE_SCALE
result = self._run_faces(
image,
global_result,
boxes,
masks,
strength=face_strength,
seed=seed,
)
if text_manifest is None:
return result
if donor is None:
raise RuntimeError("Verified text restoration requires a Qwen-VAE donor")
from remove_ai_watermarks._internal.text_restoration import (
blend_fidelity_anchor,
restore_verified_text,
)
self._progress("Blending the Qwen-VAE fidelity anchor...")
anchor = blend_fidelity_anchor(result, donor)
self._progress(f"Restoring {len(text_manifest.lines)} verified text lines...")
return restore_verified_text(image, anchor, donor, text_manifest.lines)
@@ -0,0 +1,353 @@
"""Opt-in restoration of verified text from a Qwen VAE reconstruction."""
# pyright: reportUnknownMemberType=false, reportUnknownArgumentType=false, reportUnknownVariableType=false, reportUnknownParameterType=false, reportMissingTypeArgument=false, reportMissingTypeStubs=false, reportMissingImports=false, reportArgumentType=false, reportAssignmentType=false, reportReturnType=false, reportCallIssue=false, reportIndexIssue=false, reportOperatorIssue=false
from __future__ import annotations
import hashlib
import json
import math
from dataclasses import dataclass
from typing import TYPE_CHECKING, Any
import cv2
import numpy as np
from PIL import Image
if TYPE_CHECKING:
from collections.abc import Sequence
from pathlib import Path
from numpy.typing import NDArray
TEXT_MANIFEST_SCHEMA = 1
FIDELITY_BLEND_ALPHA = 0.15
GLYPH_FEATHER = 0.5
@dataclass(frozen=True)
class VerifiedTextLine:
"""One operator-verified source line in source-pixel coordinates."""
box: tuple[int, int, int, int]
text: str
script: str
angle: float = 0.0
@dataclass(frozen=True)
class VerifiedTextManifest:
"""Text annotations cryptographically bound to one decoded RGB source."""
source_pixel_sha256: str
width: int
height: int
lines: tuple[VerifiedTextLine, ...]
def source_pixel_sha256(image: Image.Image) -> str:
"""Hash decoded RGB geometry and bytes, independent of container metadata."""
rgb = image.convert("RGB")
digest = hashlib.sha256()
digest.update(rgb.width.to_bytes(8, "big"))
digest.update(rgb.height.to_bytes(8, "big"))
digest.update(rgb.tobytes())
return digest.hexdigest()
def load_verified_text_manifest(path: Path, source: Image.Image) -> VerifiedTextManifest:
"""Load and validate a manually verified manifest for exactly ``source``."""
try:
payload = json.loads(path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError) as exc:
raise ValueError(f"Cannot read text manifest {path}: {exc}") from exc
if not isinstance(payload, dict):
raise ValueError("Text manifest must be a JSON object")
if payload.get("schema_version") != TEXT_MANIFEST_SCHEMA:
raise ValueError(f"Text manifest schema_version must be {TEXT_MANIFEST_SCHEMA}")
if payload.get("verified") is not True:
raise ValueError("Text manifest must contain verified=true after manual review")
rgb = source.convert("RGB")
width = _manifest_integer(payload, "width")
height = _manifest_integer(payload, "height")
if (width, height) != rgb.size:
raise ValueError(f"Text manifest dimensions {width}x{height} do not match source {rgb.width}x{rgb.height}")
expected_hash = payload.get("source_pixel_sha256")
if not isinstance(expected_hash, str) or len(expected_hash) != 64:
raise ValueError("Text manifest source_pixel_sha256 must be a 64-character SHA-256")
actual_hash = source_pixel_sha256(rgb)
if expected_hash.casefold() != actual_hash:
raise ValueError("Text manifest source_pixel_sha256 does not match the decoded source pixels")
raw_lines = payload.get("lines")
if not isinstance(raw_lines, list) or not raw_lines:
raise ValueError("Text manifest lines must be a non-empty list")
lines = tuple(_load_line(item, width, height, index) for index, item in enumerate(raw_lines))
if list(lines) != sorted(lines, key=lambda line: (line.box[1], line.box[0])):
raise ValueError("Text manifest lines must be in top-to-bottom, left-to-right reading order")
return VerifiedTextManifest(actual_hash, width, height, lines)
def _manifest_integer(payload: dict[str, Any], key: str) -> int:
value = payload.get(key)
if isinstance(value, bool) or not isinstance(value, int) or value <= 0:
raise ValueError(f"Text manifest {key} must be a positive integer")
return value
def _load_line(item: Any, width: int, height: int, index: int) -> VerifiedTextLine:
if not isinstance(item, dict):
raise ValueError(f"Text manifest line {index} must be an object")
raw_box = item.get("box")
if (
not isinstance(raw_box, list)
or len(raw_box) != 4
or any(isinstance(value, bool) or not isinstance(value, int) for value in raw_box)
):
raise ValueError(f"Text manifest line {index} box must contain four integers")
box = tuple(raw_box)
x1, y1, x2, y2 = box
if not (0 <= x1 < x2 <= width and 0 <= y1 < y2 <= height):
raise ValueError(f"Text manifest line {index} box is outside the source dimensions")
text = item.get("text")
script = item.get("script")
if not isinstance(text, str) or not text.strip():
raise ValueError(f"Text manifest line {index} text must be non-empty")
if not isinstance(script, str) or not script.strip():
raise ValueError(f"Text manifest line {index} script must be non-empty")
angle_value = item.get("angle", 0.0)
if isinstance(angle_value, bool) or not isinstance(angle_value, int | float):
raise ValueError(f"Text manifest line {index} angle must be numeric")
angle = float(angle_value)
if not math.isfinite(angle) or abs(angle) > 30.0:
raise ValueError(f"Text manifest line {index} angle must be between -30 and 30 degrees")
return VerifiedTextLine(box, text, script, angle)
def blend_fidelity_anchor(clean: Image.Image, donor: Image.Image) -> Image.Image:
"""Blend 15% Qwen-VAE reconstruction into the oracle-clean pipeline output."""
clean_rgb = np.asarray(clean.convert("RGB"), dtype=np.float32)
donor_rgb = np.asarray(donor.convert("RGB"), dtype=np.float32)
if clean_rgb.shape != donor_rgb.shape:
raise ValueError("Clean result and Qwen-VAE donor dimensions must match")
blended = np.rint(clean_rgb * (1.0 - FIDELITY_BLEND_ALPHA) + donor_rgb * FIDELITY_BLEND_ALPHA)
return Image.fromarray(np.clip(blended, 0, 255).astype(np.uint8))
def restore_verified_text(
source: Image.Image,
candidate: Image.Image,
donor: Image.Image,
lines: tuple[VerifiedTextLine, ...],
) -> Image.Image:
"""Erase candidate glyphs, then composite verified Qwen-VAE glyph cores."""
from remove_ai_watermarks import region_eraser
if not region_eraser.lama_available():
raise RuntimeError(
"Verified text restoration requires LaMa. Install: pip install 'remove-ai-watermarks[text-restoration]'"
)
source_rgb = np.asarray(source.convert("RGB"))
candidate_rgb = np.asarray(candidate.convert("RGB"))
donor_rgb = np.asarray(donor.convert("RGB"))
if source_rgb.shape != candidate_rgb.shape or source_rgb.shape != donor_rgb.shape:
raise ValueError("Source, candidate, and Qwen-VAE donor dimensions must match")
source_masks = [source_silhouette_mask(source_rgb, line.box, line.angle) for line in lines]
for index, mask in enumerate(source_masks):
if not np.any(mask):
raise ValueError(f"Verified text line {index} produced no source glyph pixels")
candidate_masks = [source_silhouette_mask(candidate_rgb, line.box, line.angle) for line in lines]
erase_masks = []
for line, source_mask, candidate_mask in zip(lines, source_masks, candidate_masks, strict=True):
radius = 5 if line.box[3] - line.box[1] >= 48 else 3
kernel = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (2 * radius + 1,) * 2)
erase_masks.append(cv2.dilate(np.maximum(source_mask, candidate_mask), kernel))
del candidate_masks
groups = group_text_lines(lines)
background = cv2.cvtColor(candidate_rgb, cv2.COLOR_RGB2BGR)
for group in groups:
background = region_eraser.erase_lama(background, np.maximum.reduce([erase_masks[index] for index in group]))
background_rgb = cv2.cvtColor(background, cv2.COLOR_BGR2RGB)
residual_masks = [
residual_glyph_mask(background_rgb, mask, line.box) for line, mask in zip(lines, erase_masks, strict=True)
]
for group in groups:
residual = np.maximum.reduce([residual_masks[index] for index in group])
if np.any(residual):
background = region_eraser.erase_lama(background, residual)
del erase_masks, residual_masks
restored = cv2.cvtColor(background, cv2.COLOR_BGR2RGB)
restored = composite_fresh_text_edges(source_rgb, restored, lines, source_masks)
source_glyph_mask = np.maximum.reduce(source_masks)
restored = composite_reconstructed_glyphs(donor_rgb, restored, source_glyph_mask)
return Image.fromarray(restored)
def source_silhouette_mask(
source_rgb: NDArray[Any],
box: tuple[int, int, int, int],
angle: float = 0.0,
) -> NDArray[Any]:
"""Recover a thresholded glyph shape without retaining source amplitudes."""
height, width = source_rgb.shape[:2]
x1, y1, x2, y2 = _clip_box(box, width, height)
gray = cv2.cvtColor(source_rgb[y1:y2, x1:x2], cv2.COLOR_RGB2GRAY)
support = np.ones(gray.shape, dtype=np.uint8)
if angle:
box_width, box_height = x2 - x1, y2 - y1
theta = math.radians(abs(angle))
cosine, sine = math.cos(theta), math.sin(theta)
denominator = cosine * cosine - sine * sine
rect_width = (box_width * cosine - box_height * sine) / denominator
rect_height = (box_height * cosine - box_width * sine) / denominator
rotated = cv2.boxPoints(
((box_width / 2, box_height / 2), (max(1.0, rect_width * 0.92), max(1.0, rect_height * 0.62)), -angle)
)
support.fill(0)
cv2.fillConvexPoly(support, np.rint(rotated).astype(np.int32), 1)
values = gray[support > 0]
background_luma = float(np.median(values))
else:
ring_pad = max(6, min(20, (y2 - y1) // 4))
rx1, ry1, rx2, ry2 = _clip_box((x1, y1, x2, y2), width, height, pad=ring_pad)
context = cv2.cvtColor(source_rgb[ry1:ry2, rx1:rx2], cv2.COLOR_RGB2GRAY)
ring = np.ones(context.shape, dtype=bool)
ring[y1 - ry1 : y2 - ry1, x1 - rx1 : x2 - rx1] = False
background_luma = float(np.median(context[ring])) if ring.any() else float(np.median(gray))
values = gray.reshape(-1)
low, high = float(np.percentile(values, 2)), float(np.percentile(values, 98))
dark_contrast, light_contrast = background_luma - low, high - background_luma
threshold = max(16.0, min(56.0, max(light_contrast, dark_contrast) * 0.22))
if light_contrast > dark_contrast:
crop_mask = (gray.astype(np.float32) >= background_luma + threshold).astype(np.uint8) * 255
else:
crop_mask = (gray.astype(np.float32) <= background_luma - threshold).astype(np.uint8) * 255
crop_mask[support == 0] = 0
result = np.zeros((height, width), dtype=np.uint8)
result[y1:y2, x1:x2] = crop_mask
return result
def residual_glyph_mask(
background_rgb: NDArray[Any],
original_mask: NDArray[Any],
box: tuple[int, int, int, int],
) -> NDArray[Any]:
"""Find glyph-like contrast left after the first inpaint pass."""
residual = _foreground_mask(background_rgb, box)
residual = cv2.bitwise_and(residual, original_mask)
return cv2.dilate(residual, np.ones((5, 5), np.uint8), iterations=1)
def composite_fresh_text_edges(
source_rgb: NDArray[Any],
background_rgb: NDArray[Any],
lines: tuple[VerifiedTextLine, ...],
masks: list[NDArray[Any]],
) -> NDArray[Any]:
"""Render fresh antialiased edges for source-derived glyph masks."""
restored = background_rgb
for line, mask in zip(lines, masks, strict=True):
color = _sample_text_color(source_rgb, mask, line.box)
restored = composite_fresh_silhouette(restored, mask, color)
return restored
def composite_reconstructed_glyphs(
donor_rgb: NDArray[Any],
background_rgb: NDArray[Any],
glyph_mask: NDArray[Any],
*,
feather: float = GLYPH_FEATHER,
) -> NDArray[Any]:
"""Composite an exact reconstructed core with a narrow donor edge."""
if donor_rgb.shape != background_rgb.shape or donor_rgb.shape[:2] != glyph_mask.shape:
raise ValueError("donor, background, and glyph mask dimensions must match")
blurred = cv2.GaussianBlur(glyph_mask, (0, 0), feather) if feather > 0 else glyph_mask
alpha = np.maximum(glyph_mask, blurred).astype(np.float32) / 255.0
combined = donor_rgb.astype(np.float32) * alpha[..., None] + background_rgb.astype(np.float32) * (
1.0 - alpha[..., None]
)
return np.clip(np.rint(combined), 0, 255).astype(np.uint8)
def composite_fresh_silhouette(
background_rgb: NDArray[Any],
glyph_mask: NDArray[Any],
color: tuple[int, int, int],
*,
feather: float = 0.35,
) -> NDArray[Any]:
"""Render a binary source shape with fresh color and antialiasing."""
if background_rgb.shape[:2] != glyph_mask.shape:
raise ValueError("background and glyph mask dimensions must match")
antialiased = cv2.GaussianBlur(glyph_mask, (0, 0), feather) if feather > 0 else glyph_mask
alpha = antialiased.astype(np.float32)[..., None] / 255.0
foreground = np.empty_like(background_rgb)
foreground[:, :] = color
combined = foreground.astype(np.float32) * alpha + background_rgb.astype(np.float32) * (1.0 - alpha)
return np.clip(combined, 0, 255).astype(np.uint8)
def _clip_box(box: tuple[int, int, int, int], width: int, height: int, pad: int = 0) -> tuple[int, int, int, int]:
x1, y1, x2, y2 = box
return max(0, x1 - pad), max(0, y1 - pad), min(width, x2 + pad), min(height, y2 + pad)
def _foreground_mask(source_rgb: NDArray[Any], box: tuple[int, int, int, int]) -> NDArray[Any]:
height, width = source_rgb.shape[:2]
line_height = box[3] - box[1]
x1, y1, x2, y2 = _clip_box(box, width, height, pad=max(6, int(line_height * 0.12)))
gray = cv2.cvtColor(source_rgb[y1:y2, x1:x2], cv2.COLOR_RGB2GRAY)
ring_pad = max(8, min(24, (y2 - y1) // 5))
rx1, ry1, rx2, ry2 = _clip_box((x1, y1, x2, y2), width, height, pad=ring_pad)
context = cv2.cvtColor(source_rgb[ry1:ry2, rx1:rx2], cv2.COLOR_RGB2GRAY)
ring = np.ones(context.shape, dtype=bool)
ring[y1 - ry1 : y2 - ry1, x1 - rx1 : x2 - rx1] = False
background_luma = float(np.median(context[ring])) if ring.any() else float(np.median(gray))
low, high = float(np.percentile(gray, 4)), float(np.percentile(gray, 96))
dark_contrast, light_contrast = background_luma - low, high - background_luma
threshold = max(24.0, min(72.0, max(light_contrast, dark_contrast) * 0.32))
if light_contrast > dark_contrast:
mask = (gray.astype(np.float32) >= background_luma + threshold).astype(np.uint8) * 255
else:
mask = (gray.astype(np.float32) <= background_luma - threshold).astype(np.uint8) * 255
mask = cv2.morphologyEx(mask, cv2.MORPH_OPEN, np.ones((2, 2), np.uint8))
dilation = 5 if line_height >= 48 else 3
mask = cv2.dilate(mask, cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (2 * dilation + 1,) * 2))
result = np.zeros((height, width), dtype=np.uint8)
result[y1:y2, x1:x2] = mask
return result
def _sample_text_color(
source_rgb: NDArray[Any], mask: NDArray[Any], box: tuple[int, int, int, int]
) -> tuple[int, int, int]:
height, width = source_rgb.shape[:2]
x1, y1, x2, y2 = _clip_box(box, width, height, pad=2)
crop = source_rgb[y1:y2, x1:x2]
pixels = crop[mask[y1:y2, x1:x2] > 0]
luma = pixels.mean(axis=1)
background_luma = float(crop[[0, -1], :, :].reshape(-1, 3).mean(axis=1).mean())
selected = (
pixels[luma <= np.percentile(luma, 20)] if background_luma >= 128 else pixels[luma >= np.percentile(luma, 80)]
)
return tuple(int(value) for value in np.median(selected, axis=0))
def group_text_lines(lines: Sequence[VerifiedTextLine]) -> list[list[int]]:
"""Group nearby same-script lines for a shared LaMa erase pass."""
groups: list[list[int]] = []
for index, line in enumerate(lines):
if not groups:
groups.append([index])
continue
previous = lines[groups[-1][-1]]
gap = line.box[1] - previous.box[3]
if line.script != previous.script or gap > max(60, int((previous.box[3] - previous.box[1]) * 1.1)):
groups.append([index])
else:
groups[-1].append(index)
return groups
@@ -26,6 +26,8 @@ if TYPE_CHECKING:
from collections.abc import Callable
from pathlib import Path
from remove_ai_watermarks._internal.text_restoration import VerifiedTextManifest
logger = logging.getLogger(__name__)
try:
@@ -187,6 +189,7 @@ class WatermarkRemover:
tile: bool = False,
tile_size: int = 1024,
tile_overlap: int = 128,
text_manifest: VerifiedTextManifest | None = None,
) -> Path:
"""Regenerate image pixels and write the result without AI metadata.
@@ -203,6 +206,10 @@ class WatermarkRemover:
resolved_strength = resolve_strength(strength, vendor, self.model_profile, size=source.size)
if not 0.0 <= resolved_strength <= 1.0:
raise ValueError(f"Strength must be between 0.0 and 1.0, got {resolved_strength}")
if text_manifest is not None and self.model_profile == SDXL_ZIMAGE_PROFILE:
raise ValueError("Verified text restoration is supported only by the qwen-zimage profile")
if text_manifest is not None and tile:
raise ValueError("Verified text restoration is not calibrated with tiled diffusion")
result = self._load_qwen_zimage_pipeline().run(
source,
@@ -211,6 +218,7 @@ class WatermarkRemover:
tile=tile,
tile_size=tile_size,
tile_overlap=tile_overlap,
text_manifest=text_manifest,
)
self._write_output(result, destination)
return destination
+2
View File
@@ -243,6 +243,7 @@ class InvisibleOptions:
tile: bool = False
tile_size: int = 1024
tile_overlap: int = 128
text_manifest: Path | None = None
# What the invisible stage did. "unavailable" is the one outcome the caller must
@@ -523,6 +524,7 @@ def _run_invisible(
tile=opts.tile,
tile_size=opts.tile_size,
tile_overlap=opts.tile_overlap,
text_manifest=opts.text_manifest,
)
say("invisible", "removed")
return "removed"
+34 -14
View File
@@ -311,6 +311,16 @@ _cpu_offload_option = click.option(
),
)
_text_manifest_option = click.option(
"--text-manifest",
type=click.Path(exists=True, dir_okay=False, path_type=Path),
default=None,
help=(
"Experimental verified-text restoration manifest. Requires qwen-zimage, "
"the text-restoration extra, native untiled geometry, and no postprocessing."
),
)
_visible_backend_option = click.option(
"--backend",
@@ -787,6 +797,7 @@ def cmd_erase(
@_tile_options
@_force_option
@_cpu_offload_option
@_text_manifest_option
@click.pass_context
def cmd_invisible(
ctx: click.Context,
@@ -806,6 +817,7 @@ def cmd_invisible(
tile_overlap: int,
force: bool,
cpu_offload: bool,
text_manifest: Path | None,
) -> None:
"""Remove invisible AI watermarks (SynthID, StableSignature, TreeRing).
@@ -853,20 +865,25 @@ def cmd_invisible(
console.print(f" Strength: {_resolved_strength_for_display(source, strength, vendor, pipeline)}")
t0 = time.monotonic()
result_path = engine.remove_watermark(
image_path=source,
output_path=output,
strength=strength,
seed=seed,
humanize=humanize,
unsharp=unsharp,
adaptive_polish=adaptive_polish,
max_resolution=max_resolution,
vendor=vendor,
tile=tile,
tile_size=tile_size,
tile_overlap=tile_overlap,
)
try:
result_path = engine.remove_watermark(
image_path=source,
output_path=output,
strength=strength,
seed=seed,
humanize=humanize,
unsharp=unsharp,
adaptive_polish=adaptive_polish,
max_resolution=max_resolution,
vendor=vendor,
tile=tile,
tile_size=tile_size,
tile_overlap=tile_overlap,
text_manifest=text_manifest,
)
except (OSError, RuntimeError, ValueError) as exc:
console.print(f" Error: {exc}")
raise SystemExit(1) from exc
elapsed = time.monotonic() - t0
size_kb = result_path.stat().st_size / 1024
@@ -1410,6 +1427,7 @@ def cmd_identify(ctx: click.Context, source: Path, no_visible: bool, as_json: bo
@_tile_options
@_force_option
@_cpu_offload_option
@_text_manifest_option
@click.pass_context
def cmd_all(
ctx: click.Context,
@@ -1431,6 +1449,7 @@ def cmd_all(
tile_overlap: int,
force: bool,
cpu_offload: bool,
text_manifest: Path | None,
) -> None:
"""Remove ALL watermarks: visible + invisible + metadata.
@@ -1508,6 +1527,7 @@ def cmd_all(
tile=tile,
tile_size=tile_size,
tile_overlap=tile_overlap,
text_manifest=text_manifest,
),
force=force,
progress=progress,
@@ -18,6 +18,7 @@ from typing import TYPE_CHECKING
from ._internal.watermark_profiles import (
DEFAULT_PROFILE,
QWEN_ZIMAGE_PROFILE,
REMOVAL_MODULES,
resolve_adaptive_polish,
resolve_seed,
@@ -148,6 +149,7 @@ class InvisibleEngine:
tile: bool = False,
tile_size: int = 1024,
tile_overlap: int = 128,
text_manifest: Path | None = None,
) -> Path:
"""Remove invisible watermark from an image.
@@ -180,6 +182,11 @@ class InvisibleEngine:
Engages only when the long side exceeds ``tile_size``.
tile_size: Tile dimension in px (default 1024).
tile_overlap: Overlap between adjacent tiles in px (default 128).
text_manifest: Operator-verified text lines bound to the decoded source
pixels. Enables the experimental Qwen-VAE ``vae-glyphs`` post-pass.
Requires the ``text-restoration`` extra and the ``qwen-zimage``
profile. Incompatible with tiling, downscaling, humanize, unsharp,
and adaptive polish because those combinations are not calibrated.
Returns:
Path to the cleaned image.
@@ -189,6 +196,23 @@ class InvisibleEngine:
seed = resolve_seed(seed)
adaptive_polish = resolve_adaptive_polish(adaptive_polish, self._remover.model_profile)
if text_manifest is not None:
if self._remover.model_profile != QWEN_ZIMAGE_PROFILE:
raise ValueError("--text-manifest is supported only by the qwen-zimage profile")
if max_resolution != 0:
raise ValueError("--text-manifest requires --max-resolution 0")
if tile:
raise ValueError("--text-manifest is not calibrated with --tile")
if humanize > 0.0 or unsharp > 0.0 or adaptive_polish:
raise ValueError("--text-manifest requires humanize=0, unsharp=0, and adaptive polish disabled")
from remove_ai_watermarks import region_eraser
if not region_eraser.lama_available():
raise RuntimeError(
"Verified text restoration requires LaMa. Install: "
"pip install 'remove-ai-watermarks[text-restoration]'"
)
from PIL import Image, ImageOps
# Resolution policy: a max_resolution cap (0 = none) bounds memory on huge
@@ -205,6 +229,11 @@ class InvisibleEngine:
# Full-res original, kept for the adaptive-polish detail target (image is
# reassigned to the resized copy below; PIL resize returns a new object).
reference_pil = image
verified_text = None
if text_manifest is not None:
from remove_ai_watermarks._internal.text_restoration import load_verified_text_manifest
verified_text = load_verified_text_manifest(text_manifest, reference_pil)
# Both profiles run at the input's native geometry, so only the explicit max
# cap can move it, and it can only ever scale down.
@@ -240,6 +269,7 @@ class InvisibleEngine:
tile=tile,
tile_size=tile_size,
tile_overlap=tile_overlap,
text_manifest=verified_text,
)
# Post-processing chain: decode the diffusion output ONCE, apply the
+1
View File
@@ -245,6 +245,7 @@ class TestInvisibleOptionsMirrorTheEngine:
tile=True,
tile_size=768,
tile_overlap=64,
text_manifest=tmp_path / "verified-lines.json",
)
seen: dict[str, object] = {}
+10 -2
View File
@@ -20,7 +20,7 @@ import pytest
_SCRIPTS = Path(__file__).resolve().parent.parent / "scripts"
def _load_assign():
def _load_module():
# fidelity_metrics is a standalone PEP723 script, not an installed module; load it by
# path with scripts/ on sys.path so its `_plain_console` shim import resolves.
sys.path.insert(0, str(_SCRIPTS))
@@ -35,7 +35,15 @@ def _load_assign():
pytest.skip(f"fidelity_metrics import deps missing: {exc}")
finally:
sys.path.remove(str(_SCRIPTS))
return mod.assign_faces_one_to_one
return mod
def _load_assign():
return _load_module().assign_faces_one_to_one
def test_cer_remains_case_sensitive() -> None:
assert _load_module()._cer("A", "a") == 1.0
def test_distinct_faces_match_nearest() -> None:
+27
View File
@@ -0,0 +1,27 @@
from __future__ import annotations
import importlib.util
import sys
from pathlib import Path
SCRIPT = Path(__file__).parents[1] / "scripts/infer_text_lines.py"
SPEC = importlib.util.spec_from_file_location("infer_text_lines", SCRIPT)
assert SPEC is not None
assert SPEC.loader is not None
module = importlib.util.module_from_spec(SPEC)
sys.modules[SPEC.name] = module
SPEC.loader.exec_module(module)
def test_choose_language_prefers_observed_unicode_script() -> None:
probes = {"en": ("gibberish", 0.9), "ru": ("пример", 0.9), "ch": ("example", 0.9)}
assert module.choose_language(probes) == "ru"
probes["ch"] = ("示例", 0.9)
assert module.choose_language(probes) == "ch"
def test_stable_recognition_requires_agreement_and_confidence() -> None:
assert module.stable_recognition([("Sample text", 0.9), ("sample text", 0.95)]) == "Sample text"
assert module.stable_recognition([("Sample", 0.9), ("Simple", 0.95)]) is None
assert module.stable_recognition([("Sample", 0.8), ("Sample", 0.95)]) is None
+70
View File
@@ -43,6 +43,76 @@ class TestInvisibleEngineInit:
assert engine._preload_kwargs == {"global_only": True}
class TestVerifiedTextMode:
"""The experimental mode must fail before loading models on unmeasured inputs."""
@staticmethod
def _engine(profile: str = "qwen-zimage") -> InvisibleEngine:
engine = object.__new__(InvisibleEngine)
engine._progress_callback = None
engine._remover = SimpleNamespace(model_profile=profile)
return engine
def test_rejects_incompatible_pipeline_options(self, tmp_path):
import pytest
manifest = tmp_path / "manifest.json"
manifest.write_text("{}", encoding="utf-8")
cases = (
("sdxl-zimage", {}, "qwen-zimage"),
("qwen-zimage", {"max_resolution": 1024}, "max-resolution 0"),
("qwen-zimage", {"tile": True}, "not calibrated"),
("qwen-zimage", {"humanize": 1.0}, "humanize=0"),
("qwen-zimage", {"adaptive_polish": True}, "polish disabled"),
)
for profile, kwargs, message in cases:
with pytest.raises(ValueError, match=message):
self._engine(profile).remove_watermark(
tmp_path / "unused.png",
text_manifest=manifest,
**kwargs,
)
def test_loads_and_forwards_verified_manifest(self, tmp_path, monkeypatch):
import json
from remove_ai_watermarks import region_eraser
from remove_ai_watermarks._internal.text_restoration import source_pixel_sha256
source = tmp_path / "source.png"
output = tmp_path / "output.png"
image = Image.new("RGB", (48, 32), (10, 20, 30))
image.save(source)
manifest = tmp_path / "manifest.json"
manifest.write_text(
json.dumps(
{
"schema_version": 1,
"verified": True,
"source_pixel_sha256": source_pixel_sha256(image),
"width": 48,
"height": 32,
"lines": [{"box": [8, 8, 40, 24], "text": "Exact", "script": "alphabetic"}],
}
),
encoding="utf-8",
)
seen = {}
def fake_remove(**kwargs):
seen.update(kwargs)
Image.open(kwargs["image_path"]).save(kwargs["output_path"])
return kwargs["output_path"]
engine = self._engine()
engine._remover.remove_watermark = fake_remove
monkeypatch.setattr(region_eraser, "lama_available", lambda: True)
engine.remove_watermark(source, output, text_manifest=manifest)
assert seen["text_manifest"].lines[0].text == "Exact"
class TestNativeOutputSize:
"""Model-side latent-grid rounding must not change the public output size."""
+73
View File
@@ -579,6 +579,78 @@ def test_cli_qwen_zimage_keeps_profile_postprocess_default(tmp_image_path, monke
assert mock_engine.remove_watermark.call_args.kwargs["adaptive_polish"] is True
def test_cli_forwards_verified_text_manifest(tmp_image_path, tmp_path, monkeypatch):
from remove_ai_watermarks import cli
manifest = tmp_path / "manifest.json"
manifest.write_text("{}", encoding="utf-8")
mock_engine = MagicMock()
mock_engine.remove_watermark.return_value = tmp_image_path
monkeypatch.setattr("remove_ai_watermarks.invisible_engine.is_available", lambda: True)
monkeypatch.setattr("remove_ai_watermarks.invisible_engine.InvisibleEngine", MagicMock(return_value=mock_engine))
result = CliRunner().invoke(
cli.main,
["invisible", str(tmp_image_path), "--text-manifest", str(manifest), "--force"],
)
assert result.exit_code == 0, result.output
assert mock_engine.remove_watermark.call_args.kwargs["text_manifest"] == manifest
def test_cli_reports_verified_text_manifest_errors(tmp_image_path, tmp_path, monkeypatch):
from remove_ai_watermarks import cli
manifest = tmp_path / "manifest.json"
manifest.write_text("{}", encoding="utf-8")
mock_engine = MagicMock()
mock_engine.remove_watermark.side_effect = ValueError("manifest pixels do not match")
monkeypatch.setattr("remove_ai_watermarks.invisible_engine.is_available", lambda: True)
monkeypatch.setattr("remove_ai_watermarks.invisible_engine.InvisibleEngine", MagicMock(return_value=mock_engine))
result = CliRunner().invoke(
cli.main,
["invisible", str(tmp_image_path), "--text-manifest", str(manifest), "--force"],
)
assert result.exit_code == 1
assert "manifest pixels do not match" in result.output
def test_no_face_path_still_runs_verified_text_restoration(monkeypatch):
from remove_ai_watermarks._internal import qwen_zimage_pipeline, text_restoration
from remove_ai_watermarks._internal.qwen_zimage_pipeline import QwenZImagePipeline
from remove_ai_watermarks._internal.text_restoration import VerifiedTextLine, VerifiedTextManifest
pipeline = object.__new__(QwenZImagePipeline)
pipeline.device = "cuda"
pipeline.progress_callback = None
source = Image.new("RGB", (32, 32), (10, 20, 30))
donor = Image.new("RGB", (32, 32), (40, 50, 60))
global_result = Image.new("RGB", (32, 32), (70, 80, 90))
anchor = Image.new("RGB", (32, 32), (100, 110, 120))
restored = Image.new("RGB", (32, 32), (130, 140, 150))
pipeline._qwen_vae_roundtrip = MagicMock(return_value=donor)
pipeline._run_global = MagicMock(return_value=global_result)
monkeypatch.setattr(qwen_zimage_pipeline, "detect_faces", lambda _image: [])
blend = MagicMock(return_value=anchor)
restore = MagicMock(return_value=restored)
monkeypatch.setattr(text_restoration, "blend_fidelity_anchor", blend)
monkeypatch.setattr(text_restoration, "restore_verified_text", restore)
manifest = VerifiedTextManifest(
"0" * 64,
32,
32,
(VerifiedTextLine((4, 4, 20, 16), "Exact", "alphabetic"),),
)
result = pipeline.run(source, strength=0.1, seed=0, text_manifest=manifest)
assert result is restored
blend.assert_called_once_with(global_result, donor)
restore.assert_called_once_with(source, anchor, donor, manifest.lines)
def test_watermark_remover_dispatches_to_full_pipeline(tmp_path, monkeypatch):
from remove_ai_watermarks._internal.watermark_remover import WatermarkRemover
@@ -601,6 +673,7 @@ def test_watermark_remover_dispatches_to_full_pipeline(tmp_path, monkeypatch):
_, kwargs = runtime.run.call_args
assert kwargs["strength"] == pytest.approx(0.084)
assert kwargs["seed"] == 0
assert kwargs["text_manifest"] is None
assert output.exists()
+181
View File
@@ -0,0 +1,181 @@
from __future__ import annotations
import importlib.util
import json
import sys
from pathlib import Path
import numpy as np
import pytest
from remove_ai_watermarks._internal import text_restoration
SCRIPT = Path(__file__).parents[1] / "scripts" / "selective_text_restoration.py"
SPEC = importlib.util.spec_from_file_location("selective_text_restoration", SCRIPT)
assert SPEC is not None
assert SPEC.loader is not None
module = importlib.util.module_from_spec(SPEC)
sys.modules[SPEC.name] = module
SPEC.loader.exec_module(module)
def test_normalized_edit_distance_ignores_case_and_whitespace() -> None:
assert module.normalized_edit_distance(" Sample text ", "sample\ntext") == 0.0
def test_preserve_requires_source_candidate_agreement() -> None:
assert module.should_preserve_line("clean text", "clean text", 0.9, "clean text", 0.9)
assert not module.should_preserve_line("clean text", "clean text", 0.9, "damaged", 0.9)
def test_preserve_rejects_unreliable_source_recognition() -> None:
assert not module.should_preserve_line("expected", "unrelated", 0.9, "unrelated", 0.9)
assert not module.should_preserve_line("expected", "expected", 0.7, "expected", 0.9)
def test_cjk_recognition_box_excludes_overlapping_neighbor_lines() -> None:
line = module.TextLine((1281, 650, 2357, 818), "每天都是一个新的机会。", "cjk")
assert module._recognition_box(line, 2816, 1536) == (1247, 650, 2458, 818)
def test_latin_recognition_box_keeps_context_padding() -> None:
line = module.TextLine((100, 200, 300, 260), "Sample text", "latin")
assert module._recognition_box(line, 1000, 1000) == (40, 180, 360, 280)
assert module._recognition_box(line, 1000, 1000, 0.1) == (40, 192, 360, 268)
def test_verified_lines_cover_each_ground_truth_string() -> None:
root = Path(__file__).parents[1]
lines = json.loads((root / "data/evaluations/fidelity/text-lines.json").read_text(encoding="utf-8"))
ground_truth = json.loads((root / "data/evaluations/fidelity/ground-truth.json").read_text(encoding="utf-8"))
assert lines.keys() == ground_truth.keys()
for source, expected in ground_truth.items():
observed = " ".join(line["text"] for line in lines[source])
assert module.normalize_text(observed) == module.normalize_text(expected)
def test_group_word_boxes_merges_words_but_not_neighboring_lines() -> None:
boxes = [(10, 10, 30, 30), (32, 12, 60, 29), (10, 35, 50, 55)]
assert module.group_word_boxes(boxes) == [(10, 10, 60, 30), (10, 35, 50, 55)]
def test_group_word_boxes_does_not_merge_distant_columns() -> None:
boxes = [(10, 10, 60, 30), (500, 11, 560, 31)]
assert module.group_word_boxes(boxes) == boxes
def test_source_glyph_composite_keeps_masked_pixels_exact() -> None:
source = np.zeros((9, 9, 3), dtype=np.uint8)
source[:, :] = (220, 180, 40)
background = np.zeros((9, 9, 3), dtype=np.uint8)
background[:, :] = (10, 20, 30)
mask = np.zeros((9, 9), dtype=np.uint8)
mask[3:6, 3:6] = 255
result = module.composite_source_glyphs(source, background, mask, feather=0.7)
np.testing.assert_array_equal(result[3:6, 3:6], source[3:6, 3:6])
np.testing.assert_array_equal(result[0, 0], background[0, 0])
def test_fresh_silhouette_uses_new_color_instead_of_source_pixels() -> None:
background = np.zeros((9, 9, 3), dtype=np.uint8)
background[:, :] = (10, 20, 30)
mask = np.zeros((9, 9), dtype=np.uint8)
mask[3:6, 3:6] = 255
result = text_restoration.composite_fresh_silhouette(background, mask, (220, 180, 40), feather=0)
assert np.all(result[3:6, 3:6] == (220, 180, 40))
np.testing.assert_array_equal(result[0, 0], background[0, 0])
def test_fresh_silhouette_antialiasing_softens_binary_edges() -> None:
background = np.zeros((9, 9, 3), dtype=np.uint8)
background[:, :] = (10, 20, 30)
mask = np.zeros((9, 9), dtype=np.uint8)
mask[3:6, 3:6] = 255
result = text_restoration.composite_fresh_silhouette(background, mask, (220, 180, 40), feather=1.0)
assert np.all(result[3, 3] > background[3, 3])
assert np.all(result[3, 3] < (220, 180, 40))
def test_reconstructed_glyphs_keep_exact_donor_core_and_fresh_edge() -> None:
donor = np.zeros((9, 9, 3), dtype=np.uint8)
donor[:, :] = (180, 140, 60)
background = np.zeros((9, 9, 3), dtype=np.uint8)
background[:, :] = (10, 20, 30)
mask = np.zeros((9, 9), dtype=np.uint8)
mask[3:6, 3:6] = 255
fresh_edge = text_restoration.composite_fresh_silhouette(background, mask, (220, 180, 40))
result = module.composite_reconstructed_glyphs(donor, fresh_edge, mask, feather=0.5)
np.testing.assert_array_equal(result[3:6, 3:6], donor[3:6, 3:6])
assert np.any(result[2, 3] != fresh_edge[2, 3])
np.testing.assert_array_equal(result[0, 0], background[0, 0])
def test_source_silhouette_discards_foreground_amplitudes() -> None:
source = np.full((15, 15, 3), 20, dtype=np.uint8)
source[5:10, 6:9] = 230
source[6:9, 7] = 180
mask = module.source_silhouette_mask(source, (4, 4, 11, 11))
assert mask.dtype == np.uint8
assert set(np.unique(mask)) <= {0, 255}
assert mask[7, 7] == 255
assert mask[4, 4] == 0
def test_rotated_source_silhouette_excludes_axis_aligned_corners() -> None:
source = np.full((80, 160, 3), 20, dtype=np.uint8)
source[10:70, 10:150] = 230
mask = module.source_silhouette_mask(source, (0, 0, 160, 80), angle=12)
assert mask[0, 0] == 0
assert mask[79, 159] == 0
def test_source_box_mask_pads_and_clips_boxes() -> None:
mask = module.source_box_mask((20, 30), [(1, 2, 11, 10), (25, 15, 30, 20)])
assert mask.shape == (20, 30)
assert mask[0, 0] == 255
assert mask[19, 29] == 255
assert mask[0, 22] == 0
def test_detect_line_boxes_fails_closed_on_count_mismatch() -> None:
class Engine:
def predict(self, _image):
return [{"rec_scores": [0.9], "rec_boxes": [[10, 10, 30, 30]]}]
with pytest.raises(module.click.ClickException, match="detected 1 source lines; expected exactly 2"):
module.detect_line_boxes(Engine(), np.zeros((50, 50, 3), dtype=np.uint8), expected_count=2)
def test_residual_mask_is_limited_to_original_glyph_positions(monkeypatch) -> None:
from remove_ai_watermarks._internal import text_restoration
background = np.zeros((8, 8, 3), dtype=np.uint8)
original = np.zeros((8, 8), dtype=np.uint8)
original[3, 3] = 255
detected = np.zeros((8, 8), dtype=np.uint8)
detected[3, 3] = 255
detected[6, 6] = 255
monkeypatch.setattr(text_restoration, "_foreground_mask", lambda _image, _box: detected)
residual = module.residual_glyph_mask(background, original, (0, 0, 8, 8))
assert residual[3, 3] == 255
assert residual[6, 6] == 0
+128
View File
@@ -0,0 +1,128 @@
"""Verified-text manifest and compositor tests without model downloads."""
from __future__ import annotations
import json
import numpy as np
import pytest
from PIL import Image, PngImagePlugin
from remove_ai_watermarks._internal.text_restoration import (
FIDELITY_BLEND_ALPHA,
VerifiedTextLine,
blend_fidelity_anchor,
load_verified_text_manifest,
restore_verified_text,
source_pixel_sha256,
)
def _manifest(image: Image.Image) -> dict[str, object]:
return {
"schema_version": 1,
"verified": True,
"source_pixel_sha256": source_pixel_sha256(image),
"width": image.width,
"height": image.height,
"lines": [
{
"box": [8, 8, 40, 24],
"text": "Exact text",
"script": "alphabetic",
"angle": 0.0,
}
],
}
def test_pixel_hash_ignores_container_metadata(tmp_path) -> None:
image = Image.new("RGB", (48, 32), (10, 20, 30))
plain = tmp_path / "plain.png"
tagged = tmp_path / "tagged.png"
image.save(plain)
metadata = PngImagePlugin.PngInfo()
metadata.add_text("note", "different container bytes")
image.save(tagged, pnginfo=metadata)
with Image.open(plain) as left, Image.open(tagged) as right:
assert plain.read_bytes() != tagged.read_bytes()
assert source_pixel_sha256(left) == source_pixel_sha256(right)
def test_verified_manifest_is_bound_to_source_pixels(tmp_path) -> None:
source = Image.new("RGB", (48, 32), (10, 20, 30))
path = tmp_path / "lines.json"
path.write_text(json.dumps(_manifest(source)), encoding="utf-8")
loaded = load_verified_text_manifest(path, source)
assert loaded.width == 48
assert loaded.height == 32
assert loaded.lines == (VerifiedTextLine((8, 8, 40, 24), "Exact text", "alphabetic", 0.0),)
@pytest.mark.parametrize(
("mutation", "message"),
[
({"verified": False}, "verified=true"),
({"source_pixel_sha256": "0" * 64}, "does not match"),
({"width": 49}, "dimensions"),
({"lines": []}, "non-empty"),
],
)
def test_manifest_rejects_unverified_or_unbound_input(tmp_path, mutation, message) -> None:
source = Image.new("RGB", (48, 32), (10, 20, 30))
payload = _manifest(source)
payload.update(mutation)
path = tmp_path / "lines.json"
path.write_text(json.dumps(payload), encoding="utf-8")
with pytest.raises(ValueError, match=message):
load_verified_text_manifest(path, source)
def test_fidelity_anchor_uses_the_calibrated_rounding() -> None:
clean = Image.fromarray(np.array([[[1, 2, 3], [100, 150, 200]]], dtype=np.uint8))
donor = Image.fromarray(np.array([[[255, 254, 253], [200, 100, 50]]], dtype=np.uint8))
result = np.asarray(blend_fidelity_anchor(clean, donor))
expected = np.rint(
np.asarray(clean, dtype=np.float32) * (1.0 - FIDELITY_BLEND_ALPHA)
+ np.asarray(donor, dtype=np.float32) * FIDELITY_BLEND_ALPHA
).astype(np.uint8)
assert np.array_equal(result, expected)
def test_restoration_uses_lama_and_qwen_vae_core(monkeypatch) -> None:
from remove_ai_watermarks import region_eraser
source = np.full((40, 64, 3), 20, dtype=np.uint8)
source[12:24, 12:44] = 235
candidate = np.full_like(source, 30)
candidate[12:24, 12:44] = 150
donor = np.full_like(source, 40)
donor[12:24, 12:44] = (210, 220, 230)
calls: list[np.ndarray] = []
def fake_erase(image_bgr, mask):
calls.append(mask.copy())
output = image_bgr.copy()
output[mask > 0] = (30, 30, 30)
return output
monkeypatch.setattr(region_eraser, "lama_available", lambda: True)
monkeypatch.setattr(region_eraser, "erase_lama", fake_erase)
result = restore_verified_text(
Image.fromarray(source),
Image.fromarray(candidate),
Image.fromarray(donor),
(VerifiedTextLine((8, 8, 48, 28), "Exact text", "alphabetic"),),
)
restored = np.asarray(result)
assert calls
assert np.all(restored[16, 20] == donor[16, 20])
assert np.all(restored[0, 0] == candidate[0, 0])
Generated
+75 -59
View File
@@ -597,7 +597,7 @@ name = "cuda-bindings"
version = "13.3.1"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "cuda-pathfinder" },
{ name = "cuda-pathfinder", marker = "sys_platform != 'darwin' and sys_platform != 'emscripten' and sys_platform != 'win32'" },
]
wheels = [
{ url = "https://files.pythonhosted.org/packages/51/6b/457ca12dad3ee9bfcc9a545cfd6b64b359ba49de40f776f6e028e678f262/cuda_bindings-13.3.1-cp311-cp311-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c5879712accf6e14bb01aa5e67440eb84998b8d104b509cc7a6dc0b8f656a474", size = 6053539, upload-time = "2026-05-29T23:11:43.19Z" },
@@ -630,43 +630,43 @@ wheels = [
[package.optional-dependencies]
cublas = [
{ name = "nvidia-cublas", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cuda-nvrtc", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cublas", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
{ name = "nvidia-cuda-nvrtc", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
cudart = [
{ name = "nvidia-cuda-runtime", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cuda-runtime", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
cufft = [
{ name = "nvidia-cufft", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-nvjitlink", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cufft", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
{ name = "nvidia-nvjitlink", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
cufile = [
{ name = "nvidia-cufile", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cufile", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
cupti = [
{ name = "nvidia-cuda-cupti", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cuda-cupti", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
curand = [
{ name = "nvidia-curand", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-curand", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
cusolver = [
{ name = "nvidia-cublas", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cusolver", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cusparse", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-nvjitlink", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cublas", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
{ name = "nvidia-cusolver", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
{ name = "nvidia-cusparse", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
{ name = "nvidia-nvjitlink", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
cusparse = [
{ name = "nvidia-cusparse", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-nvjitlink", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cusparse", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
{ name = "nvidia-nvjitlink", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
nvjitlink = [
{ name = "nvidia-nvjitlink", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-nvjitlink", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
nvrtc = [
{ name = "nvidia-cuda-nvrtc", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-cuda-nvrtc", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
nvtx = [
{ name = "nvidia-nvtx", marker = "platform_machine == 'aarch64' or platform_machine == 'x86_64'" },
{ name = "nvidia-nvtx", marker = "(platform_machine == 'aarch64' and sys_platform == 'linux') or (platform_machine == 'x86_64' and sys_platform == 'linux')" },
]
[[package]]
@@ -1094,15 +1094,15 @@ name = "lightning"
version = "2.6.5"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "fsspec", extra = ["http"] },
{ name = "lightning-utilities" },
{ name = "packaging" },
{ name = "pytorch-lightning" },
{ name = "pyyaml" },
{ name = "torch" },
{ name = "torchmetrics" },
{ name = "tqdm" },
{ name = "typing-extensions" },
{ name = "fsspec", extra = ["http"], marker = "python_full_version < '3.13'" },
{ name = "lightning-utilities", marker = "python_full_version < '3.13'" },
{ name = "packaging", marker = "python_full_version < '3.13'" },
{ name = "pytorch-lightning", marker = "python_full_version < '3.13'" },
{ name = "pyyaml", marker = "python_full_version < '3.13'" },
{ name = "torch", marker = "python_full_version < '3.13'" },
{ name = "torchmetrics", marker = "python_full_version < '3.13'" },
{ name = "tqdm", marker = "python_full_version < '3.13'" },
{ name = "typing-extensions", marker = "python_full_version < '3.13'" },
]
sdist = { url = "https://files.pythonhosted.org/packages/c9/1d/83be8536bec71a0173e762a9a1fd92a24a5ad0d0f74c59550c3c4e6c103b/lightning-2.6.5.tar.gz", hash = "sha256:16a30310ed69afde3748491feb5d13508908effd70390d2bfc203dc0812a4b4a", size = 659201, upload-time = "2026-05-27T14:33:41.806Z" }
wheels = [
@@ -1114,8 +1114,8 @@ name = "lightning-utilities"
version = "0.15.3"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "packaging" },
{ name = "typing-extensions" },
{ name = "packaging", marker = "python_full_version < '3.13'" },
{ name = "typing-extensions", marker = "python_full_version < '3.13'" },
]
sdist = { url = "https://files.pythonhosted.org/packages/f1/45/7fa8f56b17dc0f0a41ec70dd307ecd6787254483549843bef4c30ab5adce/lightning_utilities-0.15.3.tar.gz", hash = "sha256:792ae0204c79f6859721ac7f386c237a33b0ed06ba775009cb894e010a842033", size = 33553, upload-time = "2026-02-22T14:48:53.348Z" }
wheels = [
@@ -1518,7 +1518,7 @@ name = "nvidia-cublas"
version = "13.1.1.3"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "nvidia-cuda-nvrtc" },
{ name = "nvidia-cuda-nvrtc", marker = "sys_platform != 'darwin' and sys_platform != 'emscripten' and sys_platform != 'win32'" },
]
wheels = [
{ url = "https://files.pythonhosted.org/packages/a7/a1/0bd24ee8c8d03adac032fd2909426a00c88f8c57961b1277ded97f91119f/nvidia_cublas-13.1.1.3-py3-none-manylinux_2_27_aarch64.whl", hash = "sha256:b7a210458267ac818974c53038fbec2e969d5c99f305ab15c72522fa9f001dd5", size = 542848918, upload-time = "2026-04-08T18:46:22.985Z" },
@@ -1557,7 +1557,7 @@ name = "nvidia-cudnn-cu13"
version = "9.20.0.48"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "nvidia-cublas" },
{ name = "nvidia-cublas", marker = "sys_platform != 'darwin' and sys_platform != 'emscripten' and sys_platform != 'win32'" },
]
wheels = [
{ url = "https://files.pythonhosted.org/packages/56/c5/83384d846b2fd17c44bd499b36c75a45ed4f095fbbb2252294e89cea5c5c/nvidia_cudnn_cu13-9.20.0.48-py3-none-manylinux_2_27_aarch64.whl", hash = "sha256:e31454ae00094b0c55319d9d15b6fa2fc50a9e1c0f5c8c80fb75258234e731e1", size = 444574296, upload-time = "2026-03-09T19:28:27.751Z" },
@@ -1569,7 +1569,7 @@ name = "nvidia-cufft"
version = "12.0.0.61"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "nvidia-nvjitlink" },
{ name = "nvidia-nvjitlink", marker = "sys_platform != 'darwin' and sys_platform != 'emscripten' and sys_platform != 'win32'" },
]
wheels = [
{ url = "https://files.pythonhosted.org/packages/8b/ae/f417a75c0259e85c1d2f83ca4e960289a5f814ed0cea74d18c353d3e989d/nvidia_cufft-12.0.0.61-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:2708c852ef8cd89d1d2068bdbece0aa188813a0c934db3779b9b1faa8442e5f5", size = 214053554, upload-time = "2025-09-04T08:31:38.196Z" },
@@ -1599,9 +1599,9 @@ name = "nvidia-cusolver"
version = "12.0.4.66"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "nvidia-cublas" },
{ name = "nvidia-cusparse" },
{ name = "nvidia-nvjitlink" },
{ name = "nvidia-cublas", marker = "sys_platform != 'darwin' and sys_platform != 'emscripten' and sys_platform != 'win32'" },
{ name = "nvidia-cusparse", marker = "sys_platform != 'darwin' and sys_platform != 'emscripten' and sys_platform != 'win32'" },
{ name = "nvidia-nvjitlink", marker = "sys_platform != 'darwin' and sys_platform != 'emscripten' and sys_platform != 'win32'" },
]
wheels = [
{ url = "https://files.pythonhosted.org/packages/c8/c3/b30c9e935fc01e3da443ec0116ed1b2a009bb867f5324d3f2d7e533e776b/nvidia_cusolver-12.0.4.66-py3-none-manylinux_2_27_aarch64.whl", hash = "sha256:02c2457eaa9e39de20f880f4bd8820e6a1cfb9f9a34f820eb12a155aa5bc92d2", size = 223467760, upload-time = "2025-09-04T08:33:04.222Z" },
@@ -1613,7 +1613,7 @@ name = "nvidia-cusparse"
version = "12.6.3.3"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "nvidia-nvjitlink" },
{ name = "nvidia-nvjitlink", marker = "sys_platform != 'darwin' and sys_platform != 'emscripten' and sys_platform != 'win32'" },
]
wheels = [
{ url = "https://files.pythonhosted.org/packages/f8/94/5c26f33738ae35276672f12615a64bd008ed5be6d1ebcb23579285d960a9/nvidia_cusparse-12.6.3.3-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:80bcc4662f23f1054ee334a15c72b8940402975e0eab63178fc7e670aa59472c", size = 162155568, upload-time = "2025-09-04T08:33:42.864Z" },
@@ -1670,8 +1670,8 @@ name = "omegaconf"
version = "2.3.1"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "antlr4-python3-runtime" },
{ name = "pyyaml" },
{ name = "antlr4-python3-runtime", marker = "python_full_version < '3.13'" },
{ name = "pyyaml", marker = "python_full_version < '3.13'" },
]
sdist = { url = "https://files.pythonhosted.org/packages/ce/3d/e4b57b8d9008c6ebe0d5eff901f91d5700cf7bdb8c8863df817463a7fd5e/omegaconf-2.3.1.tar.gz", hash = "sha256:e5e7de64aeebeddaf8e6d3f7a783b32ac2a01c0fbd9c878012caecb891a1f42a", size = 3298472, upload-time = "2026-06-11T05:05:12.885Z" }
wheels = [
@@ -2273,14 +2273,14 @@ name = "pytorch-lightning"
version = "2.6.5"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "fsspec", extra = ["http"] },
{ name = "lightning-utilities" },
{ name = "packaging" },
{ name = "pyyaml" },
{ name = "torch" },
{ name = "torchmetrics" },
{ name = "tqdm" },
{ name = "typing-extensions" },
{ name = "fsspec", extra = ["http"], marker = "python_full_version < '3.13'" },
{ name = "lightning-utilities", marker = "python_full_version < '3.13'" },
{ name = "packaging", marker = "python_full_version < '3.13'" },
{ name = "pyyaml", marker = "python_full_version < '3.13'" },
{ name = "torch", marker = "python_full_version < '3.13'" },
{ name = "torchmetrics", marker = "python_full_version < '3.13'" },
{ name = "tqdm", marker = "python_full_version < '3.13'" },
{ name = "typing-extensions", marker = "python_full_version < '3.13'" },
]
sdist = { url = "https://files.pythonhosted.org/packages/52/2c/8e73a3929b4c4bd600cafd38a97aaf7242a8cf518fb9f33d27c274ec898f/pytorch_lightning-2.6.5.tar.gz", hash = "sha256:1c32cefa76a1a9c4c5250338272d961d1e48b180e68396849efe128538ddb28e", size = 661673, upload-time = "2026-05-27T14:33:41.961Z" }
wheels = [
@@ -2604,6 +2604,21 @@ qwen-zimage = [
{ name = "torchvision" },
{ name = "transformers" },
]
text-restoration = [
{ name = "accelerate" },
{ name = "diffsynth" },
{ name = "diffusers" },
{ name = "huggingface-hub" },
{ name = "numpy", version = "1.26.4", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.13'" },
{ name = "numpy", version = "2.5.2", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.13'" },
{ name = "onnxruntime" },
{ name = "opencv-python-headless" },
{ name = "safetensors" },
{ name = "tokenizers" },
{ name = "torch" },
{ name = "torchvision" },
{ name = "transformers" },
]
trustmark = [
{ name = "trustmark", marker = "python_full_version < '3.13'" },
]
@@ -2650,8 +2665,9 @@ requires-dist = [
{ name = "remove-ai-watermarks", extras = ["pixels"], marker = "extra == 'detect'" },
{ name = "remove-ai-watermarks", extras = ["pixels"], marker = "extra == 'diffusion'" },
{ name = "remove-ai-watermarks", extras = ["pixels"], marker = "extra == 'visible'" },
{ name = "remove-ai-watermarks", extras = ["qwen-zimage", "lama"], marker = "extra == 'text-restoration'" },
{ name = "remove-ai-watermarks", extras = ["video"], marker = "extra == 'dev'" },
{ name = "remove-ai-watermarks", extras = ["video", "heif", "detect", "trustmark", "qwen-zimage", "lama", "migan"], marker = "extra == 'all'" },
{ name = "remove-ai-watermarks", extras = ["video", "heif", "detect", "trustmark", "text-restoration", "migan"], marker = "extra == 'all'" },
{ name = "remove-ai-watermarks", extras = ["visible"], marker = "extra == 'lama'" },
{ name = "remove-ai-watermarks", extras = ["visible"], marker = "extra == 'migan'" },
{ name = "remove-ai-watermarks", extras = ["visible"], marker = "extra == 'video'" },
@@ -2663,7 +2679,7 @@ requires-dist = [
{ name = "transformers", marker = "extra == 'diffusion'", specifier = ">=5,<6" },
{ name = "trustmark", marker = "python_full_version < '3.13' and extra == 'trustmark'", specifier = ">=0.8.0" },
]
provides-extras = ["pixels", "heif", "visible", "video", "detect", "diffusion", "qwen-zimage", "trustmark", "lama", "migan", "dev", "all"]
provides-extras = ["pixels", "heif", "visible", "video", "detect", "diffusion", "qwen-zimage", "text-restoration", "trustmark", "lama", "migan", "dev", "all"]
[[package]]
name = "requests"
@@ -2968,10 +2984,10 @@ name = "torchmetrics"
version = "1.9.0"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "lightning-utilities" },
{ name = "numpy", version = "1.26.4", source = { registry = "https://pypi.org/simple" } },
{ name = "packaging" },
{ name = "torch" },
{ name = "lightning-utilities", marker = "python_full_version < '3.13'" },
{ name = "numpy", version = "1.26.4", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.13'" },
{ name = "packaging", marker = "python_full_version < '3.13'" },
{ name = "torch", marker = "python_full_version < '3.13'" },
]
sdist = { url = "https://files.pythonhosted.org/packages/81/34/39b8b749333db56c0585d7a11fa62a283c087bb1dfc897d69fb8cedbefb1/torchmetrics-1.9.0.tar.gz", hash = "sha256:a488609948600df52d3db4fcdab02e62aab2a85ef34da67037dc3e65b8512faa", size = 581765, upload-time = "2026-03-09T17:41:22.443Z" }
wheels = [
@@ -3066,13 +3082,13 @@ name = "trustmark"
version = "0.9.1"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "einops" },
{ name = "lightning" },
{ name = "numpy", version = "1.26.4", source = { registry = "https://pypi.org/simple" } },
{ name = "omegaconf" },
{ name = "six" },
{ name = "torch" },
{ name = "torchvision" },
{ name = "einops", marker = "python_full_version < '3.13'" },
{ name = "lightning", marker = "python_full_version < '3.13'" },
{ name = "numpy", version = "1.26.4", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.13'" },
{ name = "omegaconf", marker = "python_full_version < '3.13'" },
{ name = "six", marker = "python_full_version < '3.13'" },
{ name = "torch", marker = "python_full_version < '3.13'" },
{ name = "torchvision", marker = "python_full_version < '3.13'" },
]
sdist = { url = "https://files.pythonhosted.org/packages/87/0a/0a4232030c6a62d12b6a02ae73bdce6e99c8532bc8f05a5a2e6ce103da82/trustmark-0.9.1.tar.gz", hash = "sha256:dc79e3fb070f5d94765acf8868a51f50a612cc05b53223cf1e6b605d4ff1e0ae", size = 63949, upload-time = "2026-04-09T08:59:52.472Z" }