Add verified text restoration

This commit is contained in:
Victor Kuznetsov
2026-08-15 12:25:35 -07:00
parent 8c00525946
commit 432b63b6d7
21 changed files with 974 additions and 255 deletions
+31
View File
@@ -398,6 +398,37 @@ schedule, CFG 1.0 and CUDA, so every one of those flags existed only to be refus
several layers down. They are not parsed at all now, which fails at the point the
user can act on rather than after a model load.
### Restore operator-verified text
`--text-manifest` enables the experimental `vae-glyphs` post-pass. It reconstructs
the source with the Qwen VAE, blends 15% of that reconstruction into the normal
`qwen-zimage` result, erases the annotated candidate glyphs with LaMa, and composites
only the reconstructed glyph cores through source-derived silhouettes. It does not
run OCR or choose which strings are correct.
Install the combined extra and run only with a manually reviewed manifest:
```bash
uv tool install --force "remove-ai-watermarks[text-restoration]"
remove-ai-watermarks invisible image.png -o clean.png \
--pipeline qwen-zimage --text-manifest verified-lines.json --force
```
The manifest is a JSON object with `schema_version: 1`, `verified: true`, decoded
RGB dimensions, `source_pixel_sha256`, and a non-empty `lines` array. Each line has
an integer `[x1, y1, x2, y2]` box, exact `text`, a non-empty `script`, and an optional
angle from -30 to 30 degrees. Lines must be in top-to-bottom, left-to-right order.
The hash binds the annotations to decoded RGB geometry and pixels, so metadata-only
container changes remain valid while a resized or edited source fails closed. The
experimental helper
`remove_ai_watermarks._internal.text_restoration.source_pixel_sha256` computes it.
This mode is supported only by `qwen-zimage` at native untiled geometry with
`humanize=0`, `unsharp=0`, and adaptive polish disabled. `all` also accepts the flag,
but its manifest must match the pixels entering the invisible stage; if visible-mark
removal changes those pixels, the hash check rejects the run. One oracle verdict does
not certify another manifest, seed, model/runtime version, or output hash.
### Work with limited memory
Lower CUDA memory pressure:
+12
View File
@@ -87,6 +87,15 @@ removal, metadata stripping and every `identify` command still run anywhere.
Video SynthID regeneration is a separate VAE path and does still run on CPU or MPS;
it needs the `diffusion` extra, not this one.
The experimental verified-text post-pass additionally needs LaMa:
```bash
uv tool install --force "remove-ai-watermarks[text-restoration]"
```
That extra includes `qwen-zimage` and `lama`; it does not add OCR. Text strings and
line boxes must be reviewed before the run.
## Feature extras
Extras are composable. Install only the capabilities and file formats the
@@ -104,6 +113,7 @@ application actually uses:
| `migan` | MI-GAN ONNX fill backend | `visible`, ONNX Runtime | Model download, no Torch |
| `lama` | big-LaMa ONNX fill backend | `visible`, ONNX Runtime | Model download, no Torch |
| `qwen-zimage` | Invisible image-watermark removal, both CUDA-only profiles | `diffusion`, DiffSynth | Yes |
| `text-restoration` | Opt-in verified Qwen-VAE glyph restoration | `qwen-zimage`, `lama` | Yes |
| `all` | Every production feature available on the active Python | All compatible rows above | Yes |
| `dev` | Tests, linting, typing, and upstream parity checks | `video`, `detect`, upstream invisible-watermark | Yes, for parity tests |
@@ -118,6 +128,8 @@ flowchart LR
migan --> visible
lama --> visible
qwen["qwen-zimage"] --> diffusion
text["text-restoration"] --> qwen
text --> lama
heif
trustmark
```
+11 -10
View File
@@ -69,15 +69,15 @@ difficult faces. The measurements and their OCR and oracle caveats are tracked
in [`data/evaluations/fidelity/`](../data/evaluations/fidelity/README.md).
A global Z-Image Turbo prototype preserved text substantially better at low
strength, but it has no useful cross-provider operating point and is not a
supported profile. The evaluated text restorers also remain research-only:
fresh-font and silhouette variants visibly changed typography, while the
higher-fidelity `vae-glyphs` route still requires verified strings, line
geometry, a separately generated donor, and an independently clean global
anchor. Automatic OCR and line-box proposals are not reliable enough to remove
those requirements, and the exact oracle results do not establish a general
mask, seed, or provider operating range. Qwen-Image-2.0 is hosted-only and
exposes no equivalent low-strength denoise control. Exact experiments, controls,
and pass rates are kept in
supported profile. Automatic text restorers also remain research-only:
fresh-font and silhouette variants visibly changed typography. The higher-fidelity
`vae-glyphs` route is available only as an experimental opt-in with verified strings
and line geometry. It builds its donor internally but still requires an independently
clean global anchor. Automatic OCR and line-box proposals are not reliable enough to
remove those requirements, and exact oracle results do not establish a general mask,
seed, runtime, or provider operating range. Qwen-Image-2.0 is hosted-only and exposes
no equivalent low-strength denoise control. Exact experiments, controls, and pass
rates are kept in
[`text-protection-research.md`](text-protection-research.md) and the
[`fidelity` evaluation record](../data/evaluations/fidelity/README.md).
@@ -192,7 +192,8 @@ certified at a fixed seed. The live resolver is
| `qwen-zimage` | CUDA only, large model stack, and limited broad certification across seeds and content. |
| `sdxl-zimage` | CUDA only. Its strength ladder is flat per vendor, not a resolution curve, because flat values are what was measured. |
The evaluated text-restoration prototypes are not optional production stages.
Only manually verified `vae-glyphs` is an optional production stage, and it is
experimental rather than a default.
OCR plus LaMa recovered literal poster text but changed fonts and worsened whole-image
fidelity. Restricting it to OCR-mismatched lines improved the tradeoff but still
left a local shadow on one poster. The published AnyText2 SD1.5 checkpoint
+23
View File
@@ -949,6 +949,29 @@ orchestration, YuNet integration, SAM selection, masks, sizing helpers, and pixe
compositing are implemented for this runtime. Changing a calibrated model input
requires the same provider-oracle and identity evaluation as a model change.
#### Verified text restoration
[`_internal/text_restoration.py`](../src/remove_ai_watermarks/_internal/text_restoration.py)
implements the opt-in `vae-glyphs` stage. A versioned manifest carries manually
reviewed strings and source-space line boxes, plus a SHA-256 over decoded RGB width,
height, and pixels. Validation happens before model loading. The product never treats
OCR confidence as verification.
When enabled, `QwenZImagePipeline` reconstructs the source once through its already
loaded Qwen VAE, runs the ordinary global and face stages, blends 15% of the VAE
reconstruction into that clean result, and calls the shared restoration compositor.
The compositor derives binary source and candidate silhouettes, groups nearby lines,
uses LaMa for the initial and residual-glyph erase passes, paints fresh silhouette
edges, then copies the Qwen-VAE core with a 0.5-pixel feather. The evaluation script
imports these same mask and compositing helpers so the two implementations cannot
silently drift.
The stage is deliberately narrower than the engine: it rejects `sdxl-zimage`, tiles,
resolution caps, humanize, unsharp, and adaptive polish. Those combinations change
geometry or final pixels after the verified layer and have no measured oracle result.
It remains opt-in because annotations are manual and provider verdicts apply only to
the exact tested output hashes, not to the mechanism in general.
A matched stage-isolation check on the 18-face Gemini portrait grid confirms the
division of responsibility. The visible-cleaned, metadata-stripped control and the
Z-Image face-only output were both SynthID-positive; Qwen global-only and the full
+17
View File
@@ -596,6 +596,23 @@ engine = InvisibleEngine(pipeline="sdxl-zimage")
The `qwen-zimage` extra is required for both profiles: each runs the same
DiffSynth Z-Image face stage.
The opt-in verified-text stage uses the same `text_manifest` argument as the CLI:
```python
engine.remove_watermark(
Path("watermarked.png"),
Path("clean.png"),
text_manifest=Path("verified-lines.json"),
)
```
Install `remove-ai-watermarks[text-restoration]`. The manifest schema and safety
constraints are documented in the CLI guide. The engine verifies its decoded RGB
hash before loading the diffusion models and rejects SDXL, tiling, downscaling, and
postprocessing combinations that were not evaluated. `InvisibleOptions` exposes the
same field for `remove_all`; after a visible-stage edit, the manifest must be built
against the staged pixels rather than the pristine source.
`remove_watermark` takes strength, seed, tiling, resolution, and postprocessing
controls. It takes no model id, step count or guidance scale, and neither does the
constructor: each profile pins its model stack, its per-stage schedule and CFG