Refactor watermark detection and provenance handling

This commit is contained in:
Victor Kuznetsov
2026-07-16 17:40:46 -07:00
parent 9618ac93c8
commit a8f3536d3e
12 changed files with 654 additions and 293 deletions
+95
View File
File diff suppressed because one or more lines are too long
+2 -2
View File
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -27,7 +27,7 @@ It does **not** target watermarks that protect someone else's paid or copyrighte
## Features
- **Visible watermark removal** — a registry of known marks in their usual places: the Gemini / Nano Banana sparkle, the Doubao "豆包AI生成" text strip, the Jimeng "★ 即梦AI" wordmark, and the Samsung Galaxy AI "✦ Contenuti generati dall'AI" strip (bottom-left, locale-specific). Each mark is **localized to a footprint mask, then filled**: the engine finds the mark, builds a binary mask over its footprint, and one shared, swappable fill inpaints that region. Choose the fill with `--backend`: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN, light, the memory-tight pick where LaMa will not fit), or `lama` (big-LaMa, best quality, heavier, auto-preferred when a learned backend is available); the default `auto` uses LaMa > MI-GAN > cv2, best available. The localizer is cheap CPU (cv2/numpy), so it runs anywhere; the heavier MI-GAN/LaMa fill is opt-in. Detection keys on each mark's own shape (NCC against a captured silhouette; the alpha captures rebuilt by `scripts/visible_alpha_solve.py` are used to detect and to shape the mask, not for pixel recovery). The visual detector needs no metadata, but a borderline (faint or moved) mark is only trusted with corroboration: `--sensitivity` (default `auto`) relaxes a mark's gate when local metadata confirms the vendor or a same-product sibling mark is found; `--sensitivity assume-ai` relaxes every mark on the assertion that the image is AI, recovering the moved or re-rendered marks the conservative gate skips on a metadata-stripped screenshot (`strict` never relaxes). `visible --mark auto` finds and removes every detected mark in one pass. Fast, offline, no GPU. (For arbitrary logos/objects, see `erase`.)
- **Visible watermark removal** — a registry of known marks in their usual places: the Gemini / Nano Banana sparkle, the Doubao "豆包AI生成" text strip, the Jimeng "★ 即梦AI" wordmark, and the Samsung Galaxy AI "✦ Contenuti generati dall'AI" strip (bottom-left, locale-specific). Each mark is **localized to a footprint mask, then filled**: the engine finds the mark, builds a binary mask over its footprint, and one shared, swappable fill inpaints that region. Choose the fill with `--backend`: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN, light, the memory-tight pick where LaMa will not fit), or `lama` (big-LaMa, best quality, heavier, auto-preferred when a learned backend is available); the default `auto` uses LaMa > MI-GAN > cv2, best available. The localizer is cheap CPU (cv2/numpy), so it runs anywhere; the heavier MI-GAN/LaMa fill is opt-in. Detection keys on each mark's own shape (NCC against a captured silhouette; the alpha captures rebuilt by `scripts/visible_alpha_solve.py` are used to detect and to shape the mask, not for pixel recovery). The visual detector needs no metadata, but a borderline (faint or moved) mark is only trusted with corroboration: `--sensitivity` (default `auto`) relaxes a mark's gate when local metadata confirms the vendor or a same-product sibling mark is found; `--sensitivity assume-ai` relaxes every mark on the assertion that the image is AI, recovering the moved or re-rendered marks the conservative gate skips on a metadata-stripped screenshot (`strict` never relaxes). Because asserting "this is AI" says nothing about *which* vendor made it, a mark relaxed on that assertion alone still has to clear a confidence floor, so a clean photo is left untouched. `visible --mark auto` finds and removes every detected mark in one pass. Fast, offline, no GPU. (For arbitrary logos/objects, see `erase`.)
- **Universal region eraser (`erase`)** — remove any logo / watermark / object inside boxes you specify, regardless of position or color. Default cv2 inpainting (CPU, instant); optional big-LaMa via onnxruntime (`lama` extra) for higher quality
- **Invisible watermark removal** — SynthID, StableSignature, TreeRing via diffusion-based regeneration (needs a local GPU, or run it with no setup on [raiw.cc](https://raiw.cc))
- **AI metadata stripping** — EXIF, PNG text chunks, C2PA provenance manifests (PNG / JPEG / AVIF / HEIF / JPEG-XL, **MP4 / MOV / M4V / M4A** at the container level, and **WebM / MP3 / WAV / FLAC / OGG** losslessly via ffmpeg), XMP DigitalSourceType
+14 -3
View File
@@ -61,11 +61,22 @@ module.
`watermark_registry.py`**single catalog of known visible watermarks**, the unified "find known marks in their usual places, recognize, remove" entry.
**Localize -> fill by policy (replaced reverse-alpha):** each mark is localized to a binary full-frame footprint mask (a `Localization`), and one shared, swappable fill inpaints that mask via `fill(image, mask, backend=...)` (delegates to `region_eraser.erase`). This replaced the old reverse-alpha removal (invert a captured alpha map, `original = (wm - a*logo)/(1-a)`, plus a thin residual inpaint) for ALL marks — gemini, doubao, jimeng, samsung, and jimeng_pill. **Why it changed:** reverse-alpha depended on a fixed captured alpha map at a fixed position, so it broke whenever a vendor moved or re-rendered its mark; and it was not color-lossless even with the right map (it amplifies 8-bit quantization and JPEG-chroma error by `1/(1-a)`), which showed up as "the color just changed, not removed" reports. Localize -> fill has a benign failure mode: a slightly-off localization just inpaints a small region near-losslessly instead of leaving a color-shifted smear. The captured alpha maps are still used to DETECT the marks and to shape the mask (gemini's footprint), but NOT for pixel recovery. Fill backends: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN ONNX, light, the memory-tight pick where LaMa will not fit), `lama` (big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available); `auto` = LaMa > MI-GAN > cv2, best available. Each `KnownMark` ties a key to {usual `location`, `in_auto` flag, a `_detect` callable → uniform `MarkDetection`, a `_mask` callable → full-frame footprint mask}; `KnownMark.remove(image, *, backend="auto", provenance=False, force=False)`. Entries today: `gemini` (bottom-right sparkle), `doubao` (bottom-right "豆包AI生成"), `jimeng` (bottom-right "★ 即梦AI"), `samsung` (bottom-**LEFT** "✦ Contenuti generati dall'AI", Samsung Galaxy AI, Italian locale), and the capture-less `jimeng_pill` (top-left "AI生成"). `detect_marks(image, *, provenance=frozenset())` scans all (strict, for the identify verdict); `remove_auto_marks(image, *, sensitivity="auto", provenance=frozenset(), backend="auto")` removes every detected mark in one pass. **Sensitivity (`auto`/`strict`/`assume_ai`)** decides how hard a borderline mark is trusted: the visual detectors are pixel-based (no metadata needed) and the recall gain comes from relaxing the false-positive gate, not from metadata. `resolve_relax` turns the policy + evidence into the per-mark relaxation boolean — `strict` never relaxes, `assume_ai` always relaxes (the caller asserts AI, e.g. a metadata-stripped screenshot; ~46% -> ~92% Gemini recall, at the cost of a small near-lossless fill on some clean corners), `auto` relaxes only on same-product evidence (metadata provenance for that vendor, or a confidently strict-detected sibling of the same `_PRODUCT_OF` — Doubao and Jimeng are both bottom-right ByteDance but distinct products, so they do NOT cross-relax). **Perception / decision / action are separated three ways** (the removal path only; `identify` keeps calling `KnownMark.detect` directly, so its verdict is untouched): `_build_candidates(image)` is PERCEPTION — it runs each detector at both trust levels and packages raw verdicts + the pill's flatness feature into `Candidate`s, no policy; `decide(candidates, Context(sensitivity, provenance)) -> [Decision]` is the pure DECISION arbiter — all keep/drop policy (`resolve_relax` cross-mark corroboration + the pill gate) in one image-free, unit-testable function (`tests/test_watermark_registry.py::TestArbiter`); then `remove_auto_marks` does the ACTION, localizing -> filling each winner. The Gemini FP gate deliberately stays inside `gemini_engine` (not the arbiter) because `identify` reads that same gated confidence — pulling it out would drift the provenance verdict. Behavior is byte-identical to the pre-arbiter two-pass (corpus-verified: strict/auto 46%, assume_ai 92% Gemini recall unchanged).
**Localize -> fill by policy (replaced reverse-alpha):** each mark is localized to a binary full-frame footprint mask (a `Localization`), and one shared, swappable fill inpaints that mask via `fill(image, mask, backend=...)` (delegates to `region_eraser.erase`). This replaced the old reverse-alpha removal (invert a captured alpha map, `original = (wm - a*logo)/(1-a)`, plus a thin residual inpaint) for ALL marks — gemini, doubao, jimeng, samsung, and jimeng_pill. **Why it changed:** reverse-alpha depended on a fixed captured alpha map at a fixed position, so it broke whenever a vendor moved or re-rendered its mark; and it was not color-lossless even with the right map (it amplifies 8-bit quantization and JPEG-chroma error by `1/(1-a)`), which showed up as "the color just changed, not removed" reports. Localize -> fill has a benign failure mode: a slightly-off localization just inpaints a small region near-losslessly instead of leaving a color-shifted smear. The captured alpha maps are still used to DETECT the marks and to shape the mask (gemini's footprint), but NOT for pixel recovery. Fill backends: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN ONNX, light, the memory-tight pick where LaMa will not fit), `lama` (big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available); `auto` = LaMa > MI-GAN > cv2, best available. Each `KnownMark` ties a key to {usual `location`, `in_auto` flag, a `_detect` callable → uniform `MarkDetection`, a `_mask` callable → full-frame footprint mask}; `KnownMark.remove(image, *, backend="auto", provenance=False, force=False)`. Entries today: `gemini` (bottom-right sparkle), `doubao` (bottom-right "豆包AI生成"), `jimeng` (bottom-right "★ 即梦AI"), `samsung` (bottom-**LEFT** "✦ Contenuti generati dall'AI", Samsung Galaxy AI, Italian locale), and the capture-less `jimeng_pill` (top-left "AI生成"). `detect_marks(image, *, provenance=frozenset())` scans all (strict, for the identify verdict); `remove_auto_marks(image, *, sensitivity="auto", provenance=frozenset(), backend="auto")` removes every detected mark in one pass. **Sensitivity (`auto`/`strict`/`assume_ai`)** decides how hard a borderline mark is trusted: the visual detectors are pixel-based (no metadata needed) and the recall gain comes from relaxing the false-positive gate, not from metadata. `resolve_trust` turns the policy + evidence into the per-mark trust level the engines consume as `provenance = level != "strict"``strict` never relaxes; `auto` relaxes only on same-product evidence (metadata provenance for that vendor, or a confidently strict-detected sibling of the same `_PRODUCT_OF` — Doubao and Jimeng are both bottom-right ByteDance but distinct products, so they do NOT cross-relax); `assume_ai` relaxes every mark (the caller asserts AI, e.g. a metadata-stripped screenshot). **Three levels, not two: `strict` / `assumed` / `confirmed`.** Relaxing bypasses the engine's false-positive gate outright, and that bypass is contracted to mean the vendor is CONFIRMED (`GeminiEngine.detect_watermark`'s `trust_provenance`: "external metadata already proves this is a Google generation"). An `assume_ai` caller asserts the image is AI, which says nothing about WHICH vendor, so a mark relaxed on assumption alone must also clear `_ASSUMED_CONF_FLOOR` (`assumed_floor_ok`; gemini 0.50) — see "Assumed-trust confidence floor" below. **Perception / decision / action are separated three ways** (the removal path only; `identify` keeps calling `KnownMark.detect` directly, so its verdict is untouched): `_build_candidates(image)` is PERCEPTION — it runs each detector at both trust levels and packages raw verdicts + the pill's flatness feature into `Candidate`s, no policy; `decide(candidates, Context(sensitivity, provenance)) -> [Decision]` is the pure DECISION arbiter — all keep/drop policy (`resolve_trust` cross-mark corroboration + the assumed-trust floor + the pill gate) in one image-free, unit-testable function (`tests/test_watermark_registry.py::TestArbiter`); then `remove_auto_marks` does the ACTION, localizing -> filling each winner. The Gemini FP gate deliberately stays inside `gemini_engine` (not the arbiter) because `identify` reads that same gated confidence — pulling it out would drift the provenance verdict. Behavior was byte-identical to the pre-arbiter two-pass when the arbiter landed; `assume_ai` has since gained the assumed-trust confidence floor (see below), which deliberately changes its verdict on weak gate-bypassed matches.
**Head-to-head validation (v0.12.1 reverse-alpha vs the current localize -> fill):** run over the full labelled visible-mark set, with the cv2 / MI-GAN / LaMa fills each compared against the old reverse-alpha. **doubao and jimeng are identical** across every backend -- 100% coverage and 100% clearance either way. **gemini** strict coverage is a few points below reverse-alpha's (the deliberate false-positive tightening), but the metadata-stripped faint ones are now mostly recovered by the DEFAULT white-core rescue in the FP gate (`gemini_engine`: a bright near-WHITE core distinguishes a real faint sparkle from a colored bright corner -- ~14/20 recovered at ~1.25% clean false-fire; a learned classifier on the same features measured worse, 2026-07 tier-1), the residual under `assume_ai`; clearance is equal (~98% both), and neither version touches pixels outside the mark box (outside-box PSNR ~99). **Clearance is fill-independent** -- cv2, MI-GAN and LaMa all strip the mark's shape equally, so the re-detect metric does not separate them; the difference is purely the *visual fill quality* on the recovered region, and it is background-dependent. reverse-alpha recovered textured and especially regular/structured backgrounds (a lattice, a grid) more cleanly than any inpaint; **LaMa closes most of that gap** (the best learned backend), **MI-GAN can ghost or hallucinate structure**, and **cv2 smears** (the last-resort floor). This is why `auto` resolves `LaMa > MI-GAN > cv2` (`preferred_inpaint_backend`) and warns once on the cv2 fallback; on flat backgrounds every backend is clean.
**Provenance prior:** when local metadata already confirms the vendor, the mark's detection trust gate is relaxed (a confirmed vendor means the mark is present with high prior, so a mark the conservative detector would demote as a content false positive is trusted). `detect_marks` / `remove_auto_marks` take a `provenance` frozenset and `KnownMark.remove` a `provenance` flag. Mapping: a Google/Gemini C2PA issuer relaxes gemini (skips its false-positive gate and lowers the trust threshold from 0.5 to 0.35); a China-AIGC (TC260) label relaxes doubao/jimeng; `samsung_genai` relaxes samsung. Corpus finding: on Google-C2PA images, Gemini sparkle recall rose from ~46% (plain detector) to ~90% with the provenance prior (recovering marks the vendor moved or re-rendered). The localizer is cheap CPU (cv2/numpy), so a memory-tight caller runs it anywhere; the heavy MI-GAN/LaMa fill is opt-in and chosen by the caller.
**Assumed-trust confidence floor (`_ASSUMED_CONF_FLOOR` / `assumed_floor_ok`, 2026-07-16):** relaxing a mark bypasses the engine's false-positive gate ENTIRELY, leaving only the bare detector threshold (gemini: fused confidence >= 0.35). That is defensible when metadata names the vendor (`confirmed`) and indefensible on a bare "assume this is AI" (`assumed`), because the assertion carries no vendor information. Measured on 256 genuine camera captures (Make/Model/exposure/aperture present, no AI token -- a Gemini sparkle cannot be there) vs 697 Google-C2PA positives with the metadata used only as a label:
| bypassed threshold | recall | false fire on clean photos |
|---|---|---|
| 0.35 (the bare gate) | 82.6% | **59.8%** |
| 0.45 | 66.6% | 12.5% |
| **0.50 (chosen)** | 59.4% | **0.0%** |
| strict gate | 56.4% | 0.0% |
0.35 sat on a cliff: +26pp recall over strict bought by filling a corner on ~6 of every 10 CLEAN photos, and `api.remove_visible(sensitivity="assume_ai")` reproduced it on 8/15 of the committed verified-clean negatives. The floor is applied in the arbiter and is **monotonic over strict** -- a mark the strict gate accepted is never dropped by it, so `assume_ai` only ever adds recall. End-to-end after the fix (400 Google-C2PA positives with metadata hidden from the detector, 256 camera negatives, through the public `api.remove_visible`): recall strict 55.0% / auto 55.2% / assume_ai 62.8%; false fire 0.0% / 0.0% / 2.3%. The residual 2.3% contains **no gemini at all** (doubao 2, jimeng_pill 2, jimeng 1, samsung 1 of 256) -- the text marks' own relaxed gates (<1% each) plus the pill's flat-footprint arm, all pre-existing and benign. The superseded "~46% -> ~92%" figure measured recall only, on Google-C2PA files where the answer was always Google; false fire on non-Google content was never measured. Marks other than gemini carry no floor because their bypassed false-fire is already under 1%. Regression: `tests/test_watermark_registry.py::TestArbiter::{test_assume_ai_drops_sparkle_below_the_assumed_floor, test_assume_ai_keeps_sparkle_confirmed_by_metadata_below_the_floor, test_assume_ai_is_monotonic_over_strict}`.
**Provenance prior:** when local metadata already confirms the vendor, the mark's detection trust gate is relaxed (a confirmed vendor means the mark is present with high prior, so a mark the conservative detector would demote as a content false positive is trusted). `detect_marks` / `remove_auto_marks` take a `provenance` frozenset and `KnownMark.remove` a `provenance` flag. Mapping: a Google/Gemini C2PA issuer relaxes gemini (skips its false-positive gate and lowers the trust threshold from 0.5 to 0.35); a China-AIGC (TC260) label relaxes doubao/jimeng; `samsung_genai` relaxes samsung. Corpus finding: on Google-C2PA images, Gemini sparkle recall rose from ~46% (plain detector) to ~90% with the provenance prior (recovering marks the vendor moved or re-rendered). That gain is why the bypass exists, and it is conditional on the metadata actually naming the vendor — a caller merely ASSUMING the image is AI does not get it unconditionally (see the assumed-trust confidence floor above). The localizer is cheap CPU (cv2/numpy), so a memory-tight caller runs it anywhere; the heavy MI-GAN/LaMa fill is opt-in and chosen by the caller.
**Cross-engine confidences aren't directly comparable**, so the gemini adapter applies the corpus-validated 0.5 sparkle threshold (`_GEMINI_AUTO_MIN_CONF`) for its `detected` flag (lowered to 0.35 under the Google/Gemini provenance prior) — otherwise the gemini engine's loose internal threshold weakly fires (~0.36) on the Doubao text and hijacks `auto`. The shape-keyed Doubao/Jimeng/Samsung NCC detectors don't cross-fire (jimeng scores ~0.22 on the Doubao strip, well under its 0.45 threshold; Samsung is bottom-left so it shares no corner with the others, and scored 0.0 on Doubao/Jimeng captures and they 0.0 on a real Samsung photo), so `auto` picks the right one. `cli.cmd_visible` is registry-driven: `--mark auto``remove_auto_marks` (removes every detected mark), `--mark <key>` → that mark; `--mark` choices come from `mark_keys()`.
@@ -230,7 +241,7 @@ Diffusion SynthID removal. The `--tile/--no-tile` knob is the *lossless* alterna
### `visible`
Known-visible-mark removal by **localize -> fill**: each detected mark is localized to a binary full-frame footprint mask, then one shared, swappable fill inpaints that mask. `--backend auto|cv2|migan|lama` (default `auto`) picks the fill: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN ONNX, the memory-tight pick where LaMa will not fit), `lama` (big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available); `auto` = LaMa > MI-GAN > cv2, best available (LaMa is auto-preferred when a learned backend is present; a memory-tight deploy pins migan). `--sensitivity auto|strict|assume-ai` (default `auto`) controls how hard a borderline mark is trusted (see the registry section: the visual detectors are metadata-independent; `auto` relaxes a mark only on same-product evidence, `assume-ai` relaxes every mark on the caller's AI assertion — the only path to high recall on a metadata-stripped screenshot). `--backend` and `--sensitivity` are shared across `visible`/`all`/`batch`. Detection keys on each mark's own shape, and under `auto` the trust gate is relaxed when local metadata confirms the vendor (a Google/Gemini C2PA issuer relaxes gemini, a China-AIGC label relaxes doubao/jimeng, `samsung_genai` relaxes samsung), so a moved or re-rendered mark is still caught. `--mark auto` (default) removes EVERY detected mark in one pass (`registry.remove_auto_marks`, not the single strongest -- a Jimeng-basic image carries both the top-left pill and the bottom-right wordmark) from: the Gemini sparkle, the Doubao "豆包AI生成" text strip, the Jimeng "★ 即梦AI" wordmark, the Samsung Galaxy AI "✦ Contenuti generati dall'AI" strip (bottom-LEFT, Italian-locale detection), and the capture-less Jimeng "AI生成" pill (top-left, `pill_engine`). The pill's weak edge-NCC detector is gated in `remove_auto_marks` via `_keep_pill` (32k real-upload corpus validation 2026-07): never on Doubao, and two confirmation arms since metadata confirms the platform, not pill presence. (1) The bottom-right wordmark fired — ~94% precise and survives metadata-STRIPPED uploads (screenshots / re-saves) — removes the pill unrestricted. (2) TC260 metadata confirms Jimeng (`"jimeng" in provenance`, from `cli._visible_provenance`) OR the caller asserts AI (`sensitivity == "assume_ai"`), no wordmark — ~27% precise, its false fires are textured ceilings/walls that the fill visibly SMEARS — removes the pill ONLY when the top-left footprint is flat enough for an invisible fill (`pill_engine.footprint_is_flat`, median-Sobel ≤ `_FLAT_TEXTURE_MAX`; the flatness guard holds even under `assume_ai`). No confirmation → never removed. `--mark gemini|doubao|jimeng|samsung|jimeng_pill` forces one (choices come from the registry). Corpus validation: doubao and jimeng localize + remove at ~100% with clean footprints (the filled region blends into its surroundings within a few LAB levels, no color shift, no dark pit); clean images with no vendor signature had 0% false removal. For arbitrary logos/objects use `erase`. **When `--mark auto` finds no known mark (the common case — ~74% of real uploads carry no registered visible mark), the command does NOT silently re-serve the input as a finished result.** It runs a cheap metadata-only `identify`, prints actionable guidance (if the image carries an invisible/metadata mark, e.g. an OpenAI/Gemini C2PA image, it points to `all`; otherwise it does NOT imply the image is clean -- it warns that an invisible pixel watermark like SynthID cannot be detected once the metadata proxy is gone and routes to both `all` and `erase --region`), writes NO output file, and exits **`EXIT_NO_VISIBLE_MARK` (2)** — distinct from success (0) and a hard error (1) so a wrapping service (raiw.cc) can surface the message instead of treating the unchanged image as done (the production "it didn't work" / score-0 trap). Same handling for an explicit `--mark <name>` that is not detected. Helper `cli._no_visible_mark_exit`; regression-guarded by `tests/test_cli.py::TestVisibleCommand::test_visible_auto_no_mark_exits_two_with_eraser_hint` and `test_visible_auto_no_mark_routes_to_all_when_metadata`. `--no-detect` still forces the gemini fallback and proceeds (exit 0).
Known-visible-mark removal by **localize -> fill**: each detected mark is localized to a binary full-frame footprint mask, then one shared, swappable fill inpaints that mask. `--backend auto|cv2|migan|lama` (default `auto`) picks the fill: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN ONNX, the memory-tight pick where LaMa will not fit), `lama` (big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available); `auto` = LaMa > MI-GAN > cv2, best available (LaMa is auto-preferred when a learned backend is present; a memory-tight deploy pins migan). `--sensitivity auto|strict|assume-ai` (default `auto`) controls how hard a borderline mark is trusted (see the registry section: the visual detectors are metadata-independent; `auto` relaxes a mark only on same-product evidence, `assume-ai` relaxes every mark on the caller's AI assertion, subject to the assumed-trust confidence floor where the vendor is unconfirmed — the only path to higher recall on a metadata-stripped screenshot). `--backend` and `--sensitivity` are shared across `visible`/`all`/`batch`. Detection keys on each mark's own shape, and under `auto` the trust gate is relaxed when local metadata confirms the vendor (a Google/Gemini C2PA issuer relaxes gemini, a China-AIGC label relaxes doubao/jimeng, `samsung_genai` relaxes samsung), so a moved or re-rendered mark is still caught. `--mark auto` (default) removes EVERY detected mark in one pass (`registry.remove_auto_marks`, not the single strongest -- a Jimeng-basic image carries both the top-left pill and the bottom-right wordmark) from: the Gemini sparkle, the Doubao "豆包AI生成" text strip, the Jimeng "★ 即梦AI" wordmark, the Samsung Galaxy AI "✦ Contenuti generati dall'AI" strip (bottom-LEFT, Italian-locale detection), and the capture-less Jimeng "AI生成" pill (top-left, `pill_engine`). The pill's weak edge-NCC detector is gated in `remove_auto_marks` via `_keep_pill` (32k real-upload corpus validation 2026-07): never on Doubao, and two confirmation arms since metadata confirms the platform, not pill presence. (1) The bottom-right wordmark fired — ~94% precise and survives metadata-STRIPPED uploads (screenshots / re-saves) — removes the pill unrestricted. (2) TC260 metadata confirms Jimeng (`"jimeng" in provenance`, from `cli._visible_provenance`) OR the caller asserts AI (`sensitivity == "assume_ai"`), no wordmark — ~27% precise, its false fires are textured ceilings/walls that the fill visibly SMEARS — removes the pill ONLY when the top-left footprint is flat enough for an invisible fill (`pill_engine.footprint_is_flat`, median-Sobel ≤ `_FLAT_TEXTURE_MAX`; the flatness guard holds even under `assume_ai`). No confirmation → never removed. `--mark gemini|doubao|jimeng|samsung|jimeng_pill` forces one (choices come from the registry). Corpus validation: doubao and jimeng localize + remove at ~100% with clean footprints (the filled region blends into its surroundings within a few LAB levels, no color shift, no dark pit); clean images with no vendor signature had 0% false removal. For arbitrary logos/objects use `erase`. **When `--mark auto` finds no known mark (the common case — ~74% of real uploads carry no registered visible mark), the command does NOT silently re-serve the input as a finished result.** It runs a cheap metadata-only `identify`, prints actionable guidance (if the image carries an invisible/metadata mark, e.g. an OpenAI/Gemini C2PA image, it points to `all`; otherwise it does NOT imply the image is clean -- it warns that an invisible pixel watermark like SynthID cannot be detected once the metadata proxy is gone and routes to both `all` and `erase --region`), writes NO output file, and exits **`EXIT_NO_VISIBLE_MARK` (2)** — distinct from success (0) and a hard error (1) so a wrapping service (raiw.cc) can surface the message instead of treating the unchanged image as done (the production "it didn't work" / score-0 trap). Same handling for an explicit `--mark <name>` that is not detected. Helper `cli._no_visible_mark_exit`; regression-guarded by `tests/test_cli.py::TestVisibleCommand::test_visible_auto_no_mark_exits_two_with_eraser_hint` and `test_visible_auto_no_mark_routes_to_all_when_metadata`. `--no-detect` still forces the gemini fallback and proceeds (exit 0).
### `batch`
+83 -39
View File
@@ -16,6 +16,7 @@ Imports stay lazy (inside the functions), so ``import remove_ai_watermarks`` is
from __future__ import annotations
from dataclasses import dataclass
from pathlib import Path
from typing import TYPE_CHECKING, Any
@@ -25,6 +26,16 @@ if TYPE_CHECKING:
from remove_ai_watermarks.watermark_registry import Backend, Sensitivity
@dataclass(frozen=True)
class _VisibleInput:
"""Normalized visible-removal input with its file-only context."""
bgr: NDArray[Any]
alpha: NDArray[Any] | None = None
path: Path | None = None
provenance: frozenset[str] = frozenset()
def visible_provenance(source: str | Path) -> frozenset[str]:
"""Vendor keys that the file's local metadata confirms, the evidence that drives
the ``auto`` sensitivity (relaxing a corroborated mark's detection trust gate).
@@ -36,19 +47,70 @@ def visible_provenance(source: str | Path) -> frozenset[str]:
"""
import contextlib
keys: set[str] = set()
path = Path(source)
with contextlib.suppress(Exception):
from remove_ai_watermarks import identify, metadata
from remove_ai_watermarks import identify
rep = identify.identify(Path(source), check_visible=False, check_invisible=False)
rep = identify.identify(path, check_visible=False, check_invisible=False)
signal_names = {signal.name for signal in rep.signals}
keys: set[str] = set()
platform = (rep.platform or "").lower()
if "google" in platform or "gemini" in platform:
keys.add("gemini")
if metadata.aigc_label(Path(source)):
if "aigc" in signal_names:
keys |= {"doubao", "jimeng"}
if metadata.samsung_genai(Path(source)):
if "samsung_genai" in signal_names:
keys.add("samsung")
return frozenset(keys)
return frozenset(keys)
return frozenset()
def _load_visible_input(source: str | Path | NDArray[Any]) -> _VisibleInput:
"""Normalize a path/array source without making the public operation stateful."""
if not isinstance(source, (str, Path)):
return _VisibleInput(source)
from remove_ai_watermarks import image_io
path = Path(source)
bgr, alpha = image_io.read_bgr_and_alpha(path)
if bgr is None:
raise ValueError(f"Could not read image: {source}")
return _VisibleInput(bgr=bgr, alpha=alpha, path=path, provenance=visible_provenance(path))
def _write_visible_result(
loaded: _VisibleInput,
result: NDArray[Any],
removed: list[str],
output: str | Path,
*,
strip_metadata: bool,
write_noop: bool,
) -> None:
"""Write one visible-removal result while preserving a true no-op losslessly."""
if not removed and not write_noop:
return
from remove_ai_watermarks import image_io
out_path = Path(output)
out_path.parent.mkdir(parents=True, exist_ok=True)
source_path = loaded.path
if not removed and source_path is not None and source_path.suffix.lower() == out_path.suffix.lower():
# Copy the ORIGINAL bytes instead of lossily re-encoding a no-op. An in-place
# call needs no copy and would otherwise raise shutil.SameFileError.
if source_path.resolve() != out_path.resolve():
import shutil
shutil.copyfile(source_path, out_path)
else:
image_io.write_bgr_with_alpha(out_path, result, loaded.alpha)
if strip_metadata:
from remove_ai_watermarks import metadata
metadata.remove_ai_metadata(out_path, out_path)
def remove_visible(
@@ -88,40 +150,22 @@ def remove_visible(
output path untouched, so a caller that treats "no mark" as "produce nothing" (the CLI
``visible`` no-mark contract) does not clobber a pre-existing file at that path.
"""
from remove_ai_watermarks import image_io, watermark_registry
alpha: NDArray[Any] | None = None
provenance: frozenset[str] = frozenset()
if isinstance(source, (str, Path)):
path = Path(source)
bgr, alpha = image_io.read_bgr_and_alpha(path)
if bgr is None:
raise ValueError(f"Could not read image: {source}")
provenance = visible_provenance(path)
else:
bgr = source
from remove_ai_watermarks import watermark_registry
loaded = _load_visible_input(source)
result, removed = watermark_registry.remove_auto_marks(
bgr, sensitivity=sensitivity, provenance=provenance, backend=backend
loaded.bgr,
sensitivity=sensitivity,
provenance=loaded.provenance,
backend=backend,
)
if output is not None and (removed or write_noop):
out_path = Path(output)
out_path.parent.mkdir(parents=True, exist_ok=True)
same_format = isinstance(source, (str, Path)) and Path(source).suffix.lower() == out_path.suffix.lower()
if not removed and same_format:
# Nothing was removed: copy the ORIGINAL bytes verbatim instead of a lossy
# re-encode of its decode, so the pixels stay bit-identical (the metadata
# strip below is lossless, so it does not disturb them either). Skip the copy
# for an in-place call (output == source): the bytes are already there, and
# shutil.copyfile would raise SameFileError.
if Path(source).resolve() != out_path.resolve(): # type: ignore[arg-type]
import shutil
shutil.copyfile(source, out_path) # type: ignore[arg-type]
else:
image_io.write_bgr_with_alpha(out_path, result, alpha)
if strip_metadata:
from remove_ai_watermarks import metadata
metadata.remove_ai_metadata(out_path, out_path)
if output is not None:
_write_visible_result(
loaded,
result,
removed,
output,
strip_metadata=strip_metadata,
write_noop=write_noop,
)
return result, removed
+252 -180
View File
@@ -12,6 +12,7 @@ import contextlib
import json
import logging
import time
from dataclasses import dataclass
from pathlib import Path
from typing import TYPE_CHECKING, Any, Literal, NoReturn
@@ -311,8 +312,9 @@ _visible_sensitivity_option = click.option(
help="How hard to trust a borderline mark. auto: relax a mark only when metadata "
"or a same-product sibling mark corroborates it (safe; clean images untouched). "
"strict: high-precision visual gate only, never relaxed. assume-ai: treat the "
"image as AI and relax every mark (best recall on metadata-stripped screenshots, "
"at the cost of a small fill on some clean corners).",
"image as AI and relax every mark, keeping a confidence floor where the vendor is "
"unconfirmed (best recall on metadata-stripped screenshots; a clean image is still "
"left untouched).",
)
@@ -525,6 +527,113 @@ def main(ctx: click.Context, verbose: bool) -> None:
# ── Visible (Gemini) watermark removal ──
def _run_visible_auto(
source: Path,
output: Path,
*,
backend: watermark_registry.Backend,
sensitivity: watermark_registry.Sensitivity,
strip_metadata: bool,
) -> None:
"""Run the registry-wide visible pass and render its CLI result."""
from remove_ai_watermarks import api
t0 = time.monotonic()
try:
with console.status("Detecting & removing visible marks..."):
result, removed = api.remove_visible(
str(source),
str(output),
sensitivity=sensitivity,
backend=backend,
strip_metadata=strip_metadata,
write_noop=False,
)
except RuntimeError as e: # selected migan/lama backend whose extra is absent
console.print(f" Error: {e}")
raise SystemExit(1) from e
except (ValueError, OSError) as e: # unreadable / truncated / non-image input
console.print(f" Error: cannot read image {source.name}: {e}")
raise SystemExit(1) from e
elapsed = time.monotonic() - t0
h, w = result.shape[:2]
console.print(f" Input: {source.name} ({w}x{h})")
if not removed:
# write_noop=False means nothing was written, so a pre-existing output is intact.
console.print(" No known visible mark detected (gemini / doubao / jimeng / jimeng-pill / samsung).")
_no_visible_mark_exit(source, sensitivity=sensitivity)
console.print(f" Removed: {', '.join(removed)}")
size_kb = output.stat().st_size / 1024
console.print(f" Saved: {output} ({size_kb:.0f} KB, {elapsed:.2f}s)")
def _run_visible_explicit(
ctx: click.Context,
source: Path,
output: Path,
*,
detect: bool,
mark: str,
backend: watermark_registry.Backend,
sensitivity: watermark_registry.Sensitivity,
resolved_backend: str,
strip_metadata: bool,
) -> None:
"""Run one explicitly selected visible-mark detector/remover."""
image, alpha = image_io.read_bgr_and_alpha(source)
if image is None:
console.print(f"Error: Failed to read image: {source}")
raise SystemExit(1)
h, w = image.shape[:2]
console.print(f" Input: {source.name} ({w}x{h})")
provenance = _visible_provenance(source)
target = "gemini" if mark == "auto" else mark # --no-detect auto: gemini fallback
chosen = watermark_registry.get_mark(target)
# A single explicit mark has no sibling corroboration. Keep its trust resolution
# aligned with the registry arbiter, including the assumption-only floor.
trust = watermark_registry.resolve_trust(
chosen.key,
sensitivity=sensitivity,
provenance=provenance,
strict_keys=set(),
)
relax = trust != "strict"
detection = chosen.detect(image, provenance=relax)
if trust == "assumed" and not watermark_registry.assumed_floor_ok(chosen.key, detection.confidence):
relax = False
detection = chosen.detect(image, provenance=False)
if detect and not detection.detected:
console.print(f" {chosen.label} not detected (conf {detection.confidence:.2f}). Use --no-detect to force.")
_no_visible_mark_exit(source, sensitivity=sensitivity)
if detection.detected:
console.print(f" {chosen.label} detected ({chosen.location}, conf {detection.confidence:.2f})")
t0 = time.monotonic()
try:
with console.status(f"Removing {chosen.label}... ({resolved_backend})"):
result, _ = chosen.remove(image, backend=backend, provenance=relax, force=not detect)
except RuntimeError as e: # selected migan/lama backend whose extra is absent
console.print(f" Error: {e}")
raise SystemExit(1) from e
elapsed = time.monotonic() - t0
output.parent.mkdir(parents=True, exist_ok=True)
image_io.write_bgr_with_alpha(output, result, alpha)
if strip_metadata:
try:
from remove_ai_watermarks.metadata import remove_ai_metadata
remove_ai_metadata(output, output)
except Exception as e:
if ctx.obj.get("verbose"):
console.print(f" Warning: Failed to strip metadata: {e}")
size_kb = output.stat().st_size / 1024
console.print(f" Saved: {output} ({size_kb:.0f} KB, {elapsed:.2f}s)")
@main.command("visible")
@click.argument("source", type=click.Path(exists=True, path_type=Path))
@click.option(
@@ -560,18 +669,16 @@ def cmd_visible(
MI-GAN > cv2). ``--mark auto`` removes every detected mark in one
pass. For arbitrary logos/objects, use ``erase``.
"""
from remove_ai_watermarks import watermark_registry as registry
_banner()
source = _validate_image(source)
if output is None:
output = source.with_stem(source.stem + "_clean")
bk: registry.Backend = backend # type: ignore[assignment]
bk: watermark_registry.Backend = backend # type: ignore[assignment]
sens = _parse_sensitivity(sensitivity)
resolved_backend = registry.resolve_backend(bk)
if resolved_backend == "cv2" and not registry.inpaint_model_available():
resolved_backend = watermark_registry.resolve_backend(bk)
if resolved_backend == "cv2" and not watermark_registry.inpaint_model_available():
console.print(" Note: using cv2 fill (install the 'migan' extra for a lightweight ONNX model).")
# ``auto`` removes EVERY detected in_auto mark in one pass (a Jimeng-basic image
@@ -579,82 +686,20 @@ def cmd_visible(
# read -> provenance -> localize/fill -> write -> metadata-strip to the library
# entry point, so the CLI and the library go through ONE path (no drift).
if mark == "auto" and detect:
from remove_ai_watermarks import api
t0 = time.monotonic()
try:
with console.status("Detecting & removing visible marks..."):
result, removed = api.remove_visible(
str(source),
str(output),
sensitivity=sens,
backend=bk,
strip_metadata=strip_metadata,
write_noop=False,
)
except RuntimeError as e: # e.g. a selected migan/lama backend whose extra is absent
console.print(f" Error: {e}")
raise SystemExit(1) from e
except (ValueError, OSError) as e: # unreadable / truncated / non-image input
console.print(f" Error: cannot read image {source.name}: {e}")
raise SystemExit(1) from e
elapsed = time.monotonic() - t0
h, w = result.shape[:2]
console.print(f" Input: {source.name} ({w}x{h})")
if not removed:
# write_noop=False means nothing was written, so a pre-existing file at the
# output path is left intact (the no-mark contract writes nothing).
console.print(" No known visible mark detected (gemini / doubao / jimeng / jimeng-pill / samsung).")
_no_visible_mark_exit(source, sensitivity=sens)
console.print(f" Removed: {', '.join(removed)}")
size_kb = output.stat().st_size / 1024
console.print(f" Saved: {output} ({size_kb:.0f} KB, {elapsed:.2f}s)")
_run_visible_auto(source, output, backend=bk, sensitivity=sens, strip_metadata=strip_metadata)
return
# Explicit single mark (or --no-detect): needs the decoded array + the per-mark gate,
# so it keeps its own read/remove/write (still through the shared io + registry).
image, alpha = image_io.read_bgr_and_alpha(source)
if image is None:
console.print(f"Error: Failed to read image: {source}")
raise SystemExit(1)
h, w = image.shape[:2]
console.print(f" Input: {source.name} ({w}x{h})")
provenance = _visible_provenance(source)
target = "gemini" if mark == "auto" else mark # --no-detect auto: gemini fallback
chosen = registry.get_mark(target)
# A single explicit mark has no cross-mark pass (no sibling corroboration), so use the
# canonical arbiter policy with an empty strict-sibling set instead of re-deriving it
# inline (keeps this in lockstep with `decide`).
prov = registry.resolve_relax(chosen.key, sensitivity=sens, provenance=provenance, strict_keys=set())
det = chosen.detect(image, provenance=prov)
if detect and not det.detected:
console.print(f" {chosen.label} not detected (conf {det.confidence:.2f}). Use --no-detect to force.")
_no_visible_mark_exit(source, sensitivity=sens)
if det.detected:
console.print(f" {chosen.label} detected ({chosen.location}, conf {det.confidence:.2f})")
t0 = time.monotonic()
try:
with console.status(f"Removing {chosen.label}... ({resolved_backend})"):
result, _ = chosen.remove(image, backend=bk, provenance=prov, force=not detect)
except RuntimeError as e: # e.g. a selected migan/lama backend whose extra is absent
console.print(f" Error: {e}")
raise SystemExit(1) from e
elapsed = time.monotonic() - t0
# Save (rejoins the original alpha plane unchanged) + strip metadata.
output.parent.mkdir(parents=True, exist_ok=True)
image_io.write_bgr_with_alpha(output, result, alpha)
if strip_metadata:
try:
from remove_ai_watermarks.metadata import remove_ai_metadata
remove_ai_metadata(output, output)
except Exception as e:
if ctx.obj.get("verbose"):
console.print(f" Warning: Failed to strip metadata: {e}")
size_kb = output.stat().st_size / 1024
console.print(f" Saved: {output} ({size_kb:.0f} KB, {elapsed:.2f}s)")
_run_visible_explicit(
ctx,
source,
output,
detect=detect,
mark=mark,
backend=bk,
sensitivity=sens,
resolved_backend=resolved_backend,
strip_metadata=strip_metadata,
)
# ── Universal region eraser ──
@@ -1261,32 +1306,106 @@ def _passthrough_copy(img_path: Path, out_path: Path) -> None:
image_io.write_bgr_with_alpha(out_path, src_bgr, src_alpha)
@dataclass(frozen=True)
class _BatchOptions:
"""Validated processing options shared by every image in one batch.
Click necessarily exposes these as individual command parameters, but the
processing core should receive one coherent value instead of a 21-argument
call. Keeping the object immutable also makes it safe to reuse while the
batch caches model instances in ``ctx.obj``.
"""
strength: float | None
steps: int
pipeline: str
device: str
seed: int | None
hf_token: str | None
humanize: float
backend: str = "auto"
sensitivity: str = "auto"
unsharp: float = 0.0
max_resolution: int = 0
min_resolution: int = 1024
controlnet_scale: float = 1.0
upscaler: str = "lanczos"
model: str | None = None
guidance_scale: float | None = None
adaptive_polish: bool = False
tile: bool = False
tile_size: int = 1024
tile_overlap: int = 128
force: bool = False
def _run_batch_invisible(
ctx: click.Context,
img_path: Path,
out_path: Path,
mode: str,
options: _BatchOptions,
) -> bool:
"""Run or safely skip the invisible pass for one batch image.
Returns ``True`` only when a detectable target could not be processed because
the GPU dependencies are missing. The availability probe is intentionally
evaluated once so branching cannot observe inconsistent optional-dependency
state.
"""
from remove_ai_watermarks.invisible_engine import is_available as invisible_available
skip_no_signal = _should_skip_invisible_scrub(options.force, img_path)
available = invisible_available()
if available and not skip_no_signal:
from remove_ai_watermarks.invisible_engine import InvisibleEngine
# Cache the engine in ctx.obj so the batch builds it once (pipeline is a
# single CLI value, constant across the run).
engines = ctx.obj.setdefault("_inv_engines", {})
if options.pipeline not in engines:
engines[options.pipeline] = InvisibleEngine(
model_id=options.model,
device=None if options.device == "auto" else options.device,
pipeline=options.pipeline,
hf_token=options.hf_token,
controlnet_conditioning_scale=options.controlnet_scale,
)
engines[options.pipeline].remove_watermark(
img_path if mode == "invisible" else out_path,
out_path,
strength=options.strength,
num_inference_steps=options.steps,
guidance_scale=options.guidance_scale,
seed=options.seed,
humanize=options.humanize,
unsharp=options.unsharp,
adaptive_polish=options.adaptive_polish,
max_resolution=options.max_resolution,
min_resolution=options.min_resolution,
upscaler=options.upscaler,
tile=options.tile,
tile_size=options.tile_size,
tile_overlap=options.tile_overlap,
# Detect the vendor from the pristine original (`img_path`), not the
# visible-processed `out_path` whose C2PA is already gone.
vendor=vendor_for_strength(img_path),
)
return False
# Invisible-only mode has no preceding visible pass to create ``out_path``.
# Preserve a complete output directory while deliberately leaving pixels intact.
if mode == "invisible" and not out_path.exists():
_passthrough_copy(img_path, out_path)
return not available and not skip_no_signal
def _process_batch_image(
ctx: click.Context,
img_path: Path,
out_path: Path,
mode: str,
strength: float | None,
steps: int,
pipeline: str,
device: str,
seed: int | None,
hf_token: str | None,
humanize: float,
backend: str = "auto",
sensitivity: str = "auto",
unsharp: float = 0.0,
max_resolution: int = 0,
min_resolution: int = 1024,
controlnet_scale: float = 1.0,
upscaler: str = "lanczos",
model: str | None = None,
guidance_scale: float | None = None,
adaptive_polish: bool = False,
tile: bool = False,
tile_size: int = 1024,
tile_overlap: int = 128,
force: bool = False,
options: _BatchOptions,
) -> bool:
"""Process a single image for batch mode.
@@ -1312,71 +1431,21 @@ def _process_batch_image(
if image is None:
raise ValueError("Failed to read image")
result, _ = _remove_visible_auto(image, source_path=img_path, backend=backend, sensitivity=sensitivity)
result, _ = _remove_visible_auto(
image,
source_path=img_path,
backend=options.backend,
sensitivity=options.sensitivity,
)
image_io.write_bgr_with_alpha(out_path, result, alpha)
saved_alpha = alpha
if mode in ("invisible", "all"):
from remove_ai_watermarks.invisible_engine import (
is_available as invisible_available,
)
# Skip the destructive regeneration when no invisible watermark is locally
# detectable (would only degrade a clean image). Read the pristine `img_path`;
# `out_path` may already be the visible-processed result. --force overrides.
skip_no_signal = _should_skip_invisible_scrub(force, img_path)
if invisible_available() and not skip_no_signal:
from remove_ai_watermarks.invisible_engine import InvisibleEngine
# Cache the engine in ctx.obj so the batch builds it once (pipeline is a
# single CLI value, constant across the run).
engines = ctx.obj.setdefault("_inv_engines", {})
if pipeline not in engines:
engines[pipeline] = InvisibleEngine(
model_id=model,
device=None if device == "auto" else device,
pipeline=pipeline,
hf_token=hf_token,
controlnet_conditioning_scale=controlnet_scale,
)
engine_inv = engines[pipeline]
engine_inv.remove_watermark(
img_path if mode == "invisible" else out_path,
out_path,
strength=strength,
num_inference_steps=steps,
guidance_scale=guidance_scale,
seed=seed,
humanize=humanize,
unsharp=unsharp,
adaptive_polish=adaptive_polish,
max_resolution=max_resolution,
min_resolution=min_resolution,
upscaler=upscaler,
tile=tile,
tile_size=tile_size,
tile_overlap=tile_overlap,
# Detect the vendor from the pristine original (`img_path`), not the
# visible-processed `out_path` whose C2PA is already gone.
vendor=vendor_for_strength(img_path),
)
elif not invisible_available() and not skip_no_signal:
# An invisible signal IS present but the GPU deps are missing, so the
# SynthID scrub cannot run. Mirror the single `all` command's loud skip:
# flag it for a batch-level warning + non-zero exit (a silently retained
# SynthID watermark is the #1 "it didn't work" report). For invisible mode
# nothing wrote out_path yet -> copy the input through so the output dir is
# complete with the pixels deliberately left intact (without this, a
# signal-bearing image in a GPU-less --mode invisible run got NO output).
synthid_skipped = True
if mode == "invisible" and not out_path.exists():
_passthrough_copy(img_path, out_path)
elif skip_no_signal and mode == "invisible" and not out_path.exists():
# No invisible target and the visible/all pass did not write out_path
# (invisible mode): copy the input through so the output dir is complete
# with the pixels deliberately left intact.
_passthrough_copy(img_path, out_path)
synthid_skipped = _run_batch_invisible(ctx, img_path, out_path, mode, options)
if mode in ("metadata", "all"):
from remove_ai_watermarks.metadata import remove_ai_metadata
@@ -1485,6 +1554,29 @@ def cmd_batch(
if mode in ("invisible", "all"):
_warn_if_esrgan_unavailable(upscaler)
adaptive_polish = _resolve_auto_polish(auto, adaptive_polish)
options = _BatchOptions(
strength=strength,
steps=steps,
pipeline=pipeline,
device=device,
seed=seed,
hf_token=hf_token,
humanize=humanize,
backend=backend,
sensitivity=sensitivity,
unsharp=unsharp,
max_resolution=max_resolution,
min_resolution=min_resolution,
controlnet_scale=controlnet_scale,
upscaler=upscaler,
model=model,
guidance_scale=guidance_scale,
adaptive_polish=adaptive_polish,
tile=tile,
tile_size=tile_size,
tile_overlap=tile_overlap,
force=force,
)
processed = 0
errors = 0
@@ -1510,27 +1602,7 @@ def cmd_batch(
img_path=img_path,
out_path=out_path,
mode=mode,
strength=strength,
steps=steps,
pipeline=pipeline,
device=device,
seed=seed,
hf_token=hf_token,
humanize=humanize,
backend=backend,
sensitivity=sensitivity,
unsharp=unsharp,
max_resolution=max_resolution,
min_resolution=min_resolution,
controlnet_scale=controlnet_scale,
upscaler=upscaler,
model=model,
guidance_scale=guidance_scale,
adaptive_polish=adaptive_polish,
tile=tile,
tile_size=tile_size,
tile_overlap=tile_overlap,
force=force,
options=options,
):
synthid_skipped_count += 1
processed += 1
+37 -29
View File
@@ -505,6 +505,42 @@ def _trustmark(image_path: Path) -> str | None:
return detect_trustmark(image_path)
def _collect_visible_signals(
image_path: Path,
signals: list[Signal],
watermarks: list[str],
platform: str | None,
) -> str | None:
"""Decode once, append every trusted visible-mark signal, and return platform.
Keeping this stage separate from metadata aggregation makes the optional cv2
boundary explicit and guarantees that all visible detectors share one decoded
BGR array. A decode failure preserves the detectors' historical fallback/no-op
behavior.
"""
image: NDArray[Any] | None = None
try:
from remove_ai_watermarks.image_io import imread
image = imread(image_path)
except Exception as exc: # cv2 missing - detectors fall back / no-op
logger.debug("visible-mark decode unavailable: %s", exc)
sparkle_conf = _visible_sparkle(image_path, image=image)
if sparkle_conf is not None and sparkle_conf >= _SPARKLE_THRESHOLD:
signals.append(Signal("visible_sparkle", f"NCC confidence {sparkle_conf:.2f}", "medium"))
watermarks.append(f"Visible Gemini sparkle (confidence {sparkle_conf:.2f})")
if platform is None:
platform = "Google Gemini family (visible sparkle detected)"
for detection in _visible_text_marks(image_path, image=image):
signals.append(Signal(f"visible_{detection.key}", f"NCC confidence {detection.confidence:.2f}", "medium"))
watermarks.append(f"Visible {detection.label} (confidence {detection.confidence:.2f})")
if platform is None:
platform = _VISIBLE_MARK_PLATFORM[detection.key]
return platform
def identify(image_path: Path, *, check_visible: bool = True, check_invisible: bool = True) -> ProvenanceReport:
"""Identify an image's origin platform and watermark inventory.
@@ -755,36 +791,8 @@ def identify(image_path: Path, *, check_visible: bool = True, check_invisible: b
or xai_sig
)
# Decode the file ONCE for every visible-mark detector. The sparkle and the
# text-mark detectors both consume a BGR array; letting each re-read the file
# was two full cv2 decodes of the same bitmap, which spikes memory on a small
# worker. None (cv2 missing / unreadable container) makes each detector fall
# back to its own read, preserving the old behavior.
vis_image: NDArray[Any] | None = None
if check_visible:
try:
from remove_ai_watermarks.image_io import imread
vis_image = imread(image_path)
except Exception as exc: # cv2 missing - detectors fall back / no-op
logger.debug("visible-mark decode unavailable: %s", exc)
# ── Visible Gemini sparkle (fallback for stripped-metadata case) ─
sparkle_conf = _visible_sparkle(image_path, image=vis_image) if check_visible else None
if sparkle_conf is not None and sparkle_conf >= _SPARKLE_THRESHOLD:
signals.append(Signal("visible_sparkle", f"NCC confidence {sparkle_conf:.2f}", "medium"))
watermarks.append(f"Visible Gemini sparkle (confidence {sparkle_conf:.2f})")
if platform is None:
platform = "Google Gemini family (visible sparkle detected)"
# ── Visible Doubao / Jimeng text marks (registry; same stripped-metadata
# fallback role as the Gemini sparkle above) ─
if check_visible:
for det in _visible_text_marks(image_path, image=vis_image):
signals.append(Signal(f"visible_{det.key}", f"NCC confidence {det.confidence:.2f}", "medium"))
watermarks.append(f"Visible {det.label} (confidence {det.confidence:.2f})")
if platform is None:
platform = _VISIBLE_MARK_PLATFORM[det.key]
platform = _collect_visible_signals(image_path, signals, watermarks, platform)
visible_only = any(s.name.startswith("visible_") for s in signals) and not ai_from_metadata
hf_only = bool(hf_job) and not ai_from_metadata
+2 -2
View File
@@ -347,7 +347,7 @@ def has_ai_metadata(image_path: Path) -> bool:
return True
# China TC260 AIGC label as a PNG text chunk (the byte scan above catches
# only the XMP form; the raw-JSON tEXt chunk needs the PIL-based parse).
if aigc_label(image_path):
if aigc_label(image_path) is not None:
return True
# HuggingFace-hosted job marker (hf-job-id PNG text chunk).
if huggingface_job(image_path):
@@ -887,7 +887,7 @@ def get_ai_metadata(image_path: Path) -> dict[str, str]:
result["soft_binding"] = ", ".join(vendors)
# China TC260 AI-content label (Doubao and other China-served generators).
if aigc := aigc_label(image_path):
if (aigc := aigc_label(image_path)) is not None:
producer = aigc.get("ContentProducer", "")
result["aigc_label"] = f"China AIGC label (TC260){f'; producer {producer}' if producer else ''}"
+75 -25
View File
@@ -50,16 +50,24 @@ Backend = Literal["auto", "cv2", "migan", "lama"]
# a clean corner). Lowest recall on faint/moved marks.
# * ``auto`` (default): relax a mark's gate ONLY when the image carries same-product
# evidence the mark is there -- metadata provenance for that vendor, or a confidently
# detected sibling mark of the same product (see ``resolve_relax``). No evidence ->
# detected sibling mark of the same product (see ``resolve_trust``). No evidence ->
# stays strict. Safe: it only escalates where the mark is corroborated.
# * ``assume_ai``: relax every mark's gate regardless of evidence -- the caller asserts
# the image is AI and wants the mark gone (e.g. a metadata-stripped screenshot uploaded
# to a watermark remover). Recovers the faint/moved marks the strict gate demotes
# (~49% -> ~89% Gemini recall, corpus-measured), at the cost of a harmless small fill
# on some clean corners. The library CANNOT infer this from a stripped image -- only the
# caller's out-of-band context (the user uploaded to remove a mark) justifies it.
# to a watermark remover). Recovers the faint/moved marks the strict gate demotes. The
# library CANNOT infer this from a stripped image -- only the caller's out-of-band
# context (the user uploaded to remove a mark) justifies it. An assertion that the
# image is AI is NOT evidence of WHICH vendor made it, so a mark relaxed on assumption
# alone must still clear ``_ASSUMED_CONF_FLOOR``; see that constant for why.
Sensitivity = Literal["auto", "strict", "assume_ai"]
# The trust level a mark's detection gate is resolved to (see ``resolve_trust``). The
# split between ``assumed`` and ``confirmed`` is load-bearing: both bypass the engine's
# false-positive gate, but only ``confirmed`` has evidence naming THIS vendor, which is
# exactly what that bypass is documented to require (see GeminiEngine.detect_watermark's
# ``trust_provenance`` contract). ``assumed`` therefore carries a confidence floor.
Trust = Literal["strict", "assumed", "confirmed"]
# Product family per mark, for the ``auto`` cross-mark corroboration: a confidently
# detected mark relaxes only OTHER marks of the SAME product (different corners, one
# product -- the Jimeng wordmark + the Jimeng pill). Doubao and Jimeng are BOTH ByteDance
@@ -120,6 +128,8 @@ class Candidate:
Carries the mark's verdict at BOTH trust levels (``detected_strict`` = the
conservative gate, ``detected_relaxed`` = the gate the engine relaxes to under
provenance/assume), so the arbiter can pick per mark without re-running detection.
``relaxed_confidence`` is the gate-bypassed detection's confidence, which the arbiter
needs to apply :func:`assumed_floor_ok` when a mark is relaxed on assumption alone.
``features`` is a generic bag of physical measurements a mark's gate may need (the
mark owns which it reports via ``KnownMark._features``); e.g. the pill supplies
``footprint_flat`` (0/1). Empty for marks whose gate needs no extra evidence."""
@@ -128,6 +138,7 @@ class Candidate:
label: str
detected_strict: bool
detected_relaxed: bool
relaxed_confidence: float
features: dict[str, float] # generic; both construction sites always supply it (empty when none)
@@ -446,28 +457,59 @@ def detect_marks(
return [m.detect(image, provenance=m.key in provenance) for m in _REGISTRY if include_explicit or m.in_auto]
def resolve_relax(
# Minimum gate-bypassed confidence a mark must reach when it is relaxed on ASSUMPTION
# (``assume_ai``) rather than on evidence naming its vendor. Relaxing bypasses the
# engine's false-positive gate entirely, which is justified by vendor CONFIRMATION; an
# assumption that the image is AI says nothing about WHICH vendor, so the bypassed
# detector needs its own floor or it fires on ordinary content.
#
# Corpus-measured 2026-07-16 (256 genuine camera captures -- Make/Model/exposure/aperture
# present and no AI token, so a Gemini sparkle cannot be there -- vs 697 Google-C2PA
# positives, metadata used only as the label, never fed to the detector):
#
# bypassed threshold recall false-fire on clean photos
# 0.35 82.6% 59.8% <- the bare detector gate
# 0.45 66.6% 12.5%
# 0.50 59.4% 0.0% <- chosen
# strict gate 56.4% 0.0%
#
# So 0.35 sat on a cliff: it bought +26pp recall over strict by filling a corner on ~6
# of every 10 CLEAN photos. At 0.50 the flag is honest -- it still beats strict, for free.
# Marks absent from this dict relax identically at both levels; their bypassed false-fire
# on the same negatives is under 1% (doubao 0.8%, jimeng 0.4%, samsung 0.4%).
_ASSUMED_CONF_FLOOR: dict[str, float] = {"gemini": 0.50}
def assumed_floor_ok(key: str, confidence: float) -> bool:
"""Whether an ``assumed``-trust detection of ``key`` at ``confidence`` is trustworthy
enough to act on (see :data:`_ASSUMED_CONF_FLOOR`). Marks with no floor always pass."""
floor = _ASSUMED_CONF_FLOOR.get(key)
return floor is None or confidence >= floor
def resolve_trust(
key: str,
*,
sensitivity: Sensitivity,
provenance: frozenset[str],
strict_keys: set[str],
) -> bool:
"""Whether mark ``key``'s detection gate is relaxed (strict -> assume level).
) -> Trust:
"""The trust level mark ``key``'s detection gate is resolved to.
The single place that turns the ``sensitivity`` policy + evidence into a per-mark
boolean (which the engines consume): ``strict`` never relaxes, ``assume_ai`` always
relaxes, and ``auto`` relaxes only on same-product evidence -- the vendor confirmed
by metadata (``key in provenance``) or a confidently strict-detected sibling of the
same product (``_PRODUCT_OF``)."""
level (which the engines consume as ``provenance = level != "strict"``). ``strict``
never relaxes. A mark is ``confirmed`` only on same-product evidence -- the vendor
confirmed by metadata (``key in provenance``) or a confidently strict-detected
sibling of the same product (``_PRODUCT_OF``). Without that evidence, ``assume_ai``
yields ``assumed`` (relaxed, but subject to :func:`assumed_floor_ok`) and ``auto``
stays ``strict``."""
if sensitivity == "strict":
return False
if sensitivity == "assume_ai":
return True
if key in provenance:
return True
return "strict"
product = _PRODUCT_OF[key]
return any(_PRODUCT_OF[k] == product for k in strict_keys if k != key)
confirmed = key in provenance or any(_PRODUCT_OF[k] == product for k in strict_keys if k != key)
if confirmed:
return "confirmed"
return "assumed" if sensitivity == "assume_ai" else "strict"
def _keep_pill(keys: set[str], *, provenance: frozenset[str], sensitivity: Sensitivity, footprint_flat: bool) -> bool:
@@ -515,7 +557,7 @@ def _build_candidates(image: NDArray[Any]) -> list[Candidate]:
strict = m.detect(image, provenance=False)
relaxed = m.detect(image, provenance=True)
feats = m.features(image) if (strict.detected or relaxed.detected) else {}
cands.append(Candidate(m.key, m.label, strict.detected, relaxed.detected, feats))
cands.append(Candidate(m.key, m.label, strict.detected, relaxed.detected, relaxed.confidence, feats))
return cands
@@ -523,17 +565,25 @@ def decide(candidates: list[Candidate], context: Context) -> list[Decision]:
"""The removal ARBITER: a pure function turning perception + context into the
ordered list of marks to remove (and the trust level each was accepted at).
All policy lives here, in one place: per-mark relaxation (:func:`resolve_relax`,
which needs the strict-detected siblings for ``auto`` cross-mark corroboration) and
the capture-less pill gate (:func:`_keep_pill`). No image, no I/O -- so it is
unit-testable in isolation and the same decision drives every caller."""
All policy lives here, in one place: per-mark trust resolution (:func:`resolve_trust`,
which needs the strict-detected siblings for ``auto`` cross-mark corroboration), the
assumed-trust confidence floor (:func:`assumed_floor_ok`) and the capture-less pill
gate (:func:`_keep_pill`). No image, no I/O -- so it is unit-testable in isolation and
the same decision drives every caller."""
strict_keys = {c.key for c in candidates if c.detected_strict}
fired: list[Decision] = []
for c in candidates:
relax = resolve_relax(
trust = resolve_trust(
c.key, sensitivity=context.sensitivity, provenance=context.provenance, strict_keys=strict_keys
)
if c.detected_relaxed if relax else c.detected_strict:
relax = trust != "strict"
ok = c.detected_relaxed if relax else c.detected_strict
if trust == "assumed" and not assumed_floor_ok(c.key, c.relaxed_confidence):
# Relaxed on assumption alone and too weak to trust: fall back to the strict
# verdict rather than dropping the mark, so assume_ai is monotonic -- it only
# ever ADDS recall over strict, never removes less than strict would.
ok, relax = c.detected_strict, False
if ok:
fired.append(Decision(c, relax))
keys = {d.candidate.key for d in fired}
if "jimeng_pill" in keys:
+18
View File
@@ -104,6 +104,24 @@ class TestVisibleProvenance:
def test_unreadable_path_is_empty(self, tmp_path):
assert raiw.visible_provenance(tmp_path / "missing.png") == frozenset()
def test_uses_report_signals_for_falsy_metadata_values(self, monkeypatch, tmp_path):
"""An empty TC260 object and Samsung genAIType=0 are still present signals.
The report has already normalized those values, so the public API must not
re-read the file and accidentally discard them by truthiness.
"""
from types import SimpleNamespace
from remove_ai_watermarks import identify
report = SimpleNamespace(
platform=None,
signals=[SimpleNamespace(name="aigc"), SimpleNamespace(name="samsung_genai")],
)
monkeypatch.setattr(identify, "identify", lambda *args, **kwargs: report)
assert raiw.visible_provenance(tmp_path / "synthetic.png") == frozenset({"doubao", "jimeng", "samsung"})
class TestRemoveVisibleOutputPath:
"""Output-path robustness: in-place clean (#3) and a missing output dir (#4)."""
+13
View File
@@ -978,6 +978,19 @@ class TestAIGCLabel:
assert "aigc_label" in meta
assert "TC260" in meta["aigc_label"]
def test_empty_namespaced_label_is_still_surfaced(self, tmp_path: Path):
"""The namespaced element is unambiguous even when its JSON object is empty."""
from remove_ai_watermarks.metadata import aigc_label
p = tmp_path / "empty_aigc.png"
Image.new("RGB", (32, 32)).save(p)
with open(p, "ab") as f:
f.write(b"<TC260:AIGC>{}</TC260:AIGC>")
assert aigc_label(p) == {}
assert has_ai_metadata(p)
assert "aigc_label" in get_ai_metadata(p)
def _aigc_chunk_png(self, tmp_path: Path, producer: str = "doubao") -> Path:
"""Doubao writes the TC260 object as a PNG ``tEXt`` chunk keyed ``AIGC``
with raw JSON (no XMP, no namespaced marker)."""
+62 -12
View File
@@ -165,36 +165,60 @@ class TestLocalizeFill:
class TestSensitivity:
"""``resolve_relax`` turns the sensitivity policy + evidence into the per-mark
relaxation boolean the engines consume."""
"""``resolve_trust`` turns the sensitivity policy + evidence into the per-mark
trust level the engines consume."""
def test_strict_never_relaxes(self):
# even with metadata provenance, strict keeps the conservative gate
assert (
reg.resolve_relax("gemini", sensitivity="strict", provenance=frozenset({"gemini"}), strict_keys=set())
is False
reg.resolve_trust("gemini", sensitivity="strict", provenance=frozenset({"gemini"}), strict_keys=set())
== "strict"
)
def test_assume_ai_always_relaxes(self):
assert reg.resolve_relax("gemini", sensitivity="assume_ai", provenance=frozenset(), strict_keys=set()) is True
def test_assume_ai_without_evidence_is_assumed_not_confirmed(self):
# asserting the image is AI says nothing about WHICH vendor made it, so the mark
# is relaxed on assumption only -- it must not inherit the confirmed-vendor bypass
assert (
reg.resolve_trust("gemini", sensitivity="assume_ai", provenance=frozenset(), strict_keys=set()) == "assumed"
)
def test_assume_ai_with_metadata_is_confirmed(self):
assert (
reg.resolve_trust("gemini", sensitivity="assume_ai", provenance=frozenset({"gemini"}), strict_keys=set())
== "confirmed"
)
def test_auto_relaxes_on_own_metadata(self):
assert (
reg.resolve_relax("gemini", sensitivity="auto", provenance=frozenset({"gemini"}), strict_keys=set()) is True
reg.resolve_trust("gemini", sensitivity="auto", provenance=frozenset({"gemini"}), strict_keys=set())
== "confirmed"
)
def test_auto_strict_without_evidence(self):
assert reg.resolve_relax("gemini", sensitivity="auto", provenance=frozenset(), strict_keys=set()) is False
assert reg.resolve_trust("gemini", sensitivity="auto", provenance=frozenset(), strict_keys=set()) == "strict"
def test_auto_cross_mark_same_product(self):
# a detected Jimeng wordmark relaxes the Jimeng pill (same product, other corner)
assert (
reg.resolve_relax("jimeng_pill", sensitivity="auto", provenance=frozenset(), strict_keys={"jimeng"}) is True
reg.resolve_trust("jimeng_pill", sensitivity="auto", provenance=frozenset(), strict_keys={"jimeng"})
== "confirmed"
)
def test_auto_no_cross_mark_across_products(self):
# a detected Jimeng wordmark must NOT relax Doubao (distinct products, same corner)
assert reg.resolve_relax("doubao", sensitivity="auto", provenance=frozenset(), strict_keys={"jimeng"}) is False
assert (
reg.resolve_trust("doubao", sensitivity="auto", provenance=frozenset(), strict_keys={"jimeng"}) == "strict"
)
def test_assumed_floor_rejects_weak_sparkle_but_passes_strong(self):
# the gate-bypassed sparkle detector fires on ~60% of ordinary photos at its bare
# 0.35 threshold; only a match well clear of that floor is trustworthy on assumption
assert reg.assumed_floor_ok("gemini", 0.35) is False
assert reg.assumed_floor_ok("gemini", 0.50) is True
def test_assumed_floor_default_passes_for_unfloored_marks(self):
# text marks relax cleanly (<1% bypassed false-fire), so they carry no floor
assert reg.assumed_floor_ok("doubao", 0.36) is True
def test_remove_auto_marks_accepts_all_sensitivities(self):
blank = np.zeros((256, 256, 3), np.uint8)
@@ -209,9 +233,11 @@ class TestArbiter:
Candidates -- this is the payoff of separating decision from perception."""
@staticmethod
def _c(key, *, strict=False, relaxed=False, flat=False):
def _c(key, *, strict=False, relaxed=False, flat=False, relaxed_conf=1.0):
# relaxed_conf defaults high so a test that does not care about the assumed-trust
# confidence floor exercises the trust logic, not the floor.
feats = {"footprint_flat": 1.0} if flat else {}
return reg.Candidate(key, f"L:{key}", strict, relaxed, feats)
return reg.Candidate(key, f"L:{key}", strict, relaxed, relaxed_conf, feats)
def _keys(self, cands, ctx):
return {d.candidate.key for d in reg.decide(cands, ctx)}
@@ -228,6 +254,30 @@ class TestArbiter:
assert [d.candidate.key for d in fired] == ["gemini"]
assert fired[0].relax is True
def test_assume_ai_drops_sparkle_below_the_assumed_floor(self):
# REGRESSION (2026-07-16): assume_ai passed trust_provenance=True to the engine,
# bypassing the sparkle false-positive gate on the mere ASSERTION that the image is
# AI -- but that flag is contracted to mean "metadata proved this vendor". The bare
# bypassed gate (conf 0.35) fired on 59.8% of 256 genuine camera captures, so
# `--sensitivity assume-ai` filled a phantom sparkle on ~6 of every 10 clean photos.
weak = self._c("gemini", relaxed=True, relaxed_conf=0.40)
assert reg.decide([weak], reg.Context(sensitivity="assume_ai")) == []
def test_assume_ai_keeps_sparkle_confirmed_by_metadata_below_the_floor(self):
# the floor exists because the vendor is UNKNOWN; once metadata names Google the
# bypass is contract-legal again, so a weak match is still trusted
weak = self._c("gemini", relaxed=True, relaxed_conf=0.40)
ctx = reg.Context(sensitivity="assume_ai", provenance=frozenset({"gemini"}))
assert [d.candidate.key for d in reg.decide([weak], ctx)] == ["gemini"]
def test_assume_ai_is_monotonic_over_strict(self):
# a mark the STRICT gate accepted must never be dropped by the assumed floor:
# assume_ai only ever adds recall
weak_but_strict = self._c("gemini", strict=True, relaxed=True, relaxed_conf=0.40)
fired = reg.decide([weak_but_strict], reg.Context(sensitivity="assume_ai"))
assert [d.candidate.key for d in fired] == ["gemini"]
assert fired[0].relax is False # accepted on the strict verdict, so mask at strict
def test_auto_relaxes_on_provenance(self):
c = [self._c("gemini", relaxed=True)]
assert self._keys(c, reg.Context(provenance=frozenset({"gemini"}))) == {"gemini"}