mirror of
https://github.com/wiltodelta/remove-ai-watermarks.git
synced 2026-08-06 22:18:36 +02:00
Refactor watermark detection and provenance handling
This commit is contained in:
@@ -27,7 +27,7 @@ It does **not** target watermarks that protect someone else's paid or copyrighte
|
||||
|
||||
## Features
|
||||
|
||||
- **Visible watermark removal** — a registry of known marks in their usual places: the Gemini / Nano Banana sparkle, the Doubao "豆包AI生成" text strip, the Jimeng "★ 即梦AI" wordmark, and the Samsung Galaxy AI "✦ Contenuti generati dall'AI" strip (bottom-left, locale-specific). Each mark is **localized to a footprint mask, then filled**: the engine finds the mark, builds a binary mask over its footprint, and one shared, swappable fill inpaints that region. Choose the fill with `--backend`: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN, light, the memory-tight pick where LaMa will not fit), or `lama` (big-LaMa, best quality, heavier, auto-preferred when a learned backend is available); the default `auto` uses LaMa > MI-GAN > cv2, best available. The localizer is cheap CPU (cv2/numpy), so it runs anywhere; the heavier MI-GAN/LaMa fill is opt-in. Detection keys on each mark's own shape (NCC against a captured silhouette; the alpha captures rebuilt by `scripts/visible_alpha_solve.py` are used to detect and to shape the mask, not for pixel recovery). The visual detector needs no metadata, but a borderline (faint or moved) mark is only trusted with corroboration: `--sensitivity` (default `auto`) relaxes a mark's gate when local metadata confirms the vendor or a same-product sibling mark is found; `--sensitivity assume-ai` relaxes every mark on the assertion that the image is AI, recovering the moved or re-rendered marks the conservative gate skips on a metadata-stripped screenshot (`strict` never relaxes). `visible --mark auto` finds and removes every detected mark in one pass. Fast, offline, no GPU. (For arbitrary logos/objects, see `erase`.)
|
||||
- **Visible watermark removal** — a registry of known marks in their usual places: the Gemini / Nano Banana sparkle, the Doubao "豆包AI生成" text strip, the Jimeng "★ 即梦AI" wordmark, and the Samsung Galaxy AI "✦ Contenuti generati dall'AI" strip (bottom-left, locale-specific). Each mark is **localized to a footprint mask, then filled**: the engine finds the mark, builds a binary mask over its footprint, and one shared, swappable fill inpaints that region. Choose the fill with `--backend`: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN, light, the memory-tight pick where LaMa will not fit), or `lama` (big-LaMa, best quality, heavier, auto-preferred when a learned backend is available); the default `auto` uses LaMa > MI-GAN > cv2, best available. The localizer is cheap CPU (cv2/numpy), so it runs anywhere; the heavier MI-GAN/LaMa fill is opt-in. Detection keys on each mark's own shape (NCC against a captured silhouette; the alpha captures rebuilt by `scripts/visible_alpha_solve.py` are used to detect and to shape the mask, not for pixel recovery). The visual detector needs no metadata, but a borderline (faint or moved) mark is only trusted with corroboration: `--sensitivity` (default `auto`) relaxes a mark's gate when local metadata confirms the vendor or a same-product sibling mark is found; `--sensitivity assume-ai` relaxes every mark on the assertion that the image is AI, recovering the moved or re-rendered marks the conservative gate skips on a metadata-stripped screenshot (`strict` never relaxes). Because asserting "this is AI" says nothing about *which* vendor made it, a mark relaxed on that assertion alone still has to clear a confidence floor, so a clean photo is left untouched. `visible --mark auto` finds and removes every detected mark in one pass. Fast, offline, no GPU. (For arbitrary logos/objects, see `erase`.)
|
||||
- **Universal region eraser (`erase`)** — remove any logo / watermark / object inside boxes you specify, regardless of position or color. Default cv2 inpainting (CPU, instant); optional big-LaMa via onnxruntime (`lama` extra) for higher quality
|
||||
- **Invisible watermark removal** — SynthID, StableSignature, TreeRing via diffusion-based regeneration (needs a local GPU, or run it with no setup on [raiw.cc](https://raiw.cc))
|
||||
- **AI metadata stripping** — EXIF, PNG text chunks, C2PA provenance manifests (PNG / JPEG / AVIF / HEIF / JPEG-XL, **MP4 / MOV / M4V / M4A** at the container level, and **WebM / MP3 / WAV / FLAC / OGG** losslessly via ffmpeg), XMP DigitalSourceType
|
||||
|
||||
@@ -61,11 +61,22 @@ module.
|
||||
|
||||
`watermark_registry.py` — **single catalog of known visible watermarks**, the unified "find known marks in their usual places, recognize, remove" entry.
|
||||
|
||||
**Localize -> fill by policy (replaced reverse-alpha):** each mark is localized to a binary full-frame footprint mask (a `Localization`), and one shared, swappable fill inpaints that mask via `fill(image, mask, backend=...)` (delegates to `region_eraser.erase`). This replaced the old reverse-alpha removal (invert a captured alpha map, `original = (wm - a*logo)/(1-a)`, plus a thin residual inpaint) for ALL marks — gemini, doubao, jimeng, samsung, and jimeng_pill. **Why it changed:** reverse-alpha depended on a fixed captured alpha map at a fixed position, so it broke whenever a vendor moved or re-rendered its mark; and it was not color-lossless even with the right map (it amplifies 8-bit quantization and JPEG-chroma error by `1/(1-a)`), which showed up as "the color just changed, not removed" reports. Localize -> fill has a benign failure mode: a slightly-off localization just inpaints a small region near-losslessly instead of leaving a color-shifted smear. The captured alpha maps are still used to DETECT the marks and to shape the mask (gemini's footprint), but NOT for pixel recovery. Fill backends: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN ONNX, light, the memory-tight pick where LaMa will not fit), `lama` (big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available); `auto` = LaMa > MI-GAN > cv2, best available. Each `KnownMark` ties a key to {usual `location`, `in_auto` flag, a `_detect` callable → uniform `MarkDetection`, a `_mask` callable → full-frame footprint mask}; `KnownMark.remove(image, *, backend="auto", provenance=False, force=False)`. Entries today: `gemini` (bottom-right sparkle), `doubao` (bottom-right "豆包AI生成"), `jimeng` (bottom-right "★ 即梦AI"), `samsung` (bottom-**LEFT** "✦ Contenuti generati dall'AI", Samsung Galaxy AI, Italian locale), and the capture-less `jimeng_pill` (top-left "AI生成"). `detect_marks(image, *, provenance=frozenset())` scans all (strict, for the identify verdict); `remove_auto_marks(image, *, sensitivity="auto", provenance=frozenset(), backend="auto")` removes every detected mark in one pass. **Sensitivity (`auto`/`strict`/`assume_ai`)** decides how hard a borderline mark is trusted: the visual detectors are pixel-based (no metadata needed) and the recall gain comes from relaxing the false-positive gate, not from metadata. `resolve_relax` turns the policy + evidence into the per-mark relaxation boolean — `strict` never relaxes, `assume_ai` always relaxes (the caller asserts AI, e.g. a metadata-stripped screenshot; ~46% -> ~92% Gemini recall, at the cost of a small near-lossless fill on some clean corners), `auto` relaxes only on same-product evidence (metadata provenance for that vendor, or a confidently strict-detected sibling of the same `_PRODUCT_OF` — Doubao and Jimeng are both bottom-right ByteDance but distinct products, so they do NOT cross-relax). **Perception / decision / action are separated three ways** (the removal path only; `identify` keeps calling `KnownMark.detect` directly, so its verdict is untouched): `_build_candidates(image)` is PERCEPTION — it runs each detector at both trust levels and packages raw verdicts + the pill's flatness feature into `Candidate`s, no policy; `decide(candidates, Context(sensitivity, provenance)) -> [Decision]` is the pure DECISION arbiter — all keep/drop policy (`resolve_relax` cross-mark corroboration + the pill gate) in one image-free, unit-testable function (`tests/test_watermark_registry.py::TestArbiter`); then `remove_auto_marks` does the ACTION, localizing -> filling each winner. The Gemini FP gate deliberately stays inside `gemini_engine` (not the arbiter) because `identify` reads that same gated confidence — pulling it out would drift the provenance verdict. Behavior is byte-identical to the pre-arbiter two-pass (corpus-verified: strict/auto 46%, assume_ai 92% Gemini recall unchanged).
|
||||
**Localize -> fill by policy (replaced reverse-alpha):** each mark is localized to a binary full-frame footprint mask (a `Localization`), and one shared, swappable fill inpaints that mask via `fill(image, mask, backend=...)` (delegates to `region_eraser.erase`). This replaced the old reverse-alpha removal (invert a captured alpha map, `original = (wm - a*logo)/(1-a)`, plus a thin residual inpaint) for ALL marks — gemini, doubao, jimeng, samsung, and jimeng_pill. **Why it changed:** reverse-alpha depended on a fixed captured alpha map at a fixed position, so it broke whenever a vendor moved or re-rendered its mark; and it was not color-lossless even with the right map (it amplifies 8-bit quantization and JPEG-chroma error by `1/(1-a)`), which showed up as "the color just changed, not removed" reports. Localize -> fill has a benign failure mode: a slightly-off localization just inpaints a small region near-losslessly instead of leaving a color-shifted smear. The captured alpha maps are still used to DETECT the marks and to shape the mask (gemini's footprint), but NOT for pixel recovery. Fill backends: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN ONNX, light, the memory-tight pick where LaMa will not fit), `lama` (big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available); `auto` = LaMa > MI-GAN > cv2, best available. Each `KnownMark` ties a key to {usual `location`, `in_auto` flag, a `_detect` callable → uniform `MarkDetection`, a `_mask` callable → full-frame footprint mask}; `KnownMark.remove(image, *, backend="auto", provenance=False, force=False)`. Entries today: `gemini` (bottom-right sparkle), `doubao` (bottom-right "豆包AI生成"), `jimeng` (bottom-right "★ 即梦AI"), `samsung` (bottom-**LEFT** "✦ Contenuti generati dall'AI", Samsung Galaxy AI, Italian locale), and the capture-less `jimeng_pill` (top-left "AI生成"). `detect_marks(image, *, provenance=frozenset())` scans all (strict, for the identify verdict); `remove_auto_marks(image, *, sensitivity="auto", provenance=frozenset(), backend="auto")` removes every detected mark in one pass. **Sensitivity (`auto`/`strict`/`assume_ai`)** decides how hard a borderline mark is trusted: the visual detectors are pixel-based (no metadata needed) and the recall gain comes from relaxing the false-positive gate, not from metadata. `resolve_trust` turns the policy + evidence into the per-mark trust level the engines consume as `provenance = level != "strict"` — `strict` never relaxes; `auto` relaxes only on same-product evidence (metadata provenance for that vendor, or a confidently strict-detected sibling of the same `_PRODUCT_OF` — Doubao and Jimeng are both bottom-right ByteDance but distinct products, so they do NOT cross-relax); `assume_ai` relaxes every mark (the caller asserts AI, e.g. a metadata-stripped screenshot). **Three levels, not two: `strict` / `assumed` / `confirmed`.** Relaxing bypasses the engine's false-positive gate outright, and that bypass is contracted to mean the vendor is CONFIRMED (`GeminiEngine.detect_watermark`'s `trust_provenance`: "external metadata already proves this is a Google generation"). An `assume_ai` caller asserts the image is AI, which says nothing about WHICH vendor, so a mark relaxed on assumption alone must also clear `_ASSUMED_CONF_FLOOR` (`assumed_floor_ok`; gemini 0.50) — see "Assumed-trust confidence floor" below. **Perception / decision / action are separated three ways** (the removal path only; `identify` keeps calling `KnownMark.detect` directly, so its verdict is untouched): `_build_candidates(image)` is PERCEPTION — it runs each detector at both trust levels and packages raw verdicts + the pill's flatness feature into `Candidate`s, no policy; `decide(candidates, Context(sensitivity, provenance)) -> [Decision]` is the pure DECISION arbiter — all keep/drop policy (`resolve_trust` cross-mark corroboration + the assumed-trust floor + the pill gate) in one image-free, unit-testable function (`tests/test_watermark_registry.py::TestArbiter`); then `remove_auto_marks` does the ACTION, localizing -> filling each winner. The Gemini FP gate deliberately stays inside `gemini_engine` (not the arbiter) because `identify` reads that same gated confidence — pulling it out would drift the provenance verdict. Behavior was byte-identical to the pre-arbiter two-pass when the arbiter landed; `assume_ai` has since gained the assumed-trust confidence floor (see below), which deliberately changes its verdict on weak gate-bypassed matches.
|
||||
|
||||
**Head-to-head validation (v0.12.1 reverse-alpha vs the current localize -> fill):** run over the full labelled visible-mark set, with the cv2 / MI-GAN / LaMa fills each compared against the old reverse-alpha. **doubao and jimeng are identical** across every backend -- 100% coverage and 100% clearance either way. **gemini** strict coverage is a few points below reverse-alpha's (the deliberate false-positive tightening), but the metadata-stripped faint ones are now mostly recovered by the DEFAULT white-core rescue in the FP gate (`gemini_engine`: a bright near-WHITE core distinguishes a real faint sparkle from a colored bright corner -- ~14/20 recovered at ~1.25% clean false-fire; a learned classifier on the same features measured worse, 2026-07 tier-1), the residual under `assume_ai`; clearance is equal (~98% both), and neither version touches pixels outside the mark box (outside-box PSNR ~99). **Clearance is fill-independent** -- cv2, MI-GAN and LaMa all strip the mark's shape equally, so the re-detect metric does not separate them; the difference is purely the *visual fill quality* on the recovered region, and it is background-dependent. reverse-alpha recovered textured and especially regular/structured backgrounds (a lattice, a grid) more cleanly than any inpaint; **LaMa closes most of that gap** (the best learned backend), **MI-GAN can ghost or hallucinate structure**, and **cv2 smears** (the last-resort floor). This is why `auto` resolves `LaMa > MI-GAN > cv2` (`preferred_inpaint_backend`) and warns once on the cv2 fallback; on flat backgrounds every backend is clean.
|
||||
|
||||
**Provenance prior:** when local metadata already confirms the vendor, the mark's detection trust gate is relaxed (a confirmed vendor means the mark is present with high prior, so a mark the conservative detector would demote as a content false positive is trusted). `detect_marks` / `remove_auto_marks` take a `provenance` frozenset and `KnownMark.remove` a `provenance` flag. Mapping: a Google/Gemini C2PA issuer relaxes gemini (skips its false-positive gate and lowers the trust threshold from 0.5 to 0.35); a China-AIGC (TC260) label relaxes doubao/jimeng; `samsung_genai` relaxes samsung. Corpus finding: on Google-C2PA images, Gemini sparkle recall rose from ~46% (plain detector) to ~90% with the provenance prior (recovering marks the vendor moved or re-rendered). The localizer is cheap CPU (cv2/numpy), so a memory-tight caller runs it anywhere; the heavy MI-GAN/LaMa fill is opt-in and chosen by the caller.
|
||||
**Assumed-trust confidence floor (`_ASSUMED_CONF_FLOOR` / `assumed_floor_ok`, 2026-07-16):** relaxing a mark bypasses the engine's false-positive gate ENTIRELY, leaving only the bare detector threshold (gemini: fused confidence >= 0.35). That is defensible when metadata names the vendor (`confirmed`) and indefensible on a bare "assume this is AI" (`assumed`), because the assertion carries no vendor information. Measured on 256 genuine camera captures (Make/Model/exposure/aperture present, no AI token -- a Gemini sparkle cannot be there) vs 697 Google-C2PA positives with the metadata used only as a label:
|
||||
|
||||
| bypassed threshold | recall | false fire on clean photos |
|
||||
|---|---|---|
|
||||
| 0.35 (the bare gate) | 82.6% | **59.8%** |
|
||||
| 0.45 | 66.6% | 12.5% |
|
||||
| **0.50 (chosen)** | 59.4% | **0.0%** |
|
||||
| strict gate | 56.4% | 0.0% |
|
||||
|
||||
0.35 sat on a cliff: +26pp recall over strict bought by filling a corner on ~6 of every 10 CLEAN photos, and `api.remove_visible(sensitivity="assume_ai")` reproduced it on 8/15 of the committed verified-clean negatives. The floor is applied in the arbiter and is **monotonic over strict** -- a mark the strict gate accepted is never dropped by it, so `assume_ai` only ever adds recall. End-to-end after the fix (400 Google-C2PA positives with metadata hidden from the detector, 256 camera negatives, through the public `api.remove_visible`): recall strict 55.0% / auto 55.2% / assume_ai 62.8%; false fire 0.0% / 0.0% / 2.3%. The residual 2.3% contains **no gemini at all** (doubao 2, jimeng_pill 2, jimeng 1, samsung 1 of 256) -- the text marks' own relaxed gates (<1% each) plus the pill's flat-footprint arm, all pre-existing and benign. The superseded "~46% -> ~92%" figure measured recall only, on Google-C2PA files where the answer was always Google; false fire on non-Google content was never measured. Marks other than gemini carry no floor because their bypassed false-fire is already under 1%. Regression: `tests/test_watermark_registry.py::TestArbiter::{test_assume_ai_drops_sparkle_below_the_assumed_floor, test_assume_ai_keeps_sparkle_confirmed_by_metadata_below_the_floor, test_assume_ai_is_monotonic_over_strict}`.
|
||||
|
||||
**Provenance prior:** when local metadata already confirms the vendor, the mark's detection trust gate is relaxed (a confirmed vendor means the mark is present with high prior, so a mark the conservative detector would demote as a content false positive is trusted). `detect_marks` / `remove_auto_marks` take a `provenance` frozenset and `KnownMark.remove` a `provenance` flag. Mapping: a Google/Gemini C2PA issuer relaxes gemini (skips its false-positive gate and lowers the trust threshold from 0.5 to 0.35); a China-AIGC (TC260) label relaxes doubao/jimeng; `samsung_genai` relaxes samsung. Corpus finding: on Google-C2PA images, Gemini sparkle recall rose from ~46% (plain detector) to ~90% with the provenance prior (recovering marks the vendor moved or re-rendered). That gain is why the bypass exists, and it is conditional on the metadata actually naming the vendor — a caller merely ASSUMING the image is AI does not get it unconditionally (see the assumed-trust confidence floor above). The localizer is cheap CPU (cv2/numpy), so a memory-tight caller runs it anywhere; the heavy MI-GAN/LaMa fill is opt-in and chosen by the caller.
|
||||
|
||||
**Cross-engine confidences aren't directly comparable**, so the gemini adapter applies the corpus-validated 0.5 sparkle threshold (`_GEMINI_AUTO_MIN_CONF`) for its `detected` flag (lowered to 0.35 under the Google/Gemini provenance prior) — otherwise the gemini engine's loose internal threshold weakly fires (~0.36) on the Doubao text and hijacks `auto`. The shape-keyed Doubao/Jimeng/Samsung NCC detectors don't cross-fire (jimeng scores ~0.22 on the Doubao strip, well under its 0.45 threshold; Samsung is bottom-left so it shares no corner with the others, and scored 0.0 on Doubao/Jimeng captures and they 0.0 on a real Samsung photo), so `auto` picks the right one. `cli.cmd_visible` is registry-driven: `--mark auto` → `remove_auto_marks` (removes every detected mark), `--mark <key>` → that mark; `--mark` choices come from `mark_keys()`.
|
||||
|
||||
@@ -230,7 +241,7 @@ Diffusion SynthID removal. The `--tile/--no-tile` knob is the *lossless* alterna
|
||||
|
||||
### `visible`
|
||||
|
||||
Known-visible-mark removal by **localize -> fill**: each detected mark is localized to a binary full-frame footprint mask, then one shared, swappable fill inpaints that mask. `--backend auto|cv2|migan|lama` (default `auto`) picks the fill: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN ONNX, the memory-tight pick where LaMa will not fit), `lama` (big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available); `auto` = LaMa > MI-GAN > cv2, best available (LaMa is auto-preferred when a learned backend is present; a memory-tight deploy pins migan). `--sensitivity auto|strict|assume-ai` (default `auto`) controls how hard a borderline mark is trusted (see the registry section: the visual detectors are metadata-independent; `auto` relaxes a mark only on same-product evidence, `assume-ai` relaxes every mark on the caller's AI assertion — the only path to high recall on a metadata-stripped screenshot). `--backend` and `--sensitivity` are shared across `visible`/`all`/`batch`. Detection keys on each mark's own shape, and under `auto` the trust gate is relaxed when local metadata confirms the vendor (a Google/Gemini C2PA issuer relaxes gemini, a China-AIGC label relaxes doubao/jimeng, `samsung_genai` relaxes samsung), so a moved or re-rendered mark is still caught. `--mark auto` (default) removes EVERY detected mark in one pass (`registry.remove_auto_marks`, not the single strongest -- a Jimeng-basic image carries both the top-left pill and the bottom-right wordmark) from: the Gemini sparkle, the Doubao "豆包AI生成" text strip, the Jimeng "★ 即梦AI" wordmark, the Samsung Galaxy AI "✦ Contenuti generati dall'AI" strip (bottom-LEFT, Italian-locale detection), and the capture-less Jimeng "AI生成" pill (top-left, `pill_engine`). The pill's weak edge-NCC detector is gated in `remove_auto_marks` via `_keep_pill` (32k real-upload corpus validation 2026-07): never on Doubao, and two confirmation arms since metadata confirms the platform, not pill presence. (1) The bottom-right wordmark fired — ~94% precise and survives metadata-STRIPPED uploads (screenshots / re-saves) — removes the pill unrestricted. (2) TC260 metadata confirms Jimeng (`"jimeng" in provenance`, from `cli._visible_provenance`) OR the caller asserts AI (`sensitivity == "assume_ai"`), no wordmark — ~27% precise, its false fires are textured ceilings/walls that the fill visibly SMEARS — removes the pill ONLY when the top-left footprint is flat enough for an invisible fill (`pill_engine.footprint_is_flat`, median-Sobel ≤ `_FLAT_TEXTURE_MAX`; the flatness guard holds even under `assume_ai`). No confirmation → never removed. `--mark gemini|doubao|jimeng|samsung|jimeng_pill` forces one (choices come from the registry). Corpus validation: doubao and jimeng localize + remove at ~100% with clean footprints (the filled region blends into its surroundings within a few LAB levels, no color shift, no dark pit); clean images with no vendor signature had 0% false removal. For arbitrary logos/objects use `erase`. **When `--mark auto` finds no known mark (the common case — ~74% of real uploads carry no registered visible mark), the command does NOT silently re-serve the input as a finished result.** It runs a cheap metadata-only `identify`, prints actionable guidance (if the image carries an invisible/metadata mark, e.g. an OpenAI/Gemini C2PA image, it points to `all`; otherwise it does NOT imply the image is clean -- it warns that an invisible pixel watermark like SynthID cannot be detected once the metadata proxy is gone and routes to both `all` and `erase --region`), writes NO output file, and exits **`EXIT_NO_VISIBLE_MARK` (2)** — distinct from success (0) and a hard error (1) so a wrapping service (raiw.cc) can surface the message instead of treating the unchanged image as done (the production "it didn't work" / score-0 trap). Same handling for an explicit `--mark <name>` that is not detected. Helper `cli._no_visible_mark_exit`; regression-guarded by `tests/test_cli.py::TestVisibleCommand::test_visible_auto_no_mark_exits_two_with_eraser_hint` and `test_visible_auto_no_mark_routes_to_all_when_metadata`. `--no-detect` still forces the gemini fallback and proceeds (exit 0).
|
||||
Known-visible-mark removal by **localize -> fill**: each detected mark is localized to a binary full-frame footprint mask, then one shared, swappable fill inpaints that mask. `--backend auto|cv2|migan|lama` (default `auto`) picks the fill: `cv2` (classical inpaint, no deps, the floor), `migan` (MI-GAN ONNX, the memory-tight pick where LaMa will not fit), `lama` (big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available); `auto` = LaMa > MI-GAN > cv2, best available (LaMa is auto-preferred when a learned backend is present; a memory-tight deploy pins migan). `--sensitivity auto|strict|assume-ai` (default `auto`) controls how hard a borderline mark is trusted (see the registry section: the visual detectors are metadata-independent; `auto` relaxes a mark only on same-product evidence, `assume-ai` relaxes every mark on the caller's AI assertion, subject to the assumed-trust confidence floor where the vendor is unconfirmed — the only path to higher recall on a metadata-stripped screenshot). `--backend` and `--sensitivity` are shared across `visible`/`all`/`batch`. Detection keys on each mark's own shape, and under `auto` the trust gate is relaxed when local metadata confirms the vendor (a Google/Gemini C2PA issuer relaxes gemini, a China-AIGC label relaxes doubao/jimeng, `samsung_genai` relaxes samsung), so a moved or re-rendered mark is still caught. `--mark auto` (default) removes EVERY detected mark in one pass (`registry.remove_auto_marks`, not the single strongest -- a Jimeng-basic image carries both the top-left pill and the bottom-right wordmark) from: the Gemini sparkle, the Doubao "豆包AI生成" text strip, the Jimeng "★ 即梦AI" wordmark, the Samsung Galaxy AI "✦ Contenuti generati dall'AI" strip (bottom-LEFT, Italian-locale detection), and the capture-less Jimeng "AI生成" pill (top-left, `pill_engine`). The pill's weak edge-NCC detector is gated in `remove_auto_marks` via `_keep_pill` (32k real-upload corpus validation 2026-07): never on Doubao, and two confirmation arms since metadata confirms the platform, not pill presence. (1) The bottom-right wordmark fired — ~94% precise and survives metadata-STRIPPED uploads (screenshots / re-saves) — removes the pill unrestricted. (2) TC260 metadata confirms Jimeng (`"jimeng" in provenance`, from `cli._visible_provenance`) OR the caller asserts AI (`sensitivity == "assume_ai"`), no wordmark — ~27% precise, its false fires are textured ceilings/walls that the fill visibly SMEARS — removes the pill ONLY when the top-left footprint is flat enough for an invisible fill (`pill_engine.footprint_is_flat`, median-Sobel ≤ `_FLAT_TEXTURE_MAX`; the flatness guard holds even under `assume_ai`). No confirmation → never removed. `--mark gemini|doubao|jimeng|samsung|jimeng_pill` forces one (choices come from the registry). Corpus validation: doubao and jimeng localize + remove at ~100% with clean footprints (the filled region blends into its surroundings within a few LAB levels, no color shift, no dark pit); clean images with no vendor signature had 0% false removal. For arbitrary logos/objects use `erase`. **When `--mark auto` finds no known mark (the common case — ~74% of real uploads carry no registered visible mark), the command does NOT silently re-serve the input as a finished result.** It runs a cheap metadata-only `identify`, prints actionable guidance (if the image carries an invisible/metadata mark, e.g. an OpenAI/Gemini C2PA image, it points to `all`; otherwise it does NOT imply the image is clean -- it warns that an invisible pixel watermark like SynthID cannot be detected once the metadata proxy is gone and routes to both `all` and `erase --region`), writes NO output file, and exits **`EXIT_NO_VISIBLE_MARK` (2)** — distinct from success (0) and a hard error (1) so a wrapping service (raiw.cc) can surface the message instead of treating the unchanged image as done (the production "it didn't work" / score-0 trap). Same handling for an explicit `--mark <name>` that is not detected. Helper `cli._no_visible_mark_exit`; regression-guarded by `tests/test_cli.py::TestVisibleCommand::test_visible_auto_no_mark_exits_two_with_eraser_hint` and `test_visible_auto_no_mark_routes_to_all_when_metadata`. `--no-detect` still forces the gemini fallback and proceeds (exit 0).
|
||||
|
||||
### `batch`
|
||||
|
||||
|
||||
@@ -16,6 +16,7 @@ Imports stay lazy (inside the functions), so ``import remove_ai_watermarks`` is
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import TYPE_CHECKING, Any
|
||||
|
||||
@@ -25,6 +26,16 @@ if TYPE_CHECKING:
|
||||
from remove_ai_watermarks.watermark_registry import Backend, Sensitivity
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class _VisibleInput:
|
||||
"""Normalized visible-removal input with its file-only context."""
|
||||
|
||||
bgr: NDArray[Any]
|
||||
alpha: NDArray[Any] | None = None
|
||||
path: Path | None = None
|
||||
provenance: frozenset[str] = frozenset()
|
||||
|
||||
|
||||
def visible_provenance(source: str | Path) -> frozenset[str]:
|
||||
"""Vendor keys that the file's local metadata confirms, the evidence that drives
|
||||
the ``auto`` sensitivity (relaxing a corroborated mark's detection trust gate).
|
||||
@@ -36,19 +47,70 @@ def visible_provenance(source: str | Path) -> frozenset[str]:
|
||||
"""
|
||||
import contextlib
|
||||
|
||||
keys: set[str] = set()
|
||||
path = Path(source)
|
||||
with contextlib.suppress(Exception):
|
||||
from remove_ai_watermarks import identify, metadata
|
||||
from remove_ai_watermarks import identify
|
||||
|
||||
rep = identify.identify(Path(source), check_visible=False, check_invisible=False)
|
||||
rep = identify.identify(path, check_visible=False, check_invisible=False)
|
||||
signal_names = {signal.name for signal in rep.signals}
|
||||
keys: set[str] = set()
|
||||
platform = (rep.platform or "").lower()
|
||||
if "google" in platform or "gemini" in platform:
|
||||
keys.add("gemini")
|
||||
if metadata.aigc_label(Path(source)):
|
||||
if "aigc" in signal_names:
|
||||
keys |= {"doubao", "jimeng"}
|
||||
if metadata.samsung_genai(Path(source)):
|
||||
if "samsung_genai" in signal_names:
|
||||
keys.add("samsung")
|
||||
return frozenset(keys)
|
||||
return frozenset(keys)
|
||||
return frozenset()
|
||||
|
||||
|
||||
def _load_visible_input(source: str | Path | NDArray[Any]) -> _VisibleInput:
|
||||
"""Normalize a path/array source without making the public operation stateful."""
|
||||
if not isinstance(source, (str, Path)):
|
||||
return _VisibleInput(source)
|
||||
|
||||
from remove_ai_watermarks import image_io
|
||||
|
||||
path = Path(source)
|
||||
bgr, alpha = image_io.read_bgr_and_alpha(path)
|
||||
if bgr is None:
|
||||
raise ValueError(f"Could not read image: {source}")
|
||||
return _VisibleInput(bgr=bgr, alpha=alpha, path=path, provenance=visible_provenance(path))
|
||||
|
||||
|
||||
def _write_visible_result(
|
||||
loaded: _VisibleInput,
|
||||
result: NDArray[Any],
|
||||
removed: list[str],
|
||||
output: str | Path,
|
||||
*,
|
||||
strip_metadata: bool,
|
||||
write_noop: bool,
|
||||
) -> None:
|
||||
"""Write one visible-removal result while preserving a true no-op losslessly."""
|
||||
if not removed and not write_noop:
|
||||
return
|
||||
|
||||
from remove_ai_watermarks import image_io
|
||||
|
||||
out_path = Path(output)
|
||||
out_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
source_path = loaded.path
|
||||
if not removed and source_path is not None and source_path.suffix.lower() == out_path.suffix.lower():
|
||||
# Copy the ORIGINAL bytes instead of lossily re-encoding a no-op. An in-place
|
||||
# call needs no copy and would otherwise raise shutil.SameFileError.
|
||||
if source_path.resolve() != out_path.resolve():
|
||||
import shutil
|
||||
|
||||
shutil.copyfile(source_path, out_path)
|
||||
else:
|
||||
image_io.write_bgr_with_alpha(out_path, result, loaded.alpha)
|
||||
|
||||
if strip_metadata:
|
||||
from remove_ai_watermarks import metadata
|
||||
|
||||
metadata.remove_ai_metadata(out_path, out_path)
|
||||
|
||||
|
||||
def remove_visible(
|
||||
@@ -88,40 +150,22 @@ def remove_visible(
|
||||
output path untouched, so a caller that treats "no mark" as "produce nothing" (the CLI
|
||||
``visible`` no-mark contract) does not clobber a pre-existing file at that path.
|
||||
"""
|
||||
from remove_ai_watermarks import image_io, watermark_registry
|
||||
|
||||
alpha: NDArray[Any] | None = None
|
||||
provenance: frozenset[str] = frozenset()
|
||||
if isinstance(source, (str, Path)):
|
||||
path = Path(source)
|
||||
bgr, alpha = image_io.read_bgr_and_alpha(path)
|
||||
if bgr is None:
|
||||
raise ValueError(f"Could not read image: {source}")
|
||||
provenance = visible_provenance(path)
|
||||
else:
|
||||
bgr = source
|
||||
from remove_ai_watermarks import watermark_registry
|
||||
|
||||
loaded = _load_visible_input(source)
|
||||
result, removed = watermark_registry.remove_auto_marks(
|
||||
bgr, sensitivity=sensitivity, provenance=provenance, backend=backend
|
||||
loaded.bgr,
|
||||
sensitivity=sensitivity,
|
||||
provenance=loaded.provenance,
|
||||
backend=backend,
|
||||
)
|
||||
if output is not None and (removed or write_noop):
|
||||
out_path = Path(output)
|
||||
out_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
same_format = isinstance(source, (str, Path)) and Path(source).suffix.lower() == out_path.suffix.lower()
|
||||
if not removed and same_format:
|
||||
# Nothing was removed: copy the ORIGINAL bytes verbatim instead of a lossy
|
||||
# re-encode of its decode, so the pixels stay bit-identical (the metadata
|
||||
# strip below is lossless, so it does not disturb them either). Skip the copy
|
||||
# for an in-place call (output == source): the bytes are already there, and
|
||||
# shutil.copyfile would raise SameFileError.
|
||||
if Path(source).resolve() != out_path.resolve(): # type: ignore[arg-type]
|
||||
import shutil
|
||||
|
||||
shutil.copyfile(source, out_path) # type: ignore[arg-type]
|
||||
else:
|
||||
image_io.write_bgr_with_alpha(out_path, result, alpha)
|
||||
if strip_metadata:
|
||||
from remove_ai_watermarks import metadata
|
||||
|
||||
metadata.remove_ai_metadata(out_path, out_path)
|
||||
if output is not None:
|
||||
_write_visible_result(
|
||||
loaded,
|
||||
result,
|
||||
removed,
|
||||
output,
|
||||
strip_metadata=strip_metadata,
|
||||
write_noop=write_noop,
|
||||
)
|
||||
return result, removed
|
||||
|
||||
+252
-180
@@ -12,6 +12,7 @@ import contextlib
|
||||
import json
|
||||
import logging
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import TYPE_CHECKING, Any, Literal, NoReturn
|
||||
|
||||
@@ -311,8 +312,9 @@ _visible_sensitivity_option = click.option(
|
||||
help="How hard to trust a borderline mark. auto: relax a mark only when metadata "
|
||||
"or a same-product sibling mark corroborates it (safe; clean images untouched). "
|
||||
"strict: high-precision visual gate only, never relaxed. assume-ai: treat the "
|
||||
"image as AI and relax every mark (best recall on metadata-stripped screenshots, "
|
||||
"at the cost of a small fill on some clean corners).",
|
||||
"image as AI and relax every mark, keeping a confidence floor where the vendor is "
|
||||
"unconfirmed (best recall on metadata-stripped screenshots; a clean image is still "
|
||||
"left untouched).",
|
||||
)
|
||||
|
||||
|
||||
@@ -525,6 +527,113 @@ def main(ctx: click.Context, verbose: bool) -> None:
|
||||
|
||||
|
||||
# ── Visible (Gemini) watermark removal ──
|
||||
def _run_visible_auto(
|
||||
source: Path,
|
||||
output: Path,
|
||||
*,
|
||||
backend: watermark_registry.Backend,
|
||||
sensitivity: watermark_registry.Sensitivity,
|
||||
strip_metadata: bool,
|
||||
) -> None:
|
||||
"""Run the registry-wide visible pass and render its CLI result."""
|
||||
from remove_ai_watermarks import api
|
||||
|
||||
t0 = time.monotonic()
|
||||
try:
|
||||
with console.status("Detecting & removing visible marks..."):
|
||||
result, removed = api.remove_visible(
|
||||
str(source),
|
||||
str(output),
|
||||
sensitivity=sensitivity,
|
||||
backend=backend,
|
||||
strip_metadata=strip_metadata,
|
||||
write_noop=False,
|
||||
)
|
||||
except RuntimeError as e: # selected migan/lama backend whose extra is absent
|
||||
console.print(f" Error: {e}")
|
||||
raise SystemExit(1) from e
|
||||
except (ValueError, OSError) as e: # unreadable / truncated / non-image input
|
||||
console.print(f" Error: cannot read image {source.name}: {e}")
|
||||
raise SystemExit(1) from e
|
||||
|
||||
elapsed = time.monotonic() - t0
|
||||
h, w = result.shape[:2]
|
||||
console.print(f" Input: {source.name} ({w}x{h})")
|
||||
if not removed:
|
||||
# write_noop=False means nothing was written, so a pre-existing output is intact.
|
||||
console.print(" No known visible mark detected (gemini / doubao / jimeng / jimeng-pill / samsung).")
|
||||
_no_visible_mark_exit(source, sensitivity=sensitivity)
|
||||
console.print(f" Removed: {', '.join(removed)}")
|
||||
size_kb = output.stat().st_size / 1024
|
||||
console.print(f" Saved: {output} ({size_kb:.0f} KB, {elapsed:.2f}s)")
|
||||
|
||||
|
||||
def _run_visible_explicit(
|
||||
ctx: click.Context,
|
||||
source: Path,
|
||||
output: Path,
|
||||
*,
|
||||
detect: bool,
|
||||
mark: str,
|
||||
backend: watermark_registry.Backend,
|
||||
sensitivity: watermark_registry.Sensitivity,
|
||||
resolved_backend: str,
|
||||
strip_metadata: bool,
|
||||
) -> None:
|
||||
"""Run one explicitly selected visible-mark detector/remover."""
|
||||
image, alpha = image_io.read_bgr_and_alpha(source)
|
||||
if image is None:
|
||||
console.print(f"Error: Failed to read image: {source}")
|
||||
raise SystemExit(1)
|
||||
h, w = image.shape[:2]
|
||||
console.print(f" Input: {source.name} ({w}x{h})")
|
||||
|
||||
provenance = _visible_provenance(source)
|
||||
target = "gemini" if mark == "auto" else mark # --no-detect auto: gemini fallback
|
||||
chosen = watermark_registry.get_mark(target)
|
||||
# A single explicit mark has no sibling corroboration. Keep its trust resolution
|
||||
# aligned with the registry arbiter, including the assumption-only floor.
|
||||
trust = watermark_registry.resolve_trust(
|
||||
chosen.key,
|
||||
sensitivity=sensitivity,
|
||||
provenance=provenance,
|
||||
strict_keys=set(),
|
||||
)
|
||||
relax = trust != "strict"
|
||||
detection = chosen.detect(image, provenance=relax)
|
||||
if trust == "assumed" and not watermark_registry.assumed_floor_ok(chosen.key, detection.confidence):
|
||||
relax = False
|
||||
detection = chosen.detect(image, provenance=False)
|
||||
if detect and not detection.detected:
|
||||
console.print(f" {chosen.label} not detected (conf {detection.confidence:.2f}). Use --no-detect to force.")
|
||||
_no_visible_mark_exit(source, sensitivity=sensitivity)
|
||||
if detection.detected:
|
||||
console.print(f" {chosen.label} detected ({chosen.location}, conf {detection.confidence:.2f})")
|
||||
|
||||
t0 = time.monotonic()
|
||||
try:
|
||||
with console.status(f"Removing {chosen.label}... ({resolved_backend})"):
|
||||
result, _ = chosen.remove(image, backend=backend, provenance=relax, force=not detect)
|
||||
except RuntimeError as e: # selected migan/lama backend whose extra is absent
|
||||
console.print(f" Error: {e}")
|
||||
raise SystemExit(1) from e
|
||||
elapsed = time.monotonic() - t0
|
||||
|
||||
output.parent.mkdir(parents=True, exist_ok=True)
|
||||
image_io.write_bgr_with_alpha(output, result, alpha)
|
||||
if strip_metadata:
|
||||
try:
|
||||
from remove_ai_watermarks.metadata import remove_ai_metadata
|
||||
|
||||
remove_ai_metadata(output, output)
|
||||
except Exception as e:
|
||||
if ctx.obj.get("verbose"):
|
||||
console.print(f" Warning: Failed to strip metadata: {e}")
|
||||
|
||||
size_kb = output.stat().st_size / 1024
|
||||
console.print(f" Saved: {output} ({size_kb:.0f} KB, {elapsed:.2f}s)")
|
||||
|
||||
|
||||
@main.command("visible")
|
||||
@click.argument("source", type=click.Path(exists=True, path_type=Path))
|
||||
@click.option(
|
||||
@@ -560,18 +669,16 @@ def cmd_visible(
|
||||
MI-GAN > cv2). ``--mark auto`` removes every detected mark in one
|
||||
pass. For arbitrary logos/objects, use ``erase``.
|
||||
"""
|
||||
from remove_ai_watermarks import watermark_registry as registry
|
||||
|
||||
_banner()
|
||||
source = _validate_image(source)
|
||||
|
||||
if output is None:
|
||||
output = source.with_stem(source.stem + "_clean")
|
||||
|
||||
bk: registry.Backend = backend # type: ignore[assignment]
|
||||
bk: watermark_registry.Backend = backend # type: ignore[assignment]
|
||||
sens = _parse_sensitivity(sensitivity)
|
||||
resolved_backend = registry.resolve_backend(bk)
|
||||
if resolved_backend == "cv2" and not registry.inpaint_model_available():
|
||||
resolved_backend = watermark_registry.resolve_backend(bk)
|
||||
if resolved_backend == "cv2" and not watermark_registry.inpaint_model_available():
|
||||
console.print(" Note: using cv2 fill (install the 'migan' extra for a lightweight ONNX model).")
|
||||
|
||||
# ``auto`` removes EVERY detected in_auto mark in one pass (a Jimeng-basic image
|
||||
@@ -579,82 +686,20 @@ def cmd_visible(
|
||||
# read -> provenance -> localize/fill -> write -> metadata-strip to the library
|
||||
# entry point, so the CLI and the library go through ONE path (no drift).
|
||||
if mark == "auto" and detect:
|
||||
from remove_ai_watermarks import api
|
||||
|
||||
t0 = time.monotonic()
|
||||
try:
|
||||
with console.status("Detecting & removing visible marks..."):
|
||||
result, removed = api.remove_visible(
|
||||
str(source),
|
||||
str(output),
|
||||
sensitivity=sens,
|
||||
backend=bk,
|
||||
strip_metadata=strip_metadata,
|
||||
write_noop=False,
|
||||
)
|
||||
except RuntimeError as e: # e.g. a selected migan/lama backend whose extra is absent
|
||||
console.print(f" Error: {e}")
|
||||
raise SystemExit(1) from e
|
||||
except (ValueError, OSError) as e: # unreadable / truncated / non-image input
|
||||
console.print(f" Error: cannot read image {source.name}: {e}")
|
||||
raise SystemExit(1) from e
|
||||
elapsed = time.monotonic() - t0
|
||||
h, w = result.shape[:2]
|
||||
console.print(f" Input: {source.name} ({w}x{h})")
|
||||
if not removed:
|
||||
# write_noop=False means nothing was written, so a pre-existing file at the
|
||||
# output path is left intact (the no-mark contract writes nothing).
|
||||
console.print(" No known visible mark detected (gemini / doubao / jimeng / jimeng-pill / samsung).")
|
||||
_no_visible_mark_exit(source, sensitivity=sens)
|
||||
console.print(f" Removed: {', '.join(removed)}")
|
||||
size_kb = output.stat().st_size / 1024
|
||||
console.print(f" Saved: {output} ({size_kb:.0f} KB, {elapsed:.2f}s)")
|
||||
_run_visible_auto(source, output, backend=bk, sensitivity=sens, strip_metadata=strip_metadata)
|
||||
return
|
||||
|
||||
# Explicit single mark (or --no-detect): needs the decoded array + the per-mark gate,
|
||||
# so it keeps its own read/remove/write (still through the shared io + registry).
|
||||
image, alpha = image_io.read_bgr_and_alpha(source)
|
||||
if image is None:
|
||||
console.print(f"Error: Failed to read image: {source}")
|
||||
raise SystemExit(1)
|
||||
h, w = image.shape[:2]
|
||||
console.print(f" Input: {source.name} ({w}x{h})")
|
||||
provenance = _visible_provenance(source)
|
||||
target = "gemini" if mark == "auto" else mark # --no-detect auto: gemini fallback
|
||||
chosen = registry.get_mark(target)
|
||||
# A single explicit mark has no cross-mark pass (no sibling corroboration), so use the
|
||||
# canonical arbiter policy with an empty strict-sibling set instead of re-deriving it
|
||||
# inline (keeps this in lockstep with `decide`).
|
||||
prov = registry.resolve_relax(chosen.key, sensitivity=sens, provenance=provenance, strict_keys=set())
|
||||
det = chosen.detect(image, provenance=prov)
|
||||
if detect and not det.detected:
|
||||
console.print(f" {chosen.label} not detected (conf {det.confidence:.2f}). Use --no-detect to force.")
|
||||
_no_visible_mark_exit(source, sensitivity=sens)
|
||||
if det.detected:
|
||||
console.print(f" {chosen.label} detected ({chosen.location}, conf {det.confidence:.2f})")
|
||||
t0 = time.monotonic()
|
||||
try:
|
||||
with console.status(f"Removing {chosen.label}... ({resolved_backend})"):
|
||||
result, _ = chosen.remove(image, backend=bk, provenance=prov, force=not detect)
|
||||
except RuntimeError as e: # e.g. a selected migan/lama backend whose extra is absent
|
||||
console.print(f" Error: {e}")
|
||||
raise SystemExit(1) from e
|
||||
elapsed = time.monotonic() - t0
|
||||
|
||||
# Save (rejoins the original alpha plane unchanged) + strip metadata.
|
||||
output.parent.mkdir(parents=True, exist_ok=True)
|
||||
image_io.write_bgr_with_alpha(output, result, alpha)
|
||||
if strip_metadata:
|
||||
try:
|
||||
from remove_ai_watermarks.metadata import remove_ai_metadata
|
||||
|
||||
remove_ai_metadata(output, output)
|
||||
except Exception as e:
|
||||
if ctx.obj.get("verbose"):
|
||||
console.print(f" Warning: Failed to strip metadata: {e}")
|
||||
|
||||
size_kb = output.stat().st_size / 1024
|
||||
console.print(f" Saved: {output} ({size_kb:.0f} KB, {elapsed:.2f}s)")
|
||||
_run_visible_explicit(
|
||||
ctx,
|
||||
source,
|
||||
output,
|
||||
detect=detect,
|
||||
mark=mark,
|
||||
backend=bk,
|
||||
sensitivity=sens,
|
||||
resolved_backend=resolved_backend,
|
||||
strip_metadata=strip_metadata,
|
||||
)
|
||||
|
||||
|
||||
# ── Universal region eraser ──
|
||||
@@ -1261,32 +1306,106 @@ def _passthrough_copy(img_path: Path, out_path: Path) -> None:
|
||||
image_io.write_bgr_with_alpha(out_path, src_bgr, src_alpha)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class _BatchOptions:
|
||||
"""Validated processing options shared by every image in one batch.
|
||||
|
||||
Click necessarily exposes these as individual command parameters, but the
|
||||
processing core should receive one coherent value instead of a 21-argument
|
||||
call. Keeping the object immutable also makes it safe to reuse while the
|
||||
batch caches model instances in ``ctx.obj``.
|
||||
"""
|
||||
|
||||
strength: float | None
|
||||
steps: int
|
||||
pipeline: str
|
||||
device: str
|
||||
seed: int | None
|
||||
hf_token: str | None
|
||||
humanize: float
|
||||
backend: str = "auto"
|
||||
sensitivity: str = "auto"
|
||||
unsharp: float = 0.0
|
||||
max_resolution: int = 0
|
||||
min_resolution: int = 1024
|
||||
controlnet_scale: float = 1.0
|
||||
upscaler: str = "lanczos"
|
||||
model: str | None = None
|
||||
guidance_scale: float | None = None
|
||||
adaptive_polish: bool = False
|
||||
tile: bool = False
|
||||
tile_size: int = 1024
|
||||
tile_overlap: int = 128
|
||||
force: bool = False
|
||||
|
||||
|
||||
def _run_batch_invisible(
|
||||
ctx: click.Context,
|
||||
img_path: Path,
|
||||
out_path: Path,
|
||||
mode: str,
|
||||
options: _BatchOptions,
|
||||
) -> bool:
|
||||
"""Run or safely skip the invisible pass for one batch image.
|
||||
|
||||
Returns ``True`` only when a detectable target could not be processed because
|
||||
the GPU dependencies are missing. The availability probe is intentionally
|
||||
evaluated once so branching cannot observe inconsistent optional-dependency
|
||||
state.
|
||||
"""
|
||||
from remove_ai_watermarks.invisible_engine import is_available as invisible_available
|
||||
|
||||
skip_no_signal = _should_skip_invisible_scrub(options.force, img_path)
|
||||
available = invisible_available()
|
||||
if available and not skip_no_signal:
|
||||
from remove_ai_watermarks.invisible_engine import InvisibleEngine
|
||||
|
||||
# Cache the engine in ctx.obj so the batch builds it once (pipeline is a
|
||||
# single CLI value, constant across the run).
|
||||
engines = ctx.obj.setdefault("_inv_engines", {})
|
||||
if options.pipeline not in engines:
|
||||
engines[options.pipeline] = InvisibleEngine(
|
||||
model_id=options.model,
|
||||
device=None if options.device == "auto" else options.device,
|
||||
pipeline=options.pipeline,
|
||||
hf_token=options.hf_token,
|
||||
controlnet_conditioning_scale=options.controlnet_scale,
|
||||
)
|
||||
engines[options.pipeline].remove_watermark(
|
||||
img_path if mode == "invisible" else out_path,
|
||||
out_path,
|
||||
strength=options.strength,
|
||||
num_inference_steps=options.steps,
|
||||
guidance_scale=options.guidance_scale,
|
||||
seed=options.seed,
|
||||
humanize=options.humanize,
|
||||
unsharp=options.unsharp,
|
||||
adaptive_polish=options.adaptive_polish,
|
||||
max_resolution=options.max_resolution,
|
||||
min_resolution=options.min_resolution,
|
||||
upscaler=options.upscaler,
|
||||
tile=options.tile,
|
||||
tile_size=options.tile_size,
|
||||
tile_overlap=options.tile_overlap,
|
||||
# Detect the vendor from the pristine original (`img_path`), not the
|
||||
# visible-processed `out_path` whose C2PA is already gone.
|
||||
vendor=vendor_for_strength(img_path),
|
||||
)
|
||||
return False
|
||||
|
||||
# Invisible-only mode has no preceding visible pass to create ``out_path``.
|
||||
# Preserve a complete output directory while deliberately leaving pixels intact.
|
||||
if mode == "invisible" and not out_path.exists():
|
||||
_passthrough_copy(img_path, out_path)
|
||||
return not available and not skip_no_signal
|
||||
|
||||
|
||||
def _process_batch_image(
|
||||
ctx: click.Context,
|
||||
img_path: Path,
|
||||
out_path: Path,
|
||||
mode: str,
|
||||
strength: float | None,
|
||||
steps: int,
|
||||
pipeline: str,
|
||||
device: str,
|
||||
seed: int | None,
|
||||
hf_token: str | None,
|
||||
humanize: float,
|
||||
backend: str = "auto",
|
||||
sensitivity: str = "auto",
|
||||
unsharp: float = 0.0,
|
||||
max_resolution: int = 0,
|
||||
min_resolution: int = 1024,
|
||||
controlnet_scale: float = 1.0,
|
||||
upscaler: str = "lanczos",
|
||||
model: str | None = None,
|
||||
guidance_scale: float | None = None,
|
||||
adaptive_polish: bool = False,
|
||||
tile: bool = False,
|
||||
tile_size: int = 1024,
|
||||
tile_overlap: int = 128,
|
||||
force: bool = False,
|
||||
options: _BatchOptions,
|
||||
) -> bool:
|
||||
"""Process a single image for batch mode.
|
||||
|
||||
@@ -1312,71 +1431,21 @@ def _process_batch_image(
|
||||
if image is None:
|
||||
raise ValueError("Failed to read image")
|
||||
|
||||
result, _ = _remove_visible_auto(image, source_path=img_path, backend=backend, sensitivity=sensitivity)
|
||||
result, _ = _remove_visible_auto(
|
||||
image,
|
||||
source_path=img_path,
|
||||
backend=options.backend,
|
||||
sensitivity=options.sensitivity,
|
||||
)
|
||||
|
||||
image_io.write_bgr_with_alpha(out_path, result, alpha)
|
||||
saved_alpha = alpha
|
||||
|
||||
if mode in ("invisible", "all"):
|
||||
from remove_ai_watermarks.invisible_engine import (
|
||||
is_available as invisible_available,
|
||||
)
|
||||
|
||||
# Skip the destructive regeneration when no invisible watermark is locally
|
||||
# detectable (would only degrade a clean image). Read the pristine `img_path`;
|
||||
# `out_path` may already be the visible-processed result. --force overrides.
|
||||
skip_no_signal = _should_skip_invisible_scrub(force, img_path)
|
||||
if invisible_available() and not skip_no_signal:
|
||||
from remove_ai_watermarks.invisible_engine import InvisibleEngine
|
||||
|
||||
# Cache the engine in ctx.obj so the batch builds it once (pipeline is a
|
||||
# single CLI value, constant across the run).
|
||||
engines = ctx.obj.setdefault("_inv_engines", {})
|
||||
if pipeline not in engines:
|
||||
engines[pipeline] = InvisibleEngine(
|
||||
model_id=model,
|
||||
device=None if device == "auto" else device,
|
||||
pipeline=pipeline,
|
||||
hf_token=hf_token,
|
||||
controlnet_conditioning_scale=controlnet_scale,
|
||||
)
|
||||
engine_inv = engines[pipeline]
|
||||
engine_inv.remove_watermark(
|
||||
img_path if mode == "invisible" else out_path,
|
||||
out_path,
|
||||
strength=strength,
|
||||
num_inference_steps=steps,
|
||||
guidance_scale=guidance_scale,
|
||||
seed=seed,
|
||||
humanize=humanize,
|
||||
unsharp=unsharp,
|
||||
adaptive_polish=adaptive_polish,
|
||||
max_resolution=max_resolution,
|
||||
min_resolution=min_resolution,
|
||||
upscaler=upscaler,
|
||||
tile=tile,
|
||||
tile_size=tile_size,
|
||||
tile_overlap=tile_overlap,
|
||||
# Detect the vendor from the pristine original (`img_path`), not the
|
||||
# visible-processed `out_path` whose C2PA is already gone.
|
||||
vendor=vendor_for_strength(img_path),
|
||||
)
|
||||
elif not invisible_available() and not skip_no_signal:
|
||||
# An invisible signal IS present but the GPU deps are missing, so the
|
||||
# SynthID scrub cannot run. Mirror the single `all` command's loud skip:
|
||||
# flag it for a batch-level warning + non-zero exit (a silently retained
|
||||
# SynthID watermark is the #1 "it didn't work" report). For invisible mode
|
||||
# nothing wrote out_path yet -> copy the input through so the output dir is
|
||||
# complete with the pixels deliberately left intact (without this, a
|
||||
# signal-bearing image in a GPU-less --mode invisible run got NO output).
|
||||
synthid_skipped = True
|
||||
if mode == "invisible" and not out_path.exists():
|
||||
_passthrough_copy(img_path, out_path)
|
||||
elif skip_no_signal and mode == "invisible" and not out_path.exists():
|
||||
# No invisible target and the visible/all pass did not write out_path
|
||||
# (invisible mode): copy the input through so the output dir is complete
|
||||
# with the pixels deliberately left intact.
|
||||
_passthrough_copy(img_path, out_path)
|
||||
synthid_skipped = _run_batch_invisible(ctx, img_path, out_path, mode, options)
|
||||
|
||||
if mode in ("metadata", "all"):
|
||||
from remove_ai_watermarks.metadata import remove_ai_metadata
|
||||
@@ -1485,6 +1554,29 @@ def cmd_batch(
|
||||
if mode in ("invisible", "all"):
|
||||
_warn_if_esrgan_unavailable(upscaler)
|
||||
adaptive_polish = _resolve_auto_polish(auto, adaptive_polish)
|
||||
options = _BatchOptions(
|
||||
strength=strength,
|
||||
steps=steps,
|
||||
pipeline=pipeline,
|
||||
device=device,
|
||||
seed=seed,
|
||||
hf_token=hf_token,
|
||||
humanize=humanize,
|
||||
backend=backend,
|
||||
sensitivity=sensitivity,
|
||||
unsharp=unsharp,
|
||||
max_resolution=max_resolution,
|
||||
min_resolution=min_resolution,
|
||||
controlnet_scale=controlnet_scale,
|
||||
upscaler=upscaler,
|
||||
model=model,
|
||||
guidance_scale=guidance_scale,
|
||||
adaptive_polish=adaptive_polish,
|
||||
tile=tile,
|
||||
tile_size=tile_size,
|
||||
tile_overlap=tile_overlap,
|
||||
force=force,
|
||||
)
|
||||
|
||||
processed = 0
|
||||
errors = 0
|
||||
@@ -1510,27 +1602,7 @@ def cmd_batch(
|
||||
img_path=img_path,
|
||||
out_path=out_path,
|
||||
mode=mode,
|
||||
strength=strength,
|
||||
steps=steps,
|
||||
pipeline=pipeline,
|
||||
device=device,
|
||||
seed=seed,
|
||||
hf_token=hf_token,
|
||||
humanize=humanize,
|
||||
backend=backend,
|
||||
sensitivity=sensitivity,
|
||||
unsharp=unsharp,
|
||||
max_resolution=max_resolution,
|
||||
min_resolution=min_resolution,
|
||||
controlnet_scale=controlnet_scale,
|
||||
upscaler=upscaler,
|
||||
model=model,
|
||||
guidance_scale=guidance_scale,
|
||||
adaptive_polish=adaptive_polish,
|
||||
tile=tile,
|
||||
tile_size=tile_size,
|
||||
tile_overlap=tile_overlap,
|
||||
force=force,
|
||||
options=options,
|
||||
):
|
||||
synthid_skipped_count += 1
|
||||
processed += 1
|
||||
|
||||
@@ -505,6 +505,42 @@ def _trustmark(image_path: Path) -> str | None:
|
||||
return detect_trustmark(image_path)
|
||||
|
||||
|
||||
def _collect_visible_signals(
|
||||
image_path: Path,
|
||||
signals: list[Signal],
|
||||
watermarks: list[str],
|
||||
platform: str | None,
|
||||
) -> str | None:
|
||||
"""Decode once, append every trusted visible-mark signal, and return platform.
|
||||
|
||||
Keeping this stage separate from metadata aggregation makes the optional cv2
|
||||
boundary explicit and guarantees that all visible detectors share one decoded
|
||||
BGR array. A decode failure preserves the detectors' historical fallback/no-op
|
||||
behavior.
|
||||
"""
|
||||
image: NDArray[Any] | None = None
|
||||
try:
|
||||
from remove_ai_watermarks.image_io import imread
|
||||
|
||||
image = imread(image_path)
|
||||
except Exception as exc: # cv2 missing - detectors fall back / no-op
|
||||
logger.debug("visible-mark decode unavailable: %s", exc)
|
||||
|
||||
sparkle_conf = _visible_sparkle(image_path, image=image)
|
||||
if sparkle_conf is not None and sparkle_conf >= _SPARKLE_THRESHOLD:
|
||||
signals.append(Signal("visible_sparkle", f"NCC confidence {sparkle_conf:.2f}", "medium"))
|
||||
watermarks.append(f"Visible Gemini sparkle (confidence {sparkle_conf:.2f})")
|
||||
if platform is None:
|
||||
platform = "Google Gemini family (visible sparkle detected)"
|
||||
|
||||
for detection in _visible_text_marks(image_path, image=image):
|
||||
signals.append(Signal(f"visible_{detection.key}", f"NCC confidence {detection.confidence:.2f}", "medium"))
|
||||
watermarks.append(f"Visible {detection.label} (confidence {detection.confidence:.2f})")
|
||||
if platform is None:
|
||||
platform = _VISIBLE_MARK_PLATFORM[detection.key]
|
||||
return platform
|
||||
|
||||
|
||||
def identify(image_path: Path, *, check_visible: bool = True, check_invisible: bool = True) -> ProvenanceReport:
|
||||
"""Identify an image's origin platform and watermark inventory.
|
||||
|
||||
@@ -755,36 +791,8 @@ def identify(image_path: Path, *, check_visible: bool = True, check_invisible: b
|
||||
or xai_sig
|
||||
)
|
||||
|
||||
# Decode the file ONCE for every visible-mark detector. The sparkle and the
|
||||
# text-mark detectors both consume a BGR array; letting each re-read the file
|
||||
# was two full cv2 decodes of the same bitmap, which spikes memory on a small
|
||||
# worker. None (cv2 missing / unreadable container) makes each detector fall
|
||||
# back to its own read, preserving the old behavior.
|
||||
vis_image: NDArray[Any] | None = None
|
||||
if check_visible:
|
||||
try:
|
||||
from remove_ai_watermarks.image_io import imread
|
||||
|
||||
vis_image = imread(image_path)
|
||||
except Exception as exc: # cv2 missing - detectors fall back / no-op
|
||||
logger.debug("visible-mark decode unavailable: %s", exc)
|
||||
|
||||
# ── Visible Gemini sparkle (fallback for stripped-metadata case) ─
|
||||
sparkle_conf = _visible_sparkle(image_path, image=vis_image) if check_visible else None
|
||||
if sparkle_conf is not None and sparkle_conf >= _SPARKLE_THRESHOLD:
|
||||
signals.append(Signal("visible_sparkle", f"NCC confidence {sparkle_conf:.2f}", "medium"))
|
||||
watermarks.append(f"Visible Gemini sparkle (confidence {sparkle_conf:.2f})")
|
||||
if platform is None:
|
||||
platform = "Google Gemini family (visible sparkle detected)"
|
||||
|
||||
# ── Visible Doubao / Jimeng text marks (registry; same stripped-metadata
|
||||
# fallback role as the Gemini sparkle above) ─
|
||||
if check_visible:
|
||||
for det in _visible_text_marks(image_path, image=vis_image):
|
||||
signals.append(Signal(f"visible_{det.key}", f"NCC confidence {det.confidence:.2f}", "medium"))
|
||||
watermarks.append(f"Visible {det.label} (confidence {det.confidence:.2f})")
|
||||
if platform is None:
|
||||
platform = _VISIBLE_MARK_PLATFORM[det.key]
|
||||
platform = _collect_visible_signals(image_path, signals, watermarks, platform)
|
||||
|
||||
visible_only = any(s.name.startswith("visible_") for s in signals) and not ai_from_metadata
|
||||
hf_only = bool(hf_job) and not ai_from_metadata
|
||||
|
||||
@@ -347,7 +347,7 @@ def has_ai_metadata(image_path: Path) -> bool:
|
||||
return True
|
||||
# China TC260 AIGC label as a PNG text chunk (the byte scan above catches
|
||||
# only the XMP form; the raw-JSON tEXt chunk needs the PIL-based parse).
|
||||
if aigc_label(image_path):
|
||||
if aigc_label(image_path) is not None:
|
||||
return True
|
||||
# HuggingFace-hosted job marker (hf-job-id PNG text chunk).
|
||||
if huggingface_job(image_path):
|
||||
@@ -887,7 +887,7 @@ def get_ai_metadata(image_path: Path) -> dict[str, str]:
|
||||
result["soft_binding"] = ", ".join(vendors)
|
||||
|
||||
# China TC260 AI-content label (Doubao and other China-served generators).
|
||||
if aigc := aigc_label(image_path):
|
||||
if (aigc := aigc_label(image_path)) is not None:
|
||||
producer = aigc.get("ContentProducer", "")
|
||||
result["aigc_label"] = f"China AIGC label (TC260){f'; producer {producer}' if producer else ''}"
|
||||
|
||||
|
||||
@@ -50,16 +50,24 @@ Backend = Literal["auto", "cv2", "migan", "lama"]
|
||||
# a clean corner). Lowest recall on faint/moved marks.
|
||||
# * ``auto`` (default): relax a mark's gate ONLY when the image carries same-product
|
||||
# evidence the mark is there -- metadata provenance for that vendor, or a confidently
|
||||
# detected sibling mark of the same product (see ``resolve_relax``). No evidence ->
|
||||
# detected sibling mark of the same product (see ``resolve_trust``). No evidence ->
|
||||
# stays strict. Safe: it only escalates where the mark is corroborated.
|
||||
# * ``assume_ai``: relax every mark's gate regardless of evidence -- the caller asserts
|
||||
# the image is AI and wants the mark gone (e.g. a metadata-stripped screenshot uploaded
|
||||
# to a watermark remover). Recovers the faint/moved marks the strict gate demotes
|
||||
# (~49% -> ~89% Gemini recall, corpus-measured), at the cost of a harmless small fill
|
||||
# on some clean corners. The library CANNOT infer this from a stripped image -- only the
|
||||
# caller's out-of-band context (the user uploaded to remove a mark) justifies it.
|
||||
# to a watermark remover). Recovers the faint/moved marks the strict gate demotes. The
|
||||
# library CANNOT infer this from a stripped image -- only the caller's out-of-band
|
||||
# context (the user uploaded to remove a mark) justifies it. An assertion that the
|
||||
# image is AI is NOT evidence of WHICH vendor made it, so a mark relaxed on assumption
|
||||
# alone must still clear ``_ASSUMED_CONF_FLOOR``; see that constant for why.
|
||||
Sensitivity = Literal["auto", "strict", "assume_ai"]
|
||||
|
||||
# The trust level a mark's detection gate is resolved to (see ``resolve_trust``). The
|
||||
# split between ``assumed`` and ``confirmed`` is load-bearing: both bypass the engine's
|
||||
# false-positive gate, but only ``confirmed`` has evidence naming THIS vendor, which is
|
||||
# exactly what that bypass is documented to require (see GeminiEngine.detect_watermark's
|
||||
# ``trust_provenance`` contract). ``assumed`` therefore carries a confidence floor.
|
||||
Trust = Literal["strict", "assumed", "confirmed"]
|
||||
|
||||
# Product family per mark, for the ``auto`` cross-mark corroboration: a confidently
|
||||
# detected mark relaxes only OTHER marks of the SAME product (different corners, one
|
||||
# product -- the Jimeng wordmark + the Jimeng pill). Doubao and Jimeng are BOTH ByteDance
|
||||
@@ -120,6 +128,8 @@ class Candidate:
|
||||
Carries the mark's verdict at BOTH trust levels (``detected_strict`` = the
|
||||
conservative gate, ``detected_relaxed`` = the gate the engine relaxes to under
|
||||
provenance/assume), so the arbiter can pick per mark without re-running detection.
|
||||
``relaxed_confidence`` is the gate-bypassed detection's confidence, which the arbiter
|
||||
needs to apply :func:`assumed_floor_ok` when a mark is relaxed on assumption alone.
|
||||
``features`` is a generic bag of physical measurements a mark's gate may need (the
|
||||
mark owns which it reports via ``KnownMark._features``); e.g. the pill supplies
|
||||
``footprint_flat`` (0/1). Empty for marks whose gate needs no extra evidence."""
|
||||
@@ -128,6 +138,7 @@ class Candidate:
|
||||
label: str
|
||||
detected_strict: bool
|
||||
detected_relaxed: bool
|
||||
relaxed_confidence: float
|
||||
features: dict[str, float] # generic; both construction sites always supply it (empty when none)
|
||||
|
||||
|
||||
@@ -446,28 +457,59 @@ def detect_marks(
|
||||
return [m.detect(image, provenance=m.key in provenance) for m in _REGISTRY if include_explicit or m.in_auto]
|
||||
|
||||
|
||||
def resolve_relax(
|
||||
# Minimum gate-bypassed confidence a mark must reach when it is relaxed on ASSUMPTION
|
||||
# (``assume_ai``) rather than on evidence naming its vendor. Relaxing bypasses the
|
||||
# engine's false-positive gate entirely, which is justified by vendor CONFIRMATION; an
|
||||
# assumption that the image is AI says nothing about WHICH vendor, so the bypassed
|
||||
# detector needs its own floor or it fires on ordinary content.
|
||||
#
|
||||
# Corpus-measured 2026-07-16 (256 genuine camera captures -- Make/Model/exposure/aperture
|
||||
# present and no AI token, so a Gemini sparkle cannot be there -- vs 697 Google-C2PA
|
||||
# positives, metadata used only as the label, never fed to the detector):
|
||||
#
|
||||
# bypassed threshold recall false-fire on clean photos
|
||||
# 0.35 82.6% 59.8% <- the bare detector gate
|
||||
# 0.45 66.6% 12.5%
|
||||
# 0.50 59.4% 0.0% <- chosen
|
||||
# strict gate 56.4% 0.0%
|
||||
#
|
||||
# So 0.35 sat on a cliff: it bought +26pp recall over strict by filling a corner on ~6
|
||||
# of every 10 CLEAN photos. At 0.50 the flag is honest -- it still beats strict, for free.
|
||||
# Marks absent from this dict relax identically at both levels; their bypassed false-fire
|
||||
# on the same negatives is under 1% (doubao 0.8%, jimeng 0.4%, samsung 0.4%).
|
||||
_ASSUMED_CONF_FLOOR: dict[str, float] = {"gemini": 0.50}
|
||||
|
||||
|
||||
def assumed_floor_ok(key: str, confidence: float) -> bool:
|
||||
"""Whether an ``assumed``-trust detection of ``key`` at ``confidence`` is trustworthy
|
||||
enough to act on (see :data:`_ASSUMED_CONF_FLOOR`). Marks with no floor always pass."""
|
||||
floor = _ASSUMED_CONF_FLOOR.get(key)
|
||||
return floor is None or confidence >= floor
|
||||
|
||||
|
||||
def resolve_trust(
|
||||
key: str,
|
||||
*,
|
||||
sensitivity: Sensitivity,
|
||||
provenance: frozenset[str],
|
||||
strict_keys: set[str],
|
||||
) -> bool:
|
||||
"""Whether mark ``key``'s detection gate is relaxed (strict -> assume level).
|
||||
) -> Trust:
|
||||
"""The trust level mark ``key``'s detection gate is resolved to.
|
||||
|
||||
The single place that turns the ``sensitivity`` policy + evidence into a per-mark
|
||||
boolean (which the engines consume): ``strict`` never relaxes, ``assume_ai`` always
|
||||
relaxes, and ``auto`` relaxes only on same-product evidence -- the vendor confirmed
|
||||
by metadata (``key in provenance``) or a confidently strict-detected sibling of the
|
||||
same product (``_PRODUCT_OF``)."""
|
||||
level (which the engines consume as ``provenance = level != "strict"``). ``strict``
|
||||
never relaxes. A mark is ``confirmed`` only on same-product evidence -- the vendor
|
||||
confirmed by metadata (``key in provenance``) or a confidently strict-detected
|
||||
sibling of the same product (``_PRODUCT_OF``). Without that evidence, ``assume_ai``
|
||||
yields ``assumed`` (relaxed, but subject to :func:`assumed_floor_ok`) and ``auto``
|
||||
stays ``strict``."""
|
||||
if sensitivity == "strict":
|
||||
return False
|
||||
if sensitivity == "assume_ai":
|
||||
return True
|
||||
if key in provenance:
|
||||
return True
|
||||
return "strict"
|
||||
product = _PRODUCT_OF[key]
|
||||
return any(_PRODUCT_OF[k] == product for k in strict_keys if k != key)
|
||||
confirmed = key in provenance or any(_PRODUCT_OF[k] == product for k in strict_keys if k != key)
|
||||
if confirmed:
|
||||
return "confirmed"
|
||||
return "assumed" if sensitivity == "assume_ai" else "strict"
|
||||
|
||||
|
||||
def _keep_pill(keys: set[str], *, provenance: frozenset[str], sensitivity: Sensitivity, footprint_flat: bool) -> bool:
|
||||
@@ -515,7 +557,7 @@ def _build_candidates(image: NDArray[Any]) -> list[Candidate]:
|
||||
strict = m.detect(image, provenance=False)
|
||||
relaxed = m.detect(image, provenance=True)
|
||||
feats = m.features(image) if (strict.detected or relaxed.detected) else {}
|
||||
cands.append(Candidate(m.key, m.label, strict.detected, relaxed.detected, feats))
|
||||
cands.append(Candidate(m.key, m.label, strict.detected, relaxed.detected, relaxed.confidence, feats))
|
||||
return cands
|
||||
|
||||
|
||||
@@ -523,17 +565,25 @@ def decide(candidates: list[Candidate], context: Context) -> list[Decision]:
|
||||
"""The removal ARBITER: a pure function turning perception + context into the
|
||||
ordered list of marks to remove (and the trust level each was accepted at).
|
||||
|
||||
All policy lives here, in one place: per-mark relaxation (:func:`resolve_relax`,
|
||||
which needs the strict-detected siblings for ``auto`` cross-mark corroboration) and
|
||||
the capture-less pill gate (:func:`_keep_pill`). No image, no I/O -- so it is
|
||||
unit-testable in isolation and the same decision drives every caller."""
|
||||
All policy lives here, in one place: per-mark trust resolution (:func:`resolve_trust`,
|
||||
which needs the strict-detected siblings for ``auto`` cross-mark corroboration), the
|
||||
assumed-trust confidence floor (:func:`assumed_floor_ok`) and the capture-less pill
|
||||
gate (:func:`_keep_pill`). No image, no I/O -- so it is unit-testable in isolation and
|
||||
the same decision drives every caller."""
|
||||
strict_keys = {c.key for c in candidates if c.detected_strict}
|
||||
fired: list[Decision] = []
|
||||
for c in candidates:
|
||||
relax = resolve_relax(
|
||||
trust = resolve_trust(
|
||||
c.key, sensitivity=context.sensitivity, provenance=context.provenance, strict_keys=strict_keys
|
||||
)
|
||||
if c.detected_relaxed if relax else c.detected_strict:
|
||||
relax = trust != "strict"
|
||||
ok = c.detected_relaxed if relax else c.detected_strict
|
||||
if trust == "assumed" and not assumed_floor_ok(c.key, c.relaxed_confidence):
|
||||
# Relaxed on assumption alone and too weak to trust: fall back to the strict
|
||||
# verdict rather than dropping the mark, so assume_ai is monotonic -- it only
|
||||
# ever ADDS recall over strict, never removes less than strict would.
|
||||
ok, relax = c.detected_strict, False
|
||||
if ok:
|
||||
fired.append(Decision(c, relax))
|
||||
keys = {d.candidate.key for d in fired}
|
||||
if "jimeng_pill" in keys:
|
||||
|
||||
@@ -104,6 +104,24 @@ class TestVisibleProvenance:
|
||||
def test_unreadable_path_is_empty(self, tmp_path):
|
||||
assert raiw.visible_provenance(tmp_path / "missing.png") == frozenset()
|
||||
|
||||
def test_uses_report_signals_for_falsy_metadata_values(self, monkeypatch, tmp_path):
|
||||
"""An empty TC260 object and Samsung genAIType=0 are still present signals.
|
||||
|
||||
The report has already normalized those values, so the public API must not
|
||||
re-read the file and accidentally discard them by truthiness.
|
||||
"""
|
||||
from types import SimpleNamespace
|
||||
|
||||
from remove_ai_watermarks import identify
|
||||
|
||||
report = SimpleNamespace(
|
||||
platform=None,
|
||||
signals=[SimpleNamespace(name="aigc"), SimpleNamespace(name="samsung_genai")],
|
||||
)
|
||||
monkeypatch.setattr(identify, "identify", lambda *args, **kwargs: report)
|
||||
|
||||
assert raiw.visible_provenance(tmp_path / "synthetic.png") == frozenset({"doubao", "jimeng", "samsung"})
|
||||
|
||||
|
||||
class TestRemoveVisibleOutputPath:
|
||||
"""Output-path robustness: in-place clean (#3) and a missing output dir (#4)."""
|
||||
|
||||
@@ -978,6 +978,19 @@ class TestAIGCLabel:
|
||||
assert "aigc_label" in meta
|
||||
assert "TC260" in meta["aigc_label"]
|
||||
|
||||
def test_empty_namespaced_label_is_still_surfaced(self, tmp_path: Path):
|
||||
"""The namespaced element is unambiguous even when its JSON object is empty."""
|
||||
from remove_ai_watermarks.metadata import aigc_label
|
||||
|
||||
p = tmp_path / "empty_aigc.png"
|
||||
Image.new("RGB", (32, 32)).save(p)
|
||||
with open(p, "ab") as f:
|
||||
f.write(b"<TC260:AIGC>{}</TC260:AIGC>")
|
||||
|
||||
assert aigc_label(p) == {}
|
||||
assert has_ai_metadata(p)
|
||||
assert "aigc_label" in get_ai_metadata(p)
|
||||
|
||||
def _aigc_chunk_png(self, tmp_path: Path, producer: str = "doubao") -> Path:
|
||||
"""Doubao writes the TC260 object as a PNG ``tEXt`` chunk keyed ``AIGC``
|
||||
with raw JSON (no XMP, no namespaced marker)."""
|
||||
|
||||
@@ -165,36 +165,60 @@ class TestLocalizeFill:
|
||||
|
||||
|
||||
class TestSensitivity:
|
||||
"""``resolve_relax`` turns the sensitivity policy + evidence into the per-mark
|
||||
relaxation boolean the engines consume."""
|
||||
"""``resolve_trust`` turns the sensitivity policy + evidence into the per-mark
|
||||
trust level the engines consume."""
|
||||
|
||||
def test_strict_never_relaxes(self):
|
||||
# even with metadata provenance, strict keeps the conservative gate
|
||||
assert (
|
||||
reg.resolve_relax("gemini", sensitivity="strict", provenance=frozenset({"gemini"}), strict_keys=set())
|
||||
is False
|
||||
reg.resolve_trust("gemini", sensitivity="strict", provenance=frozenset({"gemini"}), strict_keys=set())
|
||||
== "strict"
|
||||
)
|
||||
|
||||
def test_assume_ai_always_relaxes(self):
|
||||
assert reg.resolve_relax("gemini", sensitivity="assume_ai", provenance=frozenset(), strict_keys=set()) is True
|
||||
def test_assume_ai_without_evidence_is_assumed_not_confirmed(self):
|
||||
# asserting the image is AI says nothing about WHICH vendor made it, so the mark
|
||||
# is relaxed on assumption only -- it must not inherit the confirmed-vendor bypass
|
||||
assert (
|
||||
reg.resolve_trust("gemini", sensitivity="assume_ai", provenance=frozenset(), strict_keys=set()) == "assumed"
|
||||
)
|
||||
|
||||
def test_assume_ai_with_metadata_is_confirmed(self):
|
||||
assert (
|
||||
reg.resolve_trust("gemini", sensitivity="assume_ai", provenance=frozenset({"gemini"}), strict_keys=set())
|
||||
== "confirmed"
|
||||
)
|
||||
|
||||
def test_auto_relaxes_on_own_metadata(self):
|
||||
assert (
|
||||
reg.resolve_relax("gemini", sensitivity="auto", provenance=frozenset({"gemini"}), strict_keys=set()) is True
|
||||
reg.resolve_trust("gemini", sensitivity="auto", provenance=frozenset({"gemini"}), strict_keys=set())
|
||||
== "confirmed"
|
||||
)
|
||||
|
||||
def test_auto_strict_without_evidence(self):
|
||||
assert reg.resolve_relax("gemini", sensitivity="auto", provenance=frozenset(), strict_keys=set()) is False
|
||||
assert reg.resolve_trust("gemini", sensitivity="auto", provenance=frozenset(), strict_keys=set()) == "strict"
|
||||
|
||||
def test_auto_cross_mark_same_product(self):
|
||||
# a detected Jimeng wordmark relaxes the Jimeng pill (same product, other corner)
|
||||
assert (
|
||||
reg.resolve_relax("jimeng_pill", sensitivity="auto", provenance=frozenset(), strict_keys={"jimeng"}) is True
|
||||
reg.resolve_trust("jimeng_pill", sensitivity="auto", provenance=frozenset(), strict_keys={"jimeng"})
|
||||
== "confirmed"
|
||||
)
|
||||
|
||||
def test_auto_no_cross_mark_across_products(self):
|
||||
# a detected Jimeng wordmark must NOT relax Doubao (distinct products, same corner)
|
||||
assert reg.resolve_relax("doubao", sensitivity="auto", provenance=frozenset(), strict_keys={"jimeng"}) is False
|
||||
assert (
|
||||
reg.resolve_trust("doubao", sensitivity="auto", provenance=frozenset(), strict_keys={"jimeng"}) == "strict"
|
||||
)
|
||||
|
||||
def test_assumed_floor_rejects_weak_sparkle_but_passes_strong(self):
|
||||
# the gate-bypassed sparkle detector fires on ~60% of ordinary photos at its bare
|
||||
# 0.35 threshold; only a match well clear of that floor is trustworthy on assumption
|
||||
assert reg.assumed_floor_ok("gemini", 0.35) is False
|
||||
assert reg.assumed_floor_ok("gemini", 0.50) is True
|
||||
|
||||
def test_assumed_floor_default_passes_for_unfloored_marks(self):
|
||||
# text marks relax cleanly (<1% bypassed false-fire), so they carry no floor
|
||||
assert reg.assumed_floor_ok("doubao", 0.36) is True
|
||||
|
||||
def test_remove_auto_marks_accepts_all_sensitivities(self):
|
||||
blank = np.zeros((256, 256, 3), np.uint8)
|
||||
@@ -209,9 +233,11 @@ class TestArbiter:
|
||||
Candidates -- this is the payoff of separating decision from perception."""
|
||||
|
||||
@staticmethod
|
||||
def _c(key, *, strict=False, relaxed=False, flat=False):
|
||||
def _c(key, *, strict=False, relaxed=False, flat=False, relaxed_conf=1.0):
|
||||
# relaxed_conf defaults high so a test that does not care about the assumed-trust
|
||||
# confidence floor exercises the trust logic, not the floor.
|
||||
feats = {"footprint_flat": 1.0} if flat else {}
|
||||
return reg.Candidate(key, f"L:{key}", strict, relaxed, feats)
|
||||
return reg.Candidate(key, f"L:{key}", strict, relaxed, relaxed_conf, feats)
|
||||
|
||||
def _keys(self, cands, ctx):
|
||||
return {d.candidate.key for d in reg.decide(cands, ctx)}
|
||||
@@ -228,6 +254,30 @@ class TestArbiter:
|
||||
assert [d.candidate.key for d in fired] == ["gemini"]
|
||||
assert fired[0].relax is True
|
||||
|
||||
def test_assume_ai_drops_sparkle_below_the_assumed_floor(self):
|
||||
# REGRESSION (2026-07-16): assume_ai passed trust_provenance=True to the engine,
|
||||
# bypassing the sparkle false-positive gate on the mere ASSERTION that the image is
|
||||
# AI -- but that flag is contracted to mean "metadata proved this vendor". The bare
|
||||
# bypassed gate (conf 0.35) fired on 59.8% of 256 genuine camera captures, so
|
||||
# `--sensitivity assume-ai` filled a phantom sparkle on ~6 of every 10 clean photos.
|
||||
weak = self._c("gemini", relaxed=True, relaxed_conf=0.40)
|
||||
assert reg.decide([weak], reg.Context(sensitivity="assume_ai")) == []
|
||||
|
||||
def test_assume_ai_keeps_sparkle_confirmed_by_metadata_below_the_floor(self):
|
||||
# the floor exists because the vendor is UNKNOWN; once metadata names Google the
|
||||
# bypass is contract-legal again, so a weak match is still trusted
|
||||
weak = self._c("gemini", relaxed=True, relaxed_conf=0.40)
|
||||
ctx = reg.Context(sensitivity="assume_ai", provenance=frozenset({"gemini"}))
|
||||
assert [d.candidate.key for d in reg.decide([weak], ctx)] == ["gemini"]
|
||||
|
||||
def test_assume_ai_is_monotonic_over_strict(self):
|
||||
# a mark the STRICT gate accepted must never be dropped by the assumed floor:
|
||||
# assume_ai only ever adds recall
|
||||
weak_but_strict = self._c("gemini", strict=True, relaxed=True, relaxed_conf=0.40)
|
||||
fired = reg.decide([weak_but_strict], reg.Context(sensitivity="assume_ai"))
|
||||
assert [d.candidate.key for d in fired] == ["gemini"]
|
||||
assert fired[0].relax is False # accepted on the strict verdict, so mask at strict
|
||||
|
||||
def test_auto_relaxes_on_provenance(self):
|
||||
c = [self._c("gemini", relaxed=True)]
|
||||
assert self._keys(c, reg.Context(provenance=frozenset({"gemini"}))) == {"gemini"}
|
||||
|
||||
Reference in New Issue
Block a user