56 KiB
Remove-AI-Watermarks
You are a principal Python engineer maintaining a CLI tool and library for removing visible and invisible AI watermarks from images.
Scope and non-goals
The mission is removing AI-provenance watermarks that a platform stamps onto content the user generated themselves — SynthID, the Gemini / Nano Banana sparkle, the Doubao / Jimeng / Qwen / Kling / Tencent Yuanbao / Samsung visible AI labels, the Chinese TC260 "由…AI生成" label, and C2PA / IPTC / EXIF "Made with AI" metadata. The point is user autonomy over their own generated output.
It deliberately does not remove watermarks that protect someone else's paid or copyrighted content — stock-agency overlays (Shutterstock, Getty, iStock, Adobe Stock), classifieds-site marks, or any tiled / diagonal "preview" watermark whose job is to gate a purchase. Stripping those makes a paid resource free off someone else's work; out of scope by principle, not by technical difficulty. The line: a visible mark is in scope when it labels the user's own AI generation, and out of scope when it protects a third party's paid asset.
Consequences for contributors (do not drift back into the stock niche just because it is technically feasible):
- Do not add stock / agency / classifieds watermark removal to
watermark_registry.pyor the eraser, and do not build tiled-overlay or multi-image watermark-estimation features aimed at them. erase --regionstays a generic user-driven tool (the user points at their own object); do not ship an automatic stock-watermark detector/remover on top of it.- New visible-mark templates are for AI-generation labels only.
(Established 2026-06-13 by user instruction: "Я пытаюсь сделать платные ресурсы бесплатными — это не то, против чего мы боремся.")
How to run
Per-command exit-code semantics (the no-signal / GPU-missing skip branches), test traps, and regression-guard paths live in docs/module-internals.md (section "CLI commands (cli.py)") — read it before changing any command's skip/exit behavior. Every single-image command's source argument declares dir_okay=False: without it click.Path(exists=True) accepts a directory, which then reached open() and raised IsADirectoryError (Tier E, 2026-07-20; batch's directory already declared file_okay=False). Regression: tests/test_cli_robustness.py::TestDirectoryInputIsRejected.
uv run remove-ai-watermarks all <image.png> -o <output.png>— full pipeline (visible + invisible + metadata). Same diffusion knobs asinvisible, plus the visible-pass--backend auto|cv2|migan|lama(defaultauto) and--sensitivity auto|strict(defaultauto) for the localize -> fill visible removal (see thevisiblebullet). Skips step 2 (invisible/SynthID) when the[gpu]extra is absent or no invisible signal is detectable; see the module doc for the distinct exit codes.uv run remove-ai-watermarks invisible <image.png> -o <out.png>— diffusion SynthID removal. Full knob set (kept identical acrossinvisible/all/batch):--strength(vendor-adaptive default except resolution-adaptiveqwen-zimage),--steps(interacts with--strengthon the diffusers profiles;watermark_profiles.viable_stepsprevents zero effective steps.qwen-zimageinstead fixes its Lightning stage at 4 steps),--guidance-scale,--pipeline sdxl|controlnet|qwen|qwen-zimage(defaultcontrolnet;qwenandqwen-zimageare manual opt-ins),--controlnet-scale,--model,--device,--seed,--hf-token,--max-resolution/--min-resolution,--upscaler lanczos|esrgan,--humanize,--unsharp,--adaptive-polish/--no-adaptive-polish,--tile/--no-tile+--tile-size/--tile-overlap,--cpu-offload/--no-cpu-offload,--force/--no-force.--cpu-offloadtrades speed for lower CUDA VRAM use by moving Diffusers model components between CPU and GPU; onqwen-zimageit forces the face stack to offload instead of using automatic residency. It has no effect on CPU/MPS. ControlNet is the compatibility and cost default, not the highest-fidelity mode. Recommend the CUDA-onlyqwen-zimageprofile when output quality, especially face identity, matters more than runtime and cost; it needs the separate extra, uses a fixed Qwen-Image-2512 + Z-Image stack, rejects--model, defaults to the oracle-candidate seed 0, and supports tiling only for its global Qwen pass. The full-frame face stage runs once after tile blending. Tiled outputs still need separate oracle certification.--autois deprecated and a no-op that only warns. Skips the diffusion when no invisible signal is detectable; see the module doc.uv run remove-ai-watermarks visible <image.png> -o <out.png>— known-visible-mark removal by localize -> fill: each detected mark is localized to a binary full-frame footprint mask, then one shared, swappable fill inpaints that mask.--backend auto|cv2|migan|lama(defaultauto) picks the fill:cv2(classical inpaint, no deps, the floor),migan(MI-GAN ONNX, light, the memory-tight pick where LaMa will not fit),lama(big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available);auto= LaMa > MI-GAN > cv2, best available.--mark auto(default) removes EVERY detected mark in one pass (a Jimeng-basic image carries the top-left "AI生成" pill AND the bottom-right "★ 即梦AI" wordmark) from: Gemini sparkle, Doubao "豆包AI生成", Jimeng "★ 即梦AI", Qwen "千问AI生成", Kling "可灵AI 3.0", Tencent Yuanbao "元宝 / AI生成", Samsung Galaxy AI "✦ Contenuti generati dall'AI", Baidu "百度 AI生成", LibLibAI wordmark (bottom-center), RunningHub "RunningHub AI生成" (top-left), and the capture-less Jimeng "AI生成" pill (top-left, metadata-gated);--mark gemini|doubao|jimeng|qwen|kling|yuanbao|samsung|baidu|liblib|runninghub|jimeng_pillforces one.--sensitivity auto|strict(defaultauto) sets how hard a borderline mark is trusted:autorelaxes a mark's gate only on same-product evidence (metadata provenance for that vendor, or a confidently detected sibling mark of the same product — clean images stay untouched);strictnever relaxes. Metadata provenance is read automatically and feedsauto. (assume-aiwas REMOVED in 0.16 — see the registry bullet; a user who can SEE a missed mark should point at it witherase --region, or name it with--mark <text-mark> --no-detect.) For arbitrary logos/objects useerase. When no known mark is detected the command writes no output and exits with the no-visible-mark code instead of re-serving the input;--no-detectforces the gemini fallback and proceeds. See the module doc for the routing/exit detail.--backendand--sensitivityare shared acrossvisible/all/batch.uv run remove-ai-watermarks erase <image.png> --region x,y,w,h -o <out.png>— universal region eraser (any logo/object, any position).--backend cv2(default, no deps),--backend migan(MI-GAN via onnxruntime, extramigan; ~28 MB, ~1 GB RAM, near-LaMa), or--backend lama(big-LaMa, extralama; best quality but ~4.7 GB RAM);--regionis repeatable.uv run remove-ai-watermarks identify <image>— provenance verdict (platform + watermark inventory + confidence);--jsonfor machine output,--no-visibleto skip both registered visible detectors and the optional open invisible-watermark decoderuv run remove-ai-watermarks metadata <image.png> --check— inspect AI metadata (C2PA, EXIF, PNG chunks)uv run remove-ai-watermarks metadata <image.png> --remove -o <out.png>— strip all AI metadatauv run remove-ai-watermarks batch <directory>— process every supported image in a directory (output defaults to<directory>_clean/, set with-o).--mode visible|invisible|metadata|all(defaultvisible); the invisible/all path reuses the fullinvisibleknob set above, plus--backendand--sensitivityfor the visible localize -> fill pass. Applies the same no-signal skip per image; see the module doc. Exit code: non-zero when any image errored OR (mirroring singleall) a--mode invisible/allimage carried an invisible signal but the GPU extra was absent, so its SynthID scrub was skipped — it emits a loud warning and copies the input through (invisible mode) so the output dir stays complete; a wrapping service can then detect the incomplete run instead of trusting a silent exit 0.
Test and lint
- CI (
.github/workflows/test.yml): runs on push tomain+ every PR. Alintjob (ubuntu:ruff check+ruff format --check) plus atestmatrix (ubuntu/macos/windows x py3.10/3.12) that doesuv sync --frozen --extra devthenpytest. The matrix installs only core + dev (nogpuextra), so the GPU/model-running tests skip there and it exercises the metadata/identify/visible/cv2-eraser surface on all three OSes. Keepuv.lockvalid (don't break--frozen) when editingpyproject.toml. - Dependency PR checks run the GitHub merge ref against current
main, not the contributor branch in isolation. Ifmainmoves after the dependency branch was opened, merge currentmainlocally and rerun the full gate; a new linter can expose stale directives in code that landed later. - Release flow + distribution channels (PyPI publish via
publish.yml/uv publish, the automated Homebrew-tap + HF-Space bumps indistribute.yml, conda-forge, ComfyUI Registry, the sdistdata/exclusion, hatchling pin history): seedocs/release-and-distribution.mdbefore cutting a release. bash maintain.sh— uv-outdated, uv-secure, ruff check/fix, ruff format, pyright (scopedsrc/, see the OOM note below), pytest -n auto. The helper tools live in thedevextra (pytest-xdist, plusuv-outdated/uv-securemarker-gated to py3.12+ so the py3.10 resolution stays solvable) — a bare env without--extra devdoes not have them.- Strict pyright is clean across
src/(0 errors). The cv2/torch/diffusers boundary files (gemini_engine,region_eraser,doubao_engine,humanizer,invisible_engine,noai/watermark_remover) carry a documented per-file# pyright:relax pragma that turns off only the unknown-type / untyped-third-party rules — those libs ship no usable types, so strict typing there fights the ecosystem. Pure-logic files stay fully strict;typings/piexif/__init__.pyiis a local stub sometadata.py/extractor.pyresolve piexif. Public ndarray-returning signatures on the relaxed engines are still annotatedNDArray[Any]so strict consumers (cli.py) stay clean. When touching a relaxed file, prefer fixing real issues over widening the pragma; keep the pragma scoped to genuinely-untyped boundaries. Theuv-secureCVE-resolution history (idna/aiohttp bumps, retired basicsr, the dismissed torchGHSA-rrmf-rvhw-rf47) lives indocs/release-and-distribution.md— read it before re-triaging a dependency alert. - Full-project
uv run pyright(no path) OOMs/crashes node on this ML-heavy repo (emits alibnodestack frame, no summary) — a known environment limit, not a code error. Gate withuv run --extra dev --extra gpu pyright src/(completes, authoritative) or scope to changed files; also runuv run ruff checkanduv run pytestdirectly. - Run
uv runfrom the repo root — from another cwd it falls back to a bare env without numpy/cv2/torch. - Stale
trustmarkremnant in site-packages after an extras change: thetrustmarkpackage downloads model weights INTO its own package dir, so when a narroweruv syncprunes the package, atrustmark/models/directory survives as an empty namespace package. Symptom: pyright"TrustMark" is unknown import symbolontrustmark_detector.pyandfind_spec("trustmark")returning a loader-less spec (sois_available()lies True). Fix:rm -rf .venv/lib/python3.12/site-packages/trustmark(regenerable weights cache). - To add a dev tool (pytest/ruff/pyright) into the env, use
uv sync --frozen --extra dev --extra gpu, neveruv pip install—uv pip installre-resolves and rewritesuv.lock, which silently bumpedtransformersto a build incompatible with the pinneddiffusers(cannot import name 'Qwen3VLForConditionalGeneration') and broke everyidentify/metadata import. Recovery:git checkout uv.lock && uv sync --frozen --extra gpu --extra dev. Thegpuextra holdsdiffusers/transformers/torch, so a bareuv sync(no extras) removes them;noai/__init__is now lazy (PEP 562__getattr__, so importingidentify/metadatano longer pullswatermark_remover/torch), so a bare env breaks only when the removal pipeline is actually invoked, not on import.maintain.sh'suv sync --all-extrasalso pulls the heavytrustmark/lamawheels (pytorch-lightning, onnxruntime) — fine on a good connection, but on flaky DNS sync only--extra gpu --extra devand run the lint/test steps by hand. - Metadata/C2PA tests assert against real committed fixtures in
data/fixtures/provenance/(chatgpt-*.png= OpenAI C2PA,firefly-1.png= Adobe,mj-1.png= Midjourney IPTC,doubao-1.png= ByteDance Doubao with the China TC260<TC260:AIGC>XMP label and a visible "豆包AI生成" text mark bottom-right;grok-1.jpg= xAI Grok with its EXIF-onlySignature:blob + UUIDArtistand no C2PA/SynthID/IPTC;flux-1.png/flux-1.jpg= real Black Forest Labs FLUX.2 Playground output, signed C2PA (issuer "Black Forest Labs" +trainedAlgorithmicMedia) --flux-1.jpgis the first committed JPEG-with-C2PA fixture, exercising the c2pa-python non-PNG reader path end to end; whether BFL hosted output also embeds the open DWT-DCT pixel watermark is UNRESOLVED -- our detector returns None on these fox samples, but they are high-texture carriers where even a known-embedded watermark fails the round-trip, see the content-fragility caveat indocs/watermarking-landscape.md); synthetic byte blobs cover the remaining JPEG/ISOBMFF format paths. Theclean_photoconftest fixture generates a deterministic metadata-free PNG; no real negative photo is committed for tests. - Repository data follows
data/README.md: executable provenance fixtures live indata/fixtures/, minimal detector rebuild inputs indata/calibration/, canonical provider-oracle originals indata/synthid/, and evaluation-only ground truth indata/evaluations/. Store each binary once and point every consumer at the canonical path. - SynthID oracle fixtures:
scripts/synthid_corpus.pyingests labeled originals intodata/synthid/originals/. The trackedmanifest.csvis kept in sync with the files on disk, one row per image.full-pipeline-quality.csvis the reusable full-pipeline test set: read its single canonical copy, preservesource_filenamein outputs, and keep provider groups separate for their respective oracles. Generated or cleaned outputs stay outside the repository; record their reproducible command, hash, and oracle verdict instead.
Configuration
- GPU/ML modules (invisible_engine, watermark_remover) are optional — guard imports with
is_available()checks - Optional detection extras:
detect(imwatermark — open SD/SDXL/FLUX watermark) andtrustmark(Adobe TrustMark decoder; pulls torch + downloads weights). Both are guarded byis_available()and skipped byidentifywhen absent. - Optional
esrganextra (spandrel only): Real-ESRGAN pre-diffusion super-resolution for small inputs (upscaler.py, CLI--upscaler esrganoninvisible/all/batch). Guarded byupscaler.is_available(); the default upscaler stays Lanczos (cv2, no deps) and the engine falls back to Lanczos when the extra is absent or the model errors. spandrel is MIT and pulls NO basicsr (only torch/torchvision/safetensors/numpy/einops); Real-ESRGAN weights are BSD-3-Clause and download on first use viatorch.hub(never bundled). Kept OUT ofall(heavy + model download). - Tests for the model-running paths are limited to availability checks (multi-GB downloads). But the pure helpers inside ML-adjacent modules are unit-tested without any download and must stay that way:
_target_size(native-vs-downscale-cap-vs-upscale-floor,test_invisible_engine.py),humanizer.unsharp_mask/adaptive_polish(test_humanizer.py), and the MPS->CPU fallback control flow via mocked pipelines (test_img2img_runner.py, 100% cover). Don't skip these as "ML, needs a model" — onlyremove_watermark/the diffusion bodies do.
Key modules
Compact map. The full per-module detail (design decisions, tuned thresholds, calibration history, incident records, and the regression-guard map) lives in docs/module-internals.md — read the relevant section there before changing any module below.
noai/c2pa.py— C2PA reading.extract_c2pa_info(path)uses the official c2pa-pythonReaderfirst (core dep, any container;read_manifest_store_jsonreturns the WHOLE store JSON — active + ingredient manifests — so an AI marker on a parent manifest is seen), and falls back to the hand-rolled caBX/CBOR parser (has_c2pa_metadata/extract_c2pa_chunk/_extract_c2pa_info_png) for synthetic/partial blobs the validator rejects or a broken/absent wheel. The registry scan (issuer / source-type / SynthID / soft-binding) is shared by both paths via_populate_registry_fields, so the return-dict shape is identical. Do not reimplement chunk parsing; chunk reads are clamped to the remaining file size by design.extract_c2pa_chunk/inject_c2pa_chunkstay PNG-only (raw caBX bytes, test/extractor use).noai/constants.py— the singleC2PA_AI_VENDORSregistry (+C2PA_SOFT_BINDINGS) from whichC2PA_ISSUERS/SYNTHID_C2PA_ISSUERS/C2PA_IDENTITY_AI_ORGS/identify._ISSUER_PLATFORMare all derived. Add a new vendor as one registry entry; never edit the derived dicts and never add inline. A vendor'sasserts_ai=Trueflag means its mere presence asserts AI generation even without atrainedAlgorithmicMediadigital-source-type (a pure-generator brand with a distinctive issuer/generator string, e.g. Dreamina — ByteDance's international Jimeng brand, signed as "Bytedance Pte. Ltd." with a "Dreamina/x.y" claim generator and no source-type); NEVER set it for common-word issuers (Adobe/Google/OpenAI/Microsoft) that appear incidentally in unrelated bytes — those stay source-type-gated inidentify._attribute_platform.metadata.py—scan_head(path)is the shared (memoized) input for every C2PA/AIGC/IPTC byte scan; use it instead ofopen().read(1MB)for any new marker scan. Also home tosynthid_source,xai_signature,iptc_ai_system,aigc_label,huggingface_job,samsung_genai, andremove_ai_metadata(fail-safestrip_c2pa_boxes). A caller that REPORTS an outcome must usestrip_and_verify, notremove_ai_metadatadirectly -- the stripper is deliberately fail-safe (a file PIL cannot decode is copied through UNCHANGED rather than crashing), so its return value cannot distinguish a no-op from a real strip.strip_and_verifyre-scans the output and, when metadata survived butimage_iocan still decode the raster, normalizes the container and scans again; that recovery preserves pixels but drops standard metadata. A truly undecodable file keeps the surviving-marker result.metadata --removeandbatch --mode metadata|alluse this verified path.remove_ai_metadatais the SINGLE metadata stripper (the legacy PIL-re-encodingnoai/cleanerwas deleted; the diffusion core and the publicnoai.remove_ai_metadatare-export now point here). It strips losslessly per container: ISOBMFF (HEIC/AVIF/MP4) blanks tokens / strips boxes in place; JPEG uses_strip_jpeg_metadata_lossless— a marker-segment walk that drops the AI-bearing APP segments (C2PA APP11; XMP APP1 carrying C2PA, a China-AIGC token, OR an IPTCdigitalSourceType/ 2025.1 AI-disclosure marker; IPTC-IIM APP13) and scrubs AI EXIF tags via piexif, copying the entropy-coded scan verbatim so the pixels are bit-identical (no DCT re-encode). Detection<->removal parity across every marker placement is load-bearing — anything a scanner flags, the strip must reach, or a re-served file still reads as AI: (a) the APP1-XMP branch of_jpeg_app_carries_aichecks the IPTC marker sets too, not only C2PA/AIGC (the Instagram/MidJourney/Meta "Made with AI"digitalSourceTypelives in XMP, not the APP13 IIM record); (b) a bareAIGC{...}/{"AIGC":{...}}block in ANY JPEG APP segment — the specific C2PA(APP11)/XMP(APP1)/IPTC(APP13) checks FALL THROUGH to a generic_is_aigc_exif_valuedrop, so a bare AIGC in APP11 (the supported placement, NOT a C2PA manifest) is caught, not swallowed by the C2PA-only 0xEB branch — plus the same AIGC block in a STANDARD PNG text chunk value (e.g.Description, which_is_ai_keykeeps) is dropped on the value; (c) the China TC260{"AIGC":{...}}block in EXIFUserComment/ImageDescriptionis scrubbed by_scrub_ai_exif(Doubao producer + Tencent service-provider schemas); (d) the Samsung Galaxy AIPhotoEditor_Re_Edit_Datatrailer past the JPEG EOI is truncated by_strip_samsung_trailer(andsamsung_genaireads the file tail so a multi-MB photo's trailer past the 512 KB quick-scan window is still DETECTED). Pixels stay bit-identical throughout, so a--strip-metadataon a q100 removal output does NOT crush it back to q75; PNG/WebP re-saves are pixel-lossless (WebP written at cv2 lossless mode, quality 101 — quality 1-100 is lossy). The PIL-fallback save format is chosen by the source's CONTENT, not its file extension (_sniff_image_format, and the JPEG-lossless gate is content-gated too): inputs can be misnamed (a PNG served as.jpgis the common one), and routing on the extension re-encoded a lossless PNG/WebP into a real JPEG — a silent degradation that broke "work with originals". A misnamed lossless source (source-extension format != content) is preserved in its true format; a correctly-named source still honors a deliberate output-extension conversion (e.g.source.png -> output.jpg). Not yet handled: a 16-bit PNG is downconverted to 8-bit on the PIL re-save (rare; would need a byte-level PNG chunk stripper). Regression:tests/test_metadata.py::TestHasAiMetadata::test_strip_preserves_lossless_content_with_mismatched_extension.remove_ai_metadatais fail-safe on an undecodable image: a truncated/corrupt file (PIL raisesOSErrordecoding it; some inputs) is copied through UNCHANGED rather than crashing a direct library caller (a web worker would 500 on a partial upload), mirroringstrip_c2pa_boxes— we cannot strip what we cannot parse, but we never raise. Regression:tests/test_metadata.py::TestHasAiMetadata::{test_remove_ai_metadata_failsafe_on_truncated_png,test_strip_and_verify_normalizes_decodable_copy_through,test_strip_and_verify_reports_markers_in_undecodable_copy_through}. Regression:tests/test_metadata.py::TestHasAiMetadata::{test_jpeg_metadata_strip_is_pixel_lossless, test_jpeg_strip_removes_iptc_marker_in_xmp},TestSamsungGenai::{test_remove_strips_post_eoi_trailer, test_detects_trailer_past_scan_window}, the AIGC-EXIF/bare-APP removal tests, andtests/test_noai.py::TestISOBMFF::{test_blank_aigc_block_in_exif, test_blank_xai_signature_pair_in_exif}.exif_generatormatches a VALUE againstAI_GENERATOR_TOKENSacross EXIFSoftware/Make/Artist/ImageDescription, XMPCreatorTool, AND PNGtEXtchunks (Software/Source/Title/Description— NovelAI stamps there, not EXIF). Detection and removal must stay in parity: a generator that stamps an AI-shaped VALUE under a non-AI KEY (NovelAI'sTitle/Source) is dropped on removal by_is_ai_value(value-token match, mirrorsexif_generator), NOT by_is_ai_keyalone — else the cleaned file still reads as that generator. Add a new no-C2PA generator = oneAI_GENERATOR_TOKENSentry (use a distinctive token, e.g.reve.comnot barereve); detection and removal then both follow. Regression:tests/test_metadata.py::TestExifGenerator::{test_novelai_png_text_chunk_detected,test_novelai_removal_parity}.identify.py— separates file-backed metadata extraction (extract_provenance_evidence) from metadata-only verdict logic (identify_from_evidence); the compatibleidentifywrapper adds optional pixel-backed visible and invisible checks. Detection fromProvenanceEvidencemust never reopen the source. All paths produce oneProvenanceReport;is_ai_generatedis True or None, never asserted False.ProvenanceReport.ai_source_kindexposes the C2PA digital-source-type split —"generated"(trainedAlgorithmicMedia, fully AI) vs"enhanced"(compositeWithTrainedAlgorithmicMedia, a real photo with an AI-composited region), else None — so a caller branches full-frame scrub vs region-targeted clean (seenoai/tiling.feather_region_composite+WatermarkRemover.remove_watermark(region=...)). The sparkle provenance threshold is the SHAREDwatermark_registry.GEMINI_SPARKLE_TRUST_CONF(imported, not a private copy) so the provenance "is there a sparkle" verdict and the removal "take the sparkle" decision can never drift.import identifyis deliberately light (lazynoai/__init__, fits a 512 MB host) — keep heavy imports out (thewatermark_registryconstant import stays light: engines are lazy there). Add capture-camera tokens to_DEVICE_C2PA_PLATFORMonly when verified against a real C2PA file; editing-app/AI-device signer tokens go to_SIGNER_C2PA_PLATFORM; generator/issuer platforms toC2PA_AI_VENDORSinconstants.py. The IPTCdigitalSourceTypealgorithmicMedia(bare) is PROCEDURAL (an algorithm not trained on sampled data), NOT AI/ML generation, so it is deliberately absent fromIPTC_AI_MARKERS— flagging it madeidentifyassert AI +has_invisible_targetTrue, scrubbing clean procedural content (it is a distinct token fromtrainedAlgorithmicMedia, so real "Made with AI" labels are unaffected; regressiontest_metadata.py::...test_bare_algorithmic_media_not_flagged_ai). Integrity-clash detection is high-precision by design (only hard generator stamps feed it, source-grouped independence)._vendor_ofnormalizes ByteDance/Canva/ElevenLabs/Black Forest Labs (as well as OpenAI/Google/... ) so their C2PA claims participate in the clash check; the generic China TC260 AIGC label names no specific vendor, so when a TC260-applying vendor (ByteDance,_TC260_VENDORS) is co-attributed the label is attributed to it (a legit Doubao image carrying its own TC260 label must NOT clash), while a NON-TC260 vendor next to a TC260 label still clashes as a laundering tell. The vendor normalization must not introduce clashes on clean compatibility samples.watermark_registry.py— the single catalog of known visible watermarks (gemini / doubao / jimeng / qwen / kling / samsung / runninghub / baidu / liblib / jimeng_pill). Removal is LOCALIZE -> FILL for every mark: each mark is localized to a binary full-frame footprint mask (aLocalization), then ONE shared, swappable fill inpaints that mask viafill(image, mask, backend=...)(delegates toregion_eraser.erase). Reverse-alpha (the oldoriginal = (wm - a*logo)/(1-a)inversion of a captured alpha map + thin residual inpaint) is GONE for ALL marks; why it was dropped is recorded indocs/module-internals.md. Backends:cv2(classical inpaint, no deps, the floor),migan(MI-GAN ONNX, light, the memory-tight pick where LaMa will not fit),lama(big-LaMa ONNX, best quality, heavier, auto-preferred when a learned backend is available);auto= LaMa > MI-GAN > cv2, best available. The captured alpha maps (scripts/visible_alpha_solve.py) are still used to DETECT the marks and to shape the mask, but NOT for pixel recovery.--mark autoremoves EVERY detected mark in one pass viaremove_auto_marks(image, *, sensitivity="auto", provenance=frozenset(), backend="auto")(marks coexist -- a Jimeng-basic image has the top-left pill AND the bottom-right wordmark; a single-strongest pick would leave one). Three orthogonal axes:backend(the fill),sensitivity(how hard to trust a borderline mark:auto/strict, see theSensitivityliteral), andprovenance(vendor keys metadata confirms -- the evidence that drivesauto). Perception / decision / action are separated:_build_candidates(image)runs every detector at BOTH trust levels (strict + relaxed) and packages raw verdicts + features intoCandidates (no policy); the pure arbiterdecide(candidates, Context(sensitivity, provenance)) -> [Decision]makes every keep/drop call (per-markresolve_trust+ the assumed-trust floor + the pill gate) with no image/IO, so it is unit-testable in isolation; then each winner is localized -> filled. Do NOT put policy back into the engines (the one exception, the Gemini FP gate, stays ingemini_enginebecauseidentifyshares that confidence).detect_marks(..., provenance=frozenset())stays strict (identify verdict, precision over recall);KnownMark.remove/detect/localize(..., provenance: bool)take the already-resolved boolean. Howautodecides (this is metadata-INDEPENDENT for recall): the visual detectors are pixel-based and need no metadata; the recall gain comes from RELAXING the false-positive gate, not from metadata.strictnever relaxes (clean images untouched);autorelaxes a mark only on same-product evidence -- metadata provenance for that vendor OR a confidently detected sibling mark of the SAME product (_PRODUCT_OF; Doubao and Jimeng are both bottom-right ByteDance but distinct products, so they do NOT cross-relax).resolve_trustresolves TWO levels:confirmedbypasses the engine's false-positive gate, and onlyconfirmedhas evidence naming THAT vendor, which is exactly what the bypass is contracted to require (GeminiEngine.detect_watermark'strust_provenancedocstring: "external metadata already proves this is a Google generation"). Historical, kept as the reason the third level is gone: a removedassumedlevel letassume_aibypass the gate on the bare assertion an image is AI. It produced unacceptable false positives on clean camera images and was removed. A wrong relaxation only fills a small corner near-losslessly (the localize -> fill benign failure mode), which is what made a SMALL false-fire rate arguable; it was never a licence for a 60% one. Metadata provenance mapping (feedsauto, read bycli._visible_provenance): Google/Gemini C2PA issuer -> gemini; China-AIGC (TC260) label -> doubao/jimeng;samsung_genai-> samsung. Thejimeng_pillis capture-less. It uses a synthetic silhouette and a fixed top-left footprint._keep_pillrequires either a confirmed sibling wordmark or Jimeng provenance with a flat footprint. Do not loosen the flatness guard.assume_aiwas REMOVED (2026-07-19);--sensitivityis nowauto/strictonly. It relaxed every mark's FP gate on the bare assertion an image is AI -- which names no vendor and no location, exactly what the bypass requires -- and had no place in the model (detector finds -> remove; finds nothing -> leave alone; user SEES a mark -> act on that). It took_ASSUMED_CONF_FLOOR/assumed_floor_ok/ theassumedtrust level with it, collapsing the ladder tostrict/confirmed, and_keep_pilllost itssensitivityarg. Recall/precision on the unbiased sample are unchanged, so nothing on the default path moved. Replacement advice is per mark:erase --regionis sound by construction;--mark <text-mark> --no-detectis reasonable (forced mask = the real glyph blob, non-empty 13/13);--mark gemini --no-detectis NOT -- it falls back to a fixed slot that covered the true sparkle on only 31% of 97 missed sparkles, so 69% fill a clean corner AND report a removal that did not happen.cli._no_visible_mark_exitfollows that order and no longer suggests the removed mode. Migration raises loudly (validate_sensitivity, called fromapi.remove_visibleandContext.__post_init__) because aLiteralis unenforced at runtime and would silently downgrade a 0.15 caller toauto. Detection and removal-mask extraction must use the same front-end. The continuoustophatpath can detect a mark whose binary glyph blob is empty, so its fallback mask uses the detector's own best-match box. Keep the textured regression fixture when changing this path. Detector thresholds and geometry are calibrated per mark; do not port them between vendors without a fresh evaluation. New mark assets must be synthetic. Font-render the mark, calibrate it on local evaluation inputs, and commit only the synthetic silhouette. Never derive a committed asset from source images.gemini_engine.py— visible Gemini-sparkle detector + localizer (cv2/numpy, no GPU): top-K size-weighted fusion candidate selection (_SELECT_TOPK), corner-promote, false-positive gate (the provenance prior relaxes the gate + lowers the trust threshold when a Google/Gemini C2PA issuer confirms the vendor). White-core rescue: the FP gate demotes a low-gradient match (soft edges), but a real FAINT sparkle also has soft edges -- so the gate keeps a low-grad match that is a strong (conf ≥_SPARKLE_KEEP_CONF0.52), bright (margin), near-WHITE-core sparkle (_core_saturation≤_SPARKLE_WHITE_SAT0.20): a real sparkle core is white, a clean bright corner that shape-matches (sky/sun) is colored. This recovers ~14/20 metadata-stripped faint sparkles under the DEFAULT strict/auto (no flag, no metadata) at ~1.25% clean false-fire (baseline 0.55%); the ~0.51-scoring bright-bg FPs stay demoted (below 0.52). A learned classifier on the SAME features was measured WORSE than the tuned gate (2026-07 tier-1: MLP 86.7% recall vs 90.8% at equal FP), so the heuristic stays; a patch-CNN with richer features is the only lever left (roadmapped P2, low expected value -- the wall is fundamental). Detection scores the top-K size-weighted matches by full fusion (spatial+gradient+variance) and keeps the highest — NOT the raw-NCC argmax, which re-admits the tiny-patch FPs the size weight suppresses (the osachub 2026-06-12 sub-0.85 corner-sparkle regression; seedocs/module-internals.md). Keep the 0.85 corner-promote NCC gate; a margin/chroma-gated lower promote was measured and REJECTED 2026-06-11 (~33% FP on non-Google content). Removal is localize -> fill:footprint_maskreturns the sparkle footprint (the captured alpha thresholded LOW so the faint halo is included, then dilated by a sparkle-relative margin), and the sharedwatermark_registry.fillinpaints it. The captured alpha maps are used only to detect and to shape the mask, not for pixel recovery.detect_sparkle_confidencereuses a process-wide_shared_engine()singleton (lru_cache) — the engine holds only constant assets (captures, alpha maps, a precomputed 16..118 template ladder) and takes the image as an arg, so do NOT reconstructGeminiEngine()per call: that reloaded assets + recomputed alpha maps + rebuilt the template cache on every one of ~34kidentifycalls (−24% on the sparkle path once made a singleton, output byte-identical).detect_watermark/footprint_maskguardimage.size == 0beforeto_bgr, and return an empty (detected=False) result when no template scale fits (short side < 16 px), rather than dereferencing an empty candidate list._text_mark_engine.py— shared base for the text-mark engines (extracted 2026-06-09); the per-engine modules are config-only subclasses. Detection still matches the glyph silhouette (NCC, keys on glyph shape). The removal mask is TEMPLATE-FREE: it is the bounding box of the top-hat glyph blob (extract_mask), filled solid + dilated, so the shared fill inpaints the whole wordmark rectangle. This drops the fixed alpha-template placement, so a re-rendered or differently-placed mark is still masked; the captured alpha maps are now used only for the detection silhouette, not for removal. New text mark = aTextMarkConfig+ a thin subclass + one registry row. Gemini stays a separate engine (different model). The corner anchor iscorner=br/bl/tl/bc(tl added 2026-07-22 for runninghub, bc for liblib's centered wordmark);detect_frontendisbinary/tophat/gray(gray= raw-grayscale NCC for the faint mid-gray runninghub mark, added 2026-07-22; contrast-DEPENDENT, so its gates never port). The detection scale ladder is per-mark (TextMarkConfig.ladder, default(0.8, 1.0, 1.25)-- added 2026-07-21 for qwen's two size modes; the shared default is unchanged for every other mark, and densifying the SHARED ladder was measured and rejected, seedocs/verification-plan.mdB2).pill_engine.py— detects the capture-less Jimeng "AI生成" pill with a synthetic silhouette and removes it through the shared fill path. Its weak detector is registry-gated: a sibling wordmark may confirm it, while metadata-only removal also requires a flat footprint. Do not loosen those gates.doubao_engine.py/jimeng_engine.py/samsung_engine.py— thinTextMarkEnginesubclasses: Doubao "豆包AI生成" (bottom-right), Jimeng "★ 即梦AI" (bottom-right), Samsung Galaxy AI "✦ Contenuti generati dall'AI" (bottom-LEFT, locale-specific — Italian variant calibrated). Detection matches the glyph silhouette (NCC); removal localizes the glyph blob to a solid dilated box (extract_mask) and hands it to the shared fill. Calibration confirms that doubao and jimeng localize and remove cleanly, while clean compatibility images remain unchanged. Samsung detection is calibrated only for the Italian "Contenuti generati dall'AI" string (a pre-existing limit, unchanged by the localize -> fill refactor but now surfaced because detection gates removal): non-Italian Samsung locales are not detected, and thus not removed, even though the fill mask itself is locale-independent; other locales need their own detection silhouette (the locale string font-rendered + calibrated on real positives), NOT an app capture.qwen_engine.py— detects the Qwen "千问AI生成" bottom-right text mark with a synthetic silhouette and per-mark geometry. Keep its calibration independent from similar CJK marks.kling_engine.py— detects the Kling "可灵AI 3.0" bottom-right text mark. It uses short-side geometry and a synthetic silhouette.yuanbao_engine.py— detects the standard two-line Tencent Yuanbao "元宝 / AI生成" bottom-right mark through polarity-independent local contrast. The one-line overlay variant remains unsupported.runninghub_engine.py— detects the faint top-left RunningHub mark through grayscale silhouette matching. The anchor gate and detector-owned match box are part of its false-positive control.baidu_engine.py— detects the Baidu "百度 AI生成" bottom-right mark. Rival margins separate it from similar CJK marks, and its custom footprint includes the adjacent tag.liblib_engine.py— detects the bottom-center LibLibAI wordmark with a synthetic silhouette, contrast gating, and detector-owned footprint.region_eraser.py— universal region eraser (eraseCLI) and the shared fill backend behindwatermark_registry.fillfor the visible localize -> fill removal. Three backends:cv2(default for the user-directederasecommand, no deps, the floor),migan(MI-GAN ONNX, extramigan, MIT, ~28 MB / ~0.19 s, the memory-tight learned tier),lama(big-LaMa ONNX, extralama, ~200 MB / ~4.7 GB peak, best quality but too heavy for a minimal worker). The visible registry'sautoresolution is LaMa > MI-GAN > cv2; select MI-GAN or OpenCV explicitly when memory matters. Bothmiganandlamacrop a padded region around the mask before inference and paste only masked pixels back, so peak RAM is bounded by the MARK size, not the image (migan~0.6-0.9 GB regardless of upload size — feeding the whole frame scaled it to ~2.4 GB at 25 MP;miganfeeds the crop at native resolution,lamaresizes to its fixed 512²). Measured end to end 2026-07-20 (scripts/resource_ceilings.py, fresh process per cell, 1 MP → 25 MP):migan603 → 775 MB andlama4679 → 4779 MB, both flat in input size — the crop-around-the-mask design holds and both documented figures reproduce.cv2is the only backend that GROWS with the input (74 → 440 MB, 5.9x) because it inpaints the full frame rather than a crop; still the cheapest tier, but size it for the largest upload accepted. Cold wall time 0.02-0.12 s (cv2) / ~0.6 s (migan) / ~3.8 s (lama), model load included. (The harness's own no-op check originally allocated a full-frame temp before reading peak RSS and inflated these by up to 17% at 25 MP; it now compares only the mask box. The conclusion survived re-measurement, the digits moved.) MI-GAN mask polarity is INVERTED (0=hole/255=known) vs this package's 255-erase convention;erase_miganinverts before feeding the model (feeding 255=hole regenerates the whole frame into stripes — verified). Both ONNX models download on first use, never bundled. Theerasecommand keeps its own--backend/--inpaint-method(unchanged).invisible_watermark.py— decodes the OPEN DWT-DCT watermarks (SD / SDXL / FLUX) viaimwatermark(extradetect, pulls torch). Fragile two ways: (1) does not survive JPEG re-encode/resize; (2) carrier-fragile on a broad class of pristine images -- a clean encode->decode round-trip recovers 48/48 on chatgpt/firefly/random but FAILS (28-39/48, below the_MATCH_48=44 gate) on the FLUX fox, doubao, a flat FLUX generation, AND a clean synthetic flat fill with no watermark. The failure does NOT track texture; it goes with a degenerate all-ones decode that is a CARRIER ARTIFACT, not a watermark (synthetic clean image reproduces it). Sodetect_invisible_watermarkis positive-only: trust a hit; aNoneis inconclusive unless a same-carrier positive-control embed first recovers >=44. Verified 2026-06-19; full caveat indocs/watermarking-landscape.md.trustmark_detector.py— Adobe TrustMark open decoder (extratrustmark). Do NOT remove the JPEG re-encode false-positive gate — a lone TrustMark hit without it is almost always content noise.noai/watermark_remover.py—WatermarkRemoverwith four diffusion pipelines selected by the explicitpipelinector arg, never inferred frommodel_id:sdxl(plain SDXL img2img),controlnet(SDXL + canny ControlNet, the compatibility and cost DEFAULT since 2026-06-09),qwen(Qwen-Image 20B img2img), andqwen-zimage(delegates to the fixed two-stage runtime below). Removal comes from img2img strength. Both SDXL loaders passadd_watermarker=False; diffusers otherwise re-stamps an open SDXL DWT-DCT watermark. Qwen's certified floors and fidelity results remain as documented below. The baseqwenprofile stays the manual text lane;qwen-zimageis the recommended high-quality manual mode, especially for face identity, while remaining experimental rather than an auto-router.noai/qwen_zimage_pipeline.py— CUDA-only Qwen-Image-2512 Lightning + DiffSynth Canny full-frame regeneration, followed by YuNet face boxes, SAM masks, and Z-Image Turbo regeneration from original face crops. Ports both adaptive denoise formulas from Synthid-Bypass v2, then scales the face result by 0.5 because this runtime lacks the reference latent noise-mask feather and uses a different sampler/compositing path; paired face evaluations and both provider oracles certified the scaled value while the global stage stayed unchanged. The active upstream face path is YOLO + SAM; this port keeps its center-point and box prompts, proposal selection, detector-box intersection, crop factor, and paste feather while replacing YOLO with YuNet to avoid an AGPL runtime. The port is architectural, not bit-identical: it uses full safetensors instead of GGUF, DiffSynth samplers instead of the Comfy sampler pairs, and no latent detailer feather. DiffSynth input pixels, Canny control, and explicit dimensions must share the same /16 grid. SAM pixels follow the model dtype, geometric prompts stay float32, and bfloat16 outputs convert through float32 before NumPy. The YuNet download verifies its SHA-256. Separateqwen-zimageextra; fixed four-step global and eight-step face schedules; no custom--model;--tileapplies only to the global stage, followed by one full-frame face stage; CLI adaptive polish defaults off.InvisibleEngine.preload(global_only=True)warms Qwen and YuNet while leaving Z-Image and SAM lazy until a face is detected; the defaultpreload()remains a full preload. GPUs with at least 64 GiB VRAM keep the face stack resident, while smaller devices retain CPU offload. Fixed prompt embeddings are cached only when they do not depend on an edit image. The exact seed-0 release candidate passed the corresponding provider-oracle checks; broader seeded text, face, and tiled-output certification remains open.noai/tiling.py— sliding-window tiled diffusion for large inputs (CLI--tile). The SDXL, ControlNet, and base Qwen paths branch torun_tiledwhentileis set AND the long side exceedstile_size, refactoring the single-pass_generateinto a per-tile_generate_one(the ControlNet edge map is rebuilt per tile inside it).qwen-zimageinstead callsrun_tiledonly around its global Qwen stage, blends the tiles, then runs one full-frame face stage. Pure helpersplan_tiles(uniform-size tiles, last one flush to the edge) andfeather_weights(strictly-positive separable taper -> partition-of-unity blend) are unit-tested without the model. Also home tofeather_region_composite(base, regenerated, box, *, feather)— the pure region-targeted compositor for AI-enhanced composites (ai_source_kind == "enhanced"): blends the regenerated AI box back over the original with a feathered seam, leaving the real photo OUTSIDE the box pixel-exact. It backsWatermarkRemover.remove_watermark(region=...)(regenerate ONLY the AI region, not the whole frame); the no-model lossless region path staysregion_eraser.erase. New tile/region-blend tuning goes in these pure helpers; do not inline blend math into the runner.auto_config.py+ the content-detection layer were REMOVED 2026-06-09;--autois a deprecated no-op (controlnet is the default pipeline and adaptive polish is ON by default for the original profiles, whileqwen-zimageleaves it off to preserve the upstream two-stage output).upscaler.py— optional Real-ESRGAN pre-diffusion super-resolution for small inputs (extraesrgan, spandrel only). Manual opt-in; the default--upscalerstayslanczosand the engine always falls back to Lanczos on absence/error. ESRGAN can degrade faces and thin text.image_io.py— centralizes Unicode-safe image IO, alpha preservation, content-based format sniffing, and HEIC/AVIF fallbacks. Callers must checkimwritesuccess. No-op visible removal preserves original bytes when the output format is unchanged.api.py— the high-level convenience API, re-exported lazily at the package top level via__init__.__getattr__(PEP 562, soimport remove_ai_watermarksstays cheap):remove_visible(source, output=None, *, sensitivity="auto", backend="auto", strip_metadata=True, write_noop=True) -> (result_bgr, [labels])(source = path OR BGR ndarray; a PATH auto-reads metadata provenance and preserves alpha, an ARRAY does neither;write_noop=Truewrites a clean passthrough copy when nothing is removed,Falseleavesoutputuntouched so a "no mark = produce nothing" caller like the CLIvisiblecommand does not clobber a pre-existing file there) andvisible_provenance(path) -> frozenset[str](the single metadata→vendor-keys mapper;cli._visible_provenanceis a thin None-guarded wrapper over it).remove_visibleis the ONE path the CLI and library share —cli.cmd_visible's--mark autobranch delegates entirely to it (read → provenance →remove_auto_marks→ write →strip_metadata), so there is no CLI-vs-library drift;strip_metadatadefaults True to matchvisible --strip-metadata. This is where a library caller should start — NOT the engines directly (GeminiEngine/TextMarkEnginehave noremove_watermarkany more; removal is registryremove_auto_marks/KnownMark.remove; the old single-strongestbest_auto_markis gone — removal takes EVERY mark).identifyis NOT top-level re-exported (it collides with theidentifysubmodule); usefrom remove_ai_watermarks.identify import identify.
For the Doubao alpha-distillation history (why content-image reverse-alpha distillation fails by physics and controlled captures were required), see docs/research-doubao-distillation.md.
Watermarking landscape
Who embeds what (C2PA / IPTC / EXIF / TC260 AIGC / xAI signature / open and proprietary invisible watermarks), whether each is locally detectable, the C2PA 2.4 durable-credentials implications, and the regulatory driver table live in docs/watermarking-landscape.md (research 2026-05-24, updated through 2026-06-10). Read it before adding a new identify signal, vendor token, or metadata marker. See identify.py for what we read today.
Known limitations
Compact list. Full measurements, incident history, and oracle-validation runs live in docs/known-limitations.md — read the relevant section there before changing the diffusion pipelines, strength defaults, resolution handling, or metadata coverage.
- Visible-mark fill quality is background/backend-dependent. The fill only touches the mark footprint (no outside-box damage) and whether the mark is removed is fill-independent — cv2/MI-GAN/LaMa all strip the shape; only the recovered region's quality differs. Flat backgrounds: all clean (cv2 often crispest). Textured/regular-structured (fabric, grid): cv2 smears, MI-GAN can ghost/hallucinate, LaMa best. The old reverse-alpha recovered true pixels so it was sometimes cleaner on structure, but localize -> fill trades that for robustness (moved/re-rendered marks, no per-mark capture);
auto= LaMa > MI-GAN > cv2 with a one-time cv2-fallback warning. Head-to-head vs v0.12.1 on the full visible set: doubao/jimeng identical (100%/100%), gemini strict coverage a few points lower (the metadata-stripped faint ones now mostly recovered by the default white-core rescue in the gemini FP gate), clearance ~98% both. Detail indocs/known-limitations.md. invisibleprocesses at native resolution for inputs >= 1024px long side and auto-upscales smaller inputs to a 1024px floor (--min-resolution 0disables;--max-resolution Nis an opt-in cap to bound GPU/MPS memory). MPS OOM is memory-tier dependent, not a hard limit: ~24 GB unified memory falls back to CPU (slow but weight-identical output), 32 GB runs native on MPS. The native-vs-cap-vs-floor decision lives in the pure helperinvisible_engine._target_size— keep the logic there, unit-tested without the model. For large inputs that OOM,--tileis the lossless alternative to--max-resolution: sliding-window diffusion at native resolution, each tile near SDXL's 1024 training size, feather-blended over the overlap (noai/tiling.py). It only engages when the long side exceeds--tile-size; the geometry (plan_tiles) and the blend window (feather_weights) are pure and unit-tested (tests/test_tiling.py). Caveat: each tile is an independent low-strength regeneration, so at the certified removal strengths (0.20-0.30) tile drift is minimal but not zero; tiling is a memory workaround, not a quality upgrade over a single native pass.- fp16 VAE black-output (issues #29/#41): the fp16-fixed SDXL VAE (
madebyollin/sdxl-vae-fp16-fix) is swapped in for the default SDXL checkpoint on cuda/xpu fp16, plus a model-agnostic backstop that detects a degenerate (all-black) fp16 output and re-runs once in fp32. cpu/mps run fp32 and never reproduce the bug. - Pyright first run is slow (2-3 min) due to ML deps (torch/diffusers/transformers stubs); full-project
uv run pyrightcan stall for many minutes — scope it to changed files. - A third-party PIL plugin autoload (e.g. an HEIF/AVIF plugin) can raise a non-OSError (
ModuleNotFoundError), notUnidentifiedImageError, when opening a file. Code that opens user-supplied or unknown-format files shouldexcept Exception, not justOSError/UnidentifiedImageError. - rich was dropped: the CLI + analysis scripts print plain text (
click.echo/ thescripts/_plain_console.pyshim).richis NOT a dependency — importing it breaks the core+dev CI sync; new scripts must use the shim. No Unicode glyphs / colors / progress bars in CLI output by design. - HEIC/AVIF are decodable on BOTH paths now: the pixel/removal path via the
image_io.imreadPillow fallback (+ corepillow-heif), and metadata detection via a plugin-free binary scan. C2PA removal in those containers (and MP4/MOV/M4V) isnoai/isobmff.py; JPEG-XL stays metadata/strip-only (Pillow can't decode it withoutpillow-jxl, not a dep). Non-ISOBMFF audio/video (WebM/MP3/WAV/FLAC/OGG) strips losslessly via ffmpeg on PATH. On the ISOBMFF pathremove_ai_metadataroutes to the container branch and never runs the JPEG_scrub_ai_exif, soisobmff.blank_ai_exif_tokensis the ONLY EXIF scrubber there and must stay in PARITY with it: it blanks in place (same-length space overwrite, piexif-validated so a coincidental II/MM run in pixels is ignored — noiinf/ilocsurgery, mirrorsblank_ai_xmp_packets) an AI-generator token inSoftware/Make/Artist/ImageDescription, the China TC260{"AIGC":{...}}block inImageDescription/UserComment(via_is_aigc_exif_value), AND the xAI/GrokSignature:+ UUID-Artistpair — leaving camera/editor EXIF intact. Still NOT built: Resemble PerTh audio detection (no presence/confidence flag exists). - SynthID technical reference:
docs/synthid.md— primary-source-cited doc covering mechanism (post-hoc encoder/decoder pair, 136-bit payload at 512x512, pixel-space, model weights NOT modified), robustness numbers (arXiv:2510.09263: ~99.98% TPR@0.1%FPR across 30 transforms including JPEG/crop/resize/color/noise), removal attacks and forensic detectability (arXiv:2605.09203: all 6 attacks detectable at >98% TPR@1%FPR), detectability limits (no public decoder, metadata-proxy only), oracle scope, and adoption landscape. Read that doc first before adding notes here. - SynthID detection is metadata-only. No local pixel detector is possible by design (Google's decoder is proprietary, trusted-testers only); we read the C2PA companion proxy, which goes quiet once metadata is stripped — a quiet proxy is not proof the pixel watermark is gone. Each vendor has its OWN oracle and it detects only that vendor's content: the Gemini app "Verify with SynthID" for Google,
openai.com/verifyfor OpenAI. Validate the OpenAI arm FIRST —openai.com/verifyis more accessible (fewer per-check restrictions) and the strongest automation candidate (Playwright / Chrome MCP); the Gemini flow is more manual. Ordering/throughput choice, not a substitution (seedocs/synthid.md). SynthID survives JPEG re-encode, so GitHub issue attachments remain valid pixel-watermark test subjects. Every spectral/phase detection approach evaluated (reverse-SynthID, our own probes) works only on controlled solid fills, never on real content. - External AI-vs-real classifier models are out of scope (decided 2026-05-24): per-generator, degrade off-distribution, and our own light SDXL pass would likely defeat them. Detection stays local + signal-based.
- Default strength is VENDOR-ADAPTIVE, one ladder for BOTH pipelines (since 2026-06-09):
resolve_strength(strength, vendor)picks OpenAI 0.10 / Gemini 0.15 / unknown 0.15 when--strengthis unset (the 2026-06-14 lowering from the 2026-06-04 cert floors of 0.20/0.30 — the single source of truth iswatermark_profiles.py, and the full cert/lowering history is indocs/known-limitations.md); explicit--strengthalways wins. Removal at low strength is content x pipeline dependent, and near-threshold removal is SEED-NON-DETERMINISTIC — pick a strength with margin and oracle-revalidate per content type. controlnetis the default pipeline;--pipeline sdxlis the lighter opt-down. Neither pipeline clears all content at low strength (photoreal survives controlnet, flat graphics survive sdxl — the lever is higher strength). A removal-priority caller MUST oracle-validate strength across content types; prod recipe: controlnet + per-vendor floor + FIXED seed. Forensic-stealth caveat (arXiv:2605.09203): defeating the SynthID verifier is NOT forensic invisibility — removal-processed images are flaggable at >98% TPR@1%FPR.