Victor KuznetsovandClaude Opus 5 2f72257996 Fix what the verification pass found in the knob removal
An adversarial review of 52b2c11 (five independent audits, each finding put to
two skeptics, plus a completeness critic) found four defects that commit
introduced and several stale claims it should have caught.

The install hint no longer installs -- again. Folding five hints into one
INVISIBLE_EXTRA constant dropped the shell quoting the originals had, so the
printed remediation was `pip install remove-ai-watermarks[qwen-zimage]`. Bare
brackets are a glob in zsh, the macOS default shell: it dies with "no matches
found" before pip runs. That is the exact failure 52b2c11 existed to stop
producing, reintroduced in a different form by a bulk replace. The constant is
quoted now, and a test asserts the quotes rather than the bare substring -- the
old assertions passed either way, which is why nothing caught it.

Three tests were not guarding what they claimed:

- The commit's headline behaviour change, per-profile polish resolution inside
  the engine, had no test at all. Rebinding resolve_adaptive_polish to the
  pre-commit `bool(value)` left the full suite green. Now covered by a test that
  drives the real engine and observes whether humanizer.adaptive_polish ran;
  that mutation now fails it.
- TestAvailability still asserted the pre-commit (torch, diffusers) contract, so
  in a diffusion-only environment it was simply wrong, and comparing each gate to
  a tuple copied from itself could never catch the two gates disagreeing -- the
  drift the shared REMOVAL_MODULES was introduced to prevent. Replaced with a
  test that simulates each module's absence and requires BOTH gates to close.
- Both CUDA-refusal guards skipped in every environment, including CI: they were
  gated on the diffusion stack, which no CI job installs. The refusal fires
  before any torch attribute is read, so they now run everywhere; only the dtype
  assertion keeps its skip.

Also: the retired-knob test covered `invisible` but not `all` or `batch`, though
all three declared those options separately; and smoke_matrix.py still called
remove_watermark(region=...), a parameter 52b2c11 deleted, with the resulting
TypeError swallowed into a skip by a broad except.

Stale documentation the previous sweep missed: known-limitations still described
an MPS out-of-memory fallback and a lighter-pipeline escape that no code can
produce; module-internals declared Canny thresholds of 100/200 as a compatibility
contract while the code uses 13/64, attributed enable_model_cpu_offload to
deleted profiles, and still warned that the engine and CLI defaults differ (this
commit's predecessor made them identical); cli.md gated `all` on the `diffusion`
extra; python-api claimed "cuda" was the only accepted explicit device when
"auto" is too. The claim that `device` is not a parameter was wrong in both
module-internals and .claude/rules/development.md -- it is one, deliberately, and
now says so. `--cpu-offload` help and the pipeline's CUDA guard both still
pointed at MPS.

Not fixed here, reported instead -- both are outside this repo:
- ComfyUI-remove-ai-watermarks nodes.py:332 passes num_inference_steps and
  guidance_scale (plus min_resolution/upscaler from bf4bfc1). distribute.yml's
  comfyui job runs on every release and fails the release if the node sync fails,
  so 0.25.0 needs that node updated first.
- raiw-app modal_app.py:422-425 forwards the same two kwargs into
  remove_watermark. Latent: it is pinned to 1a77e24 and nothing supplies a value
  today, so it fires on the next pin bump.

pre-commit: 1) maintain.sh - exit 0 (1093 tests, Pyright 0 errors, no
vulnerabilities); 2) /simplify - not re-run, this commit is the applied output of
a five-dimension adversarial review; 3) docs sync - grepped MPS/mps, the extras
names and every symbol touched across README, docs/, scripts/, .claude/; updated
6 docs; 4) CLAUDE.md - corrected the device claim in .claude/rules/development.md
and added the shell-quoting rule

Verified by execution, not assertion: smoke_matrix --quick 51 pass / 0 fail,
_knob_rows driven directly 10 pass / 0 fail / 7 skip (no CUDA), the install hint
rendered and round-tripped through zsh, and each new test confirmed to fail
under the mutation it is meant to catch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 16:31:39 -07:00

Remove AI Watermarks

Remove AI provenance marks from images and video you generated yourself:

  • known visible labels such as the Gemini sparkle and vendor text marks;
  • invisible pixel watermarks through diffusion regeneration;
  • C2PA, EXIF, XMP, IPTC, and related AI metadata.

Video support covers provenance identification, complete visible-plus-metadata cleaning, directory batches, visible Sora, Veo, Seedance, Dola, Hailuo, and Kling mark removal, and oracle-certified VAE regeneration for video SynthID removal.

Try it online at raiw.cc if you do not want to install Python or run diffusion models locally.

PyPI Python Downloads License Tests Sponsor

This project is for lawful use on content you own. It does not target stock agency previews or other watermarks that protect third party paid content. See scope, safety, and legal notes.

Choose what you want to do

Goal Command GPU
Find provenance signals and watermarks identify No
Remove known visible AI marks visible No
Erase a region you select erase No
Strip AI metadata metadata No
Identify supported video provenance video identify No
Remove visible marks and AI metadata from video video all No
Strip AI metadata from video video metadata No
Remove a registered visible AI mark from video video visible No
Process a directory of videos video batch Depends on mode
Remove video SynthID with the certified VAE profile video invisible Recommended
Regenerate an image to disrupt invisible watermarks invisible Recommended
Run visible, invisible, and metadata removal all Recommended
Process a directory batch Depends on mode

Installation modes

Need Install
Metadata inspection and stripping remove-ai-watermarks
Visible detection and removal remove-ai-watermarks[visible]
Visible video processing remove-ai-watermarks[video]
Video SynthID removal remove-ai-watermarks[video,diffusion]
Torch-free DWT-DCT detection remove-ai-watermarks[detect]
Invisible image removal (needs CUDA) remove-ai-watermarks[qwen-zimage]
Every production feature remove-ai-watermarks[all]

Lower-level and specialized extras include pixels, heif, trustmark, migan, lama, and diffusion. The installation guide documents their exact dependency composition and model requirements.

Quick start

Install the metadata-focused default CLI:

uv tool install remove-ai-watermarks

Inspect an image:

remove-ai-watermarks identify image.png

For visible watermark removal, install the pixel dependencies:

uv tool install --force "remove-ai-watermarks[visible]"

Then remove a known visible mark and AI metadata:

remove-ai-watermarks visible image.png -o clean.png

Strip metadata without running visible inpainting or diffusion:

remove-ai-watermarks metadata image.png --remove -o clean.png

Inspect or remove AI metadata from an MP4, MOV, M4V, WebM, MKV, AVI, or FLV file:

remove-ai-watermarks video metadata input.mp4 --check
remove-ai-watermarks video metadata input.mp4 --remove -o clean.mp4

The metadata command does not transcode video or audio streams. When -o is omitted it writes <source>_clean and preserves the original. MP4 and MOV inspection includes the native TC260 AIGC tag in moov.udta.meta.keys/ilst, including a moov placed after the media payload. MKV and WebM inspection reads the normative Segment.Tags.Tag.SimpleTag placement. AVI uses LIST/INFO/AIGC, while FLV uses script.onMetaData.AIGC. The non-ISOBMFF formats are remuxed with stream copy for removal.

Use the product-oriented video path to identify or clean a file:

uv tool install --force "remove-ai-watermarks[video]"
remove-ai-watermarks video identify input.mp4
remove-ai-watermarks video all input.mp4 -o clean.mp4

video all removes a stable registered visible mark when present and always strips verified AI metadata. If neither signal is found, it still writes a same-container passthrough, so application callers get one predictable output contract. Proprietary invisible-video removal is excluded by default. --invisible opts into the lossy, oracle-certified video SynthID profile.

Process a directory with the same contract:

remove-ai-watermarks video batch ./videos --mode all

Remove a supported visible video mark:

remove-ai-watermarks video visible input.mp4 -o clean.mp4
remove-ai-watermarks video visible veo.mp4 --mark veo -o veo_clean.mp4
remove-ai-watermarks video visible seedance.mp4 --mark seedance -o seedance_clean.mp4
remove-ai-watermarks video visible dola.mp4 --mark dola -o dola_clean.mp4
remove-ai-watermarks video visible hailuo.mp4 --mark hailuo -o hailuo_clean.mp4
remove-ai-watermarks video visible kling.mp4 --mark kling -o kling_clean.mp4

This path scans the complete sequence before changing pixels. It accepts only a mark that repeats at a stable position across adjacent frames, then reuses the same OpenCV, MI-GAN, or LaMa fill backends as image removal. Audio is copied without re-encoding and is allowed to reach its natural end; the video stream is transcoded because its pixels change. By default, a guarded optical-flow pass motion-aligns the preceding accepted fill and blends it only when the nearby source context agrees; use --no-temporal-consistency to disable it. The encoder preserves supported 8-bit source chroma sampling, color tags, and MP4/MOV track timescale instead of relying on ffmpeg's implicit raw-BGR defaults. Variable frame intervals are preserved through a timestamped in-memory NUT bridge instead of being flattened to the average frame rate. Non-zero source start timestamps are retained together with the copied audio offset. The default --mark auto scans all providers in one decode pass and selects the first stable match in the specificity order shown below. Pass an explicit mark to restrict detection to one provider. Sora covers the moving Sora 2 mascot and wordmark. Veo covers both the current four-point diamond and the legacy Veo text. Seedance covers the fixed boxed AI label, Dola covers the fixed Dola AI text, Hailuo covers the composite MINIMAX | hailuo AI label, and Kling covers the bottom-right KLING AI label with its version suffix. A completed encode is published atomically. No output is written when no stable mark is found. HDR, PQ/HLG, and greater-than-8-bit inputs are rejected before encoding rather than silently reduced through OpenCV's 8-bit BGR boundary.

Remove video SynthID:

uv tool install --force "remove-ai-watermarks[video,diffusion]"
remove-ai-watermarks video invisible input.mp4 -o clean.mp4

This path regenerates the complete sequence with one latent-noise field shared across time, copies complete audio, strips source metadata, and publishes the completed encode atomically. The default noise_std=0.15 profile passed both the two-carrier calibration and a complete public eight-second Veo oracle check. Google does not publish a local decoder, so a fresh provider check remains useful for unusually important files or after provider changes, but it is not a product result state.

For invisible watermark removal, install the qwen-zimage extra. An NVIDIA GPU is required: both profiles are CUDA-only, and there is no CPU or MPS fallback.

uv tool install --force "remove-ai-watermarks[qwen-zimage]"
remove-ai-watermarks invisible image.png -o clean.png

If the local detectors cannot confirm an invisible watermark but you know the image came from an AI generator, add --force:

remove-ai-watermarks invisible image.png -o clean.png --force

See the installation guide for Homebrew, uv, optional features, and development setup.

Examples

Visible Gemini mark

Before After
Image with a visible Gemini watermark Image after visible watermark removal

High quality invisible removal

qwen-zimage is the default profile: a Qwen-Image-2512 Lightning pass under Canny ControlNet, followed by SAM-masked Z-Image repair of any detected face. The alternative, sdxl-zimage, swaps the global stage for SDXL and keeps the same face stage. Both are CUDA only.

uv tool install --force "remove-ai-watermarks[qwen-zimage]"
remove-ai-watermarks invisible image.png -o clean.png --force
OpenAI example before OpenAI example after
OpenAI portrait grid before qwen-zimage OpenAI portrait grid after qwen-zimage
Gemini example before Gemini example after
Gemini sign before qwen-zimage Gemini sign after qwen-zimage

These exact output files were checked with the matching provider verifiers. That result applies to these files, not to every seed, image, or future watermark version.

Common recipes

Remove every detected visible mark

remove-ai-watermarks visible image.png -o clean.png

The default --mark auto checks all registered visible marks and removes every match. If the mark is visible to you but the detector misses it, select its region explicitly:

remove-ai-watermarks erase image.png \
  --region 1640,1930,400,100 \
  -o clean.png

--region uses x,y,width,height and may be repeated.

Use a learned fill backend

The visible extra uses OpenCV inpainting when no learned backend is installed. For more difficult backgrounds, the learned-backend extras include the same pixel dependencies automatically:

uv tool install --force "remove-ai-watermarks[migan]"
remove-ai-watermarks visible image.png -o clean.png --backend migan
uv tool install --force "remove-ai-watermarks[lama]"
remove-ai-watermarks visible image.png -o clean.png --backend lama

Reduce CUDA memory use

remove-ai-watermarks invisible image.png -o clean.png \
  --cpu-offload --force

CPU offload lowers CUDA memory pressure by moving model components between CPU and GPU, at the cost of speed.

Process a directory

remove-ai-watermarks batch ./images --mode visible
remove-ai-watermarks batch ./images --mode all

What the tool can recognize

Visible mark support includes:

  • Google Gemini and Nano Banana sparkle;
  • Doubao, Jimeng, Qwen, Kling, Baidu, LibLibAI, and RunningHub labels;
  • one calibrated Samsung Galaxy AI label variant.

Metadata and provenance inspection covers C2PA, EXIF, XMP, IPTC, common generator parameters, China TC260 AIGC labels, and several vendor specific signals. Optional decoders add support for open DWT-DCT watermarks and Adobe TrustMark.

The exact support matrix, including important locale and detector limits, lives in supported signals.

How it works

Visible removal follows three steps:

  1. Detect a registered mark in its expected area.
  2. Build a mask around the mark.
  3. Fill only the masked region with OpenCV, MI-GAN, or LaMa.

Metadata removal uses format aware stripping. JPEG metadata removal preserves the encoded image scan instead of recompressing it. Native MP4/MOV TC260 values are blanked without changing box sizes or media offsets. Other supported containers use their corresponding metadata path.

Invisible removal is different. It regenerates the image through a diffusion pipeline to disrupt pixel and frequency domain watermarks. This changes the image and cannot guarantee that a proprietary verifier will reject every output.

See supported signals and known limitations for the full technical boundary.

Python API

The visible-removal API requires remove-ai-watermarks[visible].

import remove_ai_watermarks as raiw

result, removed = raiw.remove_visible("watermarked.png", "clean.png")
print(removed)

provenance = raiw.identify_video("input.mp4")
report = raiw.inspect_video_metadata("input.mp4")
complete = raiw.remove_video_all("input.mp4", "clean.mp4")
batch = raiw.remove_video_batch("videos", "videos_clean")
cleaned = raiw.remove_video_metadata("input.mp4")
synthid_cleaned = raiw.remove_video_invisible("input.mp4", "synthid_clean.mp4")
visible = raiw.remove_video_visible("input.mp4", "clean.mp4")
print(visible.mark)
veo = raiw.remove_video_visible("veo.mp4", "veo_clean.mp4", mark="veo")
seedance = raiw.remove_video_visible(
    "seedance.mp4",
    "seedance_clean.mp4",
    mark="seedance",
)
dola = raiw.remove_video_visible("dola.mp4", "dola_clean.mp4", mark="dola")

The high level API accepts a file path or a BGR NumPy array. For path inputs it also reads provenance metadata, preserves alpha, and can strip AI metadata from the written result.

See the Python API guide for visible removal, provenance inspection, metadata stripping, and diffusion usage.

ComfyUI

The separate ComfyUI Remove AI Watermarks package provides nodes for visible removal, detection, region erasing, and invisible removal.

Important limitations

  • A missing local signal means unknown, not clean. Proprietary pixel watermarks may remain after metadata has been stripped.
  • Visible removal reconstructs a small region. Results depend on the background and selected fill backend.
  • Invisible removal changes the whole image and may alter faces, text, or fine detail.
  • Visible video removal recognizes the moving Sora 2 wordmark, the current Veo diamond plus legacy Veo text, the Seedance boxed AI label, and the fixed Dola, Hailuo, and Kling labels. It does not recognize the older Sora Turbo corner swirl or unregistered layouts from those providers. The classical OpenCV backend can smear structured backgrounds; use MI-GAN or LaMa when recovery quality matters.
  • Video SynthID regeneration changes resolution, frame rate, and image detail. The shipped profile is oracle-certified, but no public local decoder can certify an arbitrary output at runtime. Recheck unusually important outputs after provider changes.
  • Invisible-watermark removal requires CUDA. Both profiles refuse any other device at construction rather than falling back to one that cannot run them. Visible removal, metadata stripping and identify still run anywhere.
  • Provider watermark systems can change. Validate important outputs with the provider's own verifier when one is available.

The shipped video invisible command uses the certified noise_std=0.15 profile. The companion scripts/video_synthid_sweep.py research harness builds a matched re-encode control plus VAE-regenerated candidates and leaves the verifier verdict blank:

uv run --extra video --extra diffusion python scripts/video_synthid_sweep.py input.mp4 -o sweep/

The control must still be SynthID-positive before a negative candidate can count as removal evidence. In the 2026-07-29 two-clip calibration, both matched controls were positive in Gemini's built-in SynthID verifier; the stronger candidate was negative on both carriers, while a weaker candidate was negative on one. A later adversarial follow-up that asked ordinary Gemini to reinterpret the pixel result returned UNAVAILABLE; that follow-up was not a verifier rerun and does not invalidate the built-in verdicts. A 2026-07-31 full-clip check on a public eight-second Veo sample found 0.10 still detected and 0.15 not detected, so 0.15 is now the certified default. The reproducible hashes and verdicts live in data/evaluations/video-synthid-oracle.csv.

Documentation

Start with the documentation index.

Research notes and historical experiments are listed separately in the documentation index. They explain past decisions but do not define the current public API.

Contributing

Install the development environment and run the project gate:

uv sync --frozen --extra dev
bash maintain.sh

See module internals before changing a subsystem with documented invariants.

License

Apache 2.0. Copyright 2025-2026 wiltodelta.

Languages
Python 99.9%