The visible-mark path had grown three copies of one ladder sweep, four
near-identical `detect` arms, and four hand-rolled `footprint_mask` overrides;
mark knowledge sat in five hand-maintained tables across three modules; and the
flagship `all`/`batch` pipeline existed only in cli.py, written twice with
divergent behavior.
Detection is now one measurement. `_ladder_best` replaces the three sweeps,
`_scan`/`_verdict` replace the four arms, and the winning box travels to the
mask on `TextMarkDetection.match_box` instead of being swept a second time.
`detect_both` returns the strict and relaxed verdicts from one scan, which
halves the arbiter's perception cost (260 -> 130 matchTemplate calls on a 2048²
image, verdicts identical field for field). A per-mark demotion goes in the new
`_post_gate` hook, never in a `detect` override -- an override is invisible to
the single-pass path, which is how the RunningHub and Yuanbao anchor gates
briefly stopped applying.
Everything about a mark is now one registry row: product, label regime, the
platform sentence `identify` reports, the metadata signals that confirm it, and
its TC260 producer codes. `identify._VISIBLE_MARK_PLATFORM`, the signal mapping
in `api.visible_provenance`, `_PRODUCT_OF` and the pill veto are derived from
those rows.
`api.remove_all` / `api.remove_batch` are the library form of the `all` and
`batch` commands; the CLI is a wrapper that owns console text and exit codes.
Progress is a `(stage, detail)` pair of stable tokens, so the CLI keys its
wording off structure rather than parsing the library's prose back.
Two intentional behavior changes, both verified against a recorded 811-image
sample of detector verdicts, removal-mask hashes, arbiter decisions and
`identify` reports:
* A TC260 label now relaxes the vendor its `ContentProducer` names rather than
ByteDance's pair on every China-AIGC image. 333 of 811 samples move; on 185
of them the previously relaxed pair was simply the wrong vendor, and the
mark actually present never reached the relaxed gate its own
`provenance_ncc_factor` was calibrated for.
* A confident LibLibAI detection suppresses the Jimeng pill, like every other
TC260 product's mark. It was registered alongside RunningHub and Baidu, both
of which were added to the hand-written veto list, and it was not. 1 sample
moves, and it is exactly the co-firing case.
Nothing else in that record changes: detector verdicts, mask hashes and
`identify` verdicts are byte-identical, and all 200 calibration constants are
untouched.
Also: `aigc_label` and friends plus `extract_c2pa_info` are memoized on
(path, mtime_ns, size) -- size because this package rewrites in place; the
native TC260 container readers route on magic bytes instead of the file
extension, so a mislabeled AVI or FLV is no longer invisible; `identify` shares
one pixel decode between the DWT-DCT and visible stages (TrustMark keeps its own
Pillow decode, which is not substitutable); and the six `stabilize_*` video
wrappers collapse into one policy table.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
17 KiB
CLI guide
The command line interface is organized around the type of work you want to do.
remove-ai-watermarks [OPTIONS] COMMAND [ARGS]
Run remove-ai-watermarks COMMAND --help for the complete option list and
defaults. This page focuses on choosing the right command.
Command dependency map
| Command or signal | Required installation |
|---|---|
metadata and metadata-only identify |
Default package |
Visible signals in identify |
remove-ai-watermarks[visible] (pixels is the minimal runtime) |
Open DWT-DCT signals in identify |
remove-ai-watermarks[detect] |
Adobe TrustMark signals in identify |
remove-ai-watermarks[trustmark] |
visible and erase with OpenCV |
remove-ai-watermarks[visible] (pixels is the minimal runtime) |
visible or erase with MI-GAN |
remove-ai-watermarks[migan] |
visible or erase with big-LaMa |
remove-ai-watermarks[lama] |
invisible and all (needs CUDA) |
remove-ai-watermarks[qwen-zimage] |
video metadata and video identify --no-visible |
Default package |
video identify, video visible, and visible/all batch modes |
remove-ai-watermarks[video] |
video invisible and video all --invisible |
remove-ai-watermarks[video,diffusion] |
| HEIC/HEIF/AVIF pixel input | Add remove-ai-watermarks[heif] |
| Every production command and backend | remove-ai-watermarks[all] |
batch requires the same extra as its selected mode. Extras can be combined in
one installation, for example remove-ai-watermarks[visible,detect,heif].
Inspect an image
remove-ai-watermarks identify image.png
identify always inspects supported metadata. When pixel extras are installed,
it also evaluates supported visible and invisible pixel signals. When no signal
is found, it reports the origin as unknown. It does not claim the image is
clean.
Machine readable output:
remove-ai-watermarks identify image.png --json
Metadata only inspection:
remove-ai-watermarks identify image.png --no-visible
Despite the historical option name, --no-visible skips both visible and open
invisible pixel detectors. Metadata inspection still runs.
Remove known visible marks
Install remove-ai-watermarks[visible] before using visible or erase.
remove-ai-watermarks visible image.png -o clean.png
The default behavior:
- checks every registered visible mark;
- removes every detected match;
- selects the best installed fill backend;
- strips AI metadata from the output.
Use a specific mark:
remove-ai-watermarks visible image.png --mark gemini -o clean.png
Available mark names are printed by:
remove-ai-watermarks visible --help
Keep metadata:
remove-ai-watermarks visible image.png --keep-metadata -o clean.png
Use the strict visual gate without metadata or sibling corroboration:
remove-ai-watermarks visible image.png --sensitivity strict -o clean.png
When no known mark is detected, the command does not write a new output. Use
erase if you can identify the affected region yourself.
Erase a region
remove-ai-watermarks erase image.png \
--region 1640,1930,400,100 \
-o clean.png
The region format is x,y,width,height. Repeat --region to erase more than
one box:
remove-ai-watermarks erase image.png \
--region 20,20,180,60 \
--region 1640,1930,400,100 \
-o clean.png
Choose the fill backend:
remove-ai-watermarks erase image.png \
--region 1640,1930,400,100 \
--backend migan \
-o clean.png
erase accepts cv2, migan, and lama. The corresponding optional extra
must be installed for a learned backend.
Two more knobs tune the fill. --dilate N (default 3) grows every box by N
pixels before inpainting, which helps when a mark has a soft edge or a drop
shadow just outside the box you measured; it applies to every backend because it
shapes the mask. --inpaint-method telea|ns selects the classical algorithm and
only affects the cv2 backend.
Strip AI metadata
Inspect metadata:
remove-ai-watermarks metadata image.png --check
Remove AI metadata and write a new file:
remove-ai-watermarks metadata image.png --remove -o clean.png
When -o is omitted, removal overwrites the source. Standard metadata is kept
unless you pass --remove-all.
The command also supports the audio and video containers listed in supported signals. ffmpeg must be available for the non-ISOBMFF audio and video path.
Identify and clean video
Install the video pixel and timestamp runtime for visible identification, removal, and the complete pipeline:
uv tool install --force "remove-ai-watermarks[video]"
Inspect every locally supported video signal:
remove-ai-watermarks video identify input.mp4
remove-ai-watermarks video identify input.mp4 --json
remove-ai-watermarks video identify input.mp4 --no-visible
The default scans the complete clip for stable registered visible marks and
inspects supported metadata. A result with no signals is reported as unknown,
not clean, because proprietary pixel watermarks have no public local decoder.
--no-visible performs metadata-only inspection.
Use the complete locally verifiable cleaning path:
remove-ai-watermarks video all input.mp4 -o clean.mp4
It removes a stable supported visible mark when found and always strips verified AI metadata. When neither signal is found, it writes a same-container passthrough instead of returning a missing output. The source is never overwritten.
Invisible regeneration is deliberately opt-in:
remove-ai-watermarks video all input.mp4 -o clean.mp4 --invisible
That option is supported only for MP4, MOV, and M4V. It is lossy and uses the
same oracle-certified profile as video invisible.
Process all supported files in a top-level directory:
remove-ai-watermarks video batch ./videos --mode all
remove-ai-watermarks video batch ./videos --mode visible
remove-ai-watermarks video batch ./videos --mode metadata
The batch runs sequentially, preserves successful outputs when another file
fails, and exits nonzero if any item failed. Visible no-op files are copied
byte-for-byte so the output directory remains complete. --invisible is
available only with --mode all.
Strip AI metadata from video
Metadata inspection and removal are also available as an isolated operation:
remove-ai-watermarks video metadata input.mp4 --check
remove-ai-watermarks video metadata input.mp4 --remove -o clean.mp4
Supported containers are MP4, MOV, M4V, WebM, MKV, AVI, and FLV. The operation
delegates to the same verified metadata scanner and stripper as the generic
metadata command, so detection and removal stay in parity. Video and audio
streams are not transcoded. For MP4 and MOV, this includes the native TC260
AIGC key and JSON value stored in moov.udta.meta.keys/ilst. The inspector
seeks past a large mdat to find a tail moov. Removal stream-copies the
container in bounded chunks, converts supported top-level provenance boxes to
same-size free boxes, and blanks the TC260 key/value in place. Box sizes,
media offsets, and encoded stream bytes do not move; the result is atomically
published only after the complete copy succeeds.
For MKV and WebM, the inspector reads the native TC260
Segment.Tags.Tag.SimpleTag entry. Removal uses ffmpeg stream copying to
discard container tags and chapters without transcoding the streams.
AVI uses the normative LIST/INFO/AIGC chunk, while FLV uses the
script.onMetaData.AIGC AMF0 string. Their bounded readers skip media payloads,
and removal also uses ffmpeg stream copying.
When -o is omitted, the command writes <source>_clean with the same
extension. It never overwrites the source, and it rejects an output with a
different container extension.
Visible video labels and invisible video watermarks are not handled by this command.
Remove video SynthID
uv tool install --force "remove-ai-watermarks[video,diffusion]"
remove-ai-watermarks video invisible input.mp4 -o clean.mp4
The command supports MP4, MOV, and M4V. It samples the complete sequence at the configured frame rate, resizes frames to the configured long side, regenerates them through a VAE, and applies one deterministic latent-noise field to every frame. Reusing one spatial field avoids the unnecessary flicker caused by independent per-frame noise. Frames are regenerated in bounded batches and streamed directly to ffmpeg, which encodes the result, copies audio, and drops source metadata.
The default noise_std=0.15 profile is oracle-certified. The project has no
local video SynthID decoder, so an optional per-file recheck is still useful
for unusually important files or after provider changes. In a new Gemini chat,
upload the original first, invoke the built-in verifier with @synthid, and ask:
For the video attached to this message, was it created or edited by Google AI? Use the built-in SynthID content verification result.
The source must be positive. Then upload the processed result in a separate new chat and repeat the same built-in check. Only a source-positive, output-negative pair is a fresh per-file verification. Do not ask an adversarial follow-up that tells the chat model to ignore the verifier and reason about raw pixels: that is ordinary Gemini reasoning, not a second oracle check.
The default output is <source>_clean in the same container. The
source is never overwritten. Use --noise-std, --long-side, --fps,
--batch-size, --seed, and --device to control the regeneration. The
default noise level is 0.15. It cleared both carriers in the 2026-07-29
short-clip calibration and the complete public eight-second Veo carrier in the
2026-07-31 full-clip check; 0.10 remained detected on that complete clip. This
calibration certifies the shipped operating point; the paired check above is an
optional runtime audit, not a separate result state.
Remove a supported visible video mark
remove-ai-watermarks video visible input.mp4 -o clean.mp4
remove-ai-watermarks video visible veo.mp4 --mark veo -o veo_clean.mp4
remove-ai-watermarks video visible seedance.mp4 --mark seedance -o seedance_clean.mp4
remove-ai-watermarks video visible dola.mp4 --mark dola -o dola_clean.mp4
remove-ai-watermarks video visible hailuo.mp4 --mark hailuo -o hailuo_clean.mp4
remove-ai-watermarks video visible kling.mp4 --mark kling -o kling_clean.mp4
The command supports the moving Sora mascot and wordmark, two Veo
corner variants, the Seedance boxed AI label, the Dola AI text label, the
composite MINIMAX | hailuo AI label, and the bottom-right Kling label. Sora
searches the whole frame at multiple scales. The other detectors search bounded
lower-frame regions with separate synthetic silhouettes. Kling additionally
requires its bright low-saturation label near the frame edge. Every mark
requires a spatially recurring candidate across adjacent frames. Fixed marks
must also remain anchored instead of drifting with a scene object. Matching
provider provenance may relax the visual score only for registered
provenance-aware marks; metadata alone never creates a detection.
--mark auto is the default. It evaluates all providers in one decode pass and
selects the first stable match in specificity order: Sora, Veo, Seedance, Dola,
Hailuo, then Kling. Their confidence scores are independently calibrated and
are not compared across providers. Pass an explicit --mark to scan only that
provider.
The video stream is transcoded and the complete original audio stream is
copied without truncating an audio tail that extends beyond the final video
frame. The encoder probes the source stream and preserves supported 8-bit
chroma sampling, color range/matrix/transfer/primaries tags, and MP4/MOV track
timescale. For a variable-frame-rate source, decoded PTS are carried through a
timestamped in-memory NUT bridge so the output retains the source frame
intervals instead of flattening them to a constant rate. A non-zero source
start PTS and the copied audio start offset are preserved as well.
Supported input and output containers are MP4, MOV, M4V, WebM, MKV, AVI, and
FLV; the output extension must match the input. The default cv2 backend is
fast but can smear structured backgrounds. Select --backend migan or
--backend lama for a learned fill, or --backend auto to choose the best
installed backend.
--temporal-consistency is enabled by default. It motion-aligns the preceding
accepted fill, requires overlapping removal masks and matching source context,
and blends only the safely covered pixels. Scene cuts, disjoint moving marks,
or a poor motion match keep the independent current-frame fill. Use
--no-temporal-consistency for an exact frame-local baseline.
The pixel path is intentionally limited to SDR 8-bit video. A high-bit-depth, PQ, or HLG source is rejected before ffmpeg starts, preserving any existing output instead of silently downconverting it through OpenCV's 8-bit boundary. On CPU, MI-GAN is the practical learned tier. LaMa remains an explicit offline quality option because full-sequence inference is too slow and memory-heavy for an online worker.
AI metadata is stripped from the encoded output by default. Use
--keep-metadata to retain mapped container metadata. When no temporally stable
mark is found, the command writes no output and exits with the no-visible-mark
status. The final path is replaced atomically only after ffmpeg completes, so a
failed encode does not overwrite an existing result.
Remove invisible watermarks
Install the removal dependencies first. Both profiles are CUDA-only and both run the DiffSynth Z-Image face stage, so this is the extra either one needs:
uv tool install --force "remove-ai-watermarks[qwen-zimage]"
Then run:
remove-ai-watermarks invisible image.png -o clean.png
The command normally skips regeneration when no supported local signal is
detected. Use --force when you know the image should be processed:
remove-ai-watermarks invisible image.png -o clean.png --force
Choose a pipeline
| Pipeline | When to use it |
|---|---|
qwen-zimage |
Default. Qwen-Image-2512 global pass plus a SAM-masked Z-Image face stage |
sdxl-zimage |
The same recipe and face stage on an SDXL global pass, at a higher denoise |
Both are CUDA-only. There is no CPU or MPS profile for invisible-watermark
removal. The former controlnet, sdxl, qwen and default profiles were removed
rather than kept as a CPU path: none of them matched this recipe's face preservation,
so offering them implied a quality the library no longer delivers. Passing a retired
name is rejected at parse time rather than remapped. Visible-mark removal and every
identify path still run anywhere.
Example:
remove-ai-watermarks invisible image.png -o clean.png \
--pipeline qwen-zimage --force
There is no --model, --steps, --guidance-scale or --device option, and the
deprecated --auto is gone. Each profile pins its model stack, its per-stage
schedule, CFG 1.0 and CUDA, so every one of those flags existed only to be refused
several layers down. They are not parsed at all now, which fails at the point the
user can act on rather than after a model load.
Work with limited memory
Lower CUDA memory pressure:
remove-ai-watermarks invisible image.png -o clean.png \
--cpu-offload --force
Keep large images at native resolution while processing them in overlapping tiles:
remove-ai-watermarks invisible image.png -o clean.png \
--tile --max-resolution 0 --force
Or set a resolution cap:
remove-ai-watermarks invisible image.png -o clean.png \
--max-resolution 2048 --force
Tiling avoids the explicit downscale but each tile is regenerated separately. It is a memory strategy, not a guarantee of better quality.
Run the full pipeline
The all command and the all installation extra are separate concepts. The
command runs every applicable stage. Installing remove-ai-watermarks[all]
makes every production backend available; a smaller installation such as
remove-ai-watermarks[visible,qwen-zimage] can also run the command with fewer
optional backends.
remove-ai-watermarks all image.png -o clean.png
The command runs:
- visible mark removal;
- invisible watermark removal when available and applicable;
- AI metadata stripping.
The visible options and diffusion options are also available on all.
If invisible removal is required but the qwen-zimage extra is unavailable, all still
writes the result of the visible and metadata stages, prints a prominent
warning, and exits with code 1. This prevents a partial result from being
reported as complete.
Process a directory
remove-ai-watermarks batch ./images --mode visible
Modes:
visible;invisible;metadata;all.
Set an output directory:
remove-ai-watermarks batch ./images \
--mode all \
--output-dir ./clean
The invisible and full modes accept the same main diffusion controls as their
single image counterparts. Run batch --help for the authoritative option
list.
Exit behavior
The CLI uses nonzero exit codes for meaningful incomplete outcomes, including no detected target on commands that would otherwise regenerate or create a misleading unchanged result, processing errors, and a required invisible step that could not run.
Scripts should check the process exit code and the output path. The detailed per-command contract is maintained in module internals.