mirror of
https://github.com/wiltodelta/remove-ai-watermarks.git
synced 2026-08-06 22:18:36 +02:00
Register the Kling 可灵AI 3.0 visible text mark; park Yuanbao and cat-logo (measured)
Kling (USCC cohort 91110108335469089C, n=30): kling_engine.py, gate 0.35 (clean p99 0.304 / max 0.320), strict-only, unimodal 0.12/short on the shared ladder, fitted locate box, no rival margin (crossfire 1/400 doubao below gate, 0 jimeng, 0 clean), parity 9/9 detect->fill->re-detect. Suppresses the jimeng pill like doubao/qwen. identify gains visible_kling. Yuanbao: measured negative -- the two-line italic block does not separate from clean corners on either front-end at any render/box/font setting; the fitted recipe stays in render_vendor_silhouettes.py MARK_OPTS. cat-logo: cohort has only 2 unique carriers, parked on evidence; the draw_catlogo silhouette already separates (0.50 vs clean max 0.333), so registration is a gate pick once more uniques arrive. vendor_mark_calibrate: --fit-geometry takes locate-box overrides (two-line marks were clipped by the inherited box) and the aspect sweep reaches 0.62.
This commit is contained in:
@@ -745,6 +745,49 @@ priority order:
|
||||
config -- samsung is `bl` but Latin-script and width-based). 星绘/百度 are NOT in the
|
||||
corpus in labelable quantity -- verified, do not hunt them again.
|
||||
|
||||
**STATUS 2026-07-21 (same day): 可灵 REGISTERED, 元宝 measured and PARKED.**
|
||||
* **可灵 (`kling_engine.py`)** -- "可灵AI 3.0" bottom-right, strict-only, gate
|
||||
0.35, no rival margin, shared 3-rung ladder (the mark is UNIMODAL at 0.12 of the
|
||||
short side). Cohort-vs-clean (286 guarded clean frames): clean p99 0.304 / max
|
||||
0.320; 9 of ~19 eyeballed visible marks fire = ~47% recall of visible marks, all
|
||||
9 true (precision 9/9). The misses are the faint "Omni"-suffix release, the
|
||||
latin "KlingAI 3.0" release and the version-less "可灵AI" (0.17-0.25, inside the
|
||||
clean arm's top tail -- unreachable). Crossfire: 1/400 doubao (a 豆包 frame
|
||||
INSIDE the kling cohort, still below gate), 0/298 jimeng, 0/286 clean. Parity
|
||||
9/9 detect->cv2 fill->re-detect clean. A confident kling detection suppresses
|
||||
the jimeng pill exactly like doubao's/qwen's does. The bottom-LEFT `AI生成`
|
||||
pill variant was NOT seen in this cohort's contact sheet at registration time
|
||||
and stays unhandled.
|
||||
* **元宝 -- MEASURED NEGATIVE, parked.** The mark is a TWO-LINE, italic-slanted
|
||||
block (元宝 over AI生成), ~5% of the short side. After fitting the render
|
||||
against real tophat responses (left-align, tight gap, stroke dilation, shear
|
||||
-0.75 -- which lifted marked frames to 0.65-0.70, at the real-vs-real ceiling
|
||||
~0.6), the CLEAN arm rose in lockstep (clean p99 0.643 vs cohort p50 0.472):
|
||||
the slanted two-line template correlates with generic corner texture at the
|
||||
same rate it gains on the mark, on BOTH the tophat and binary front-ends, in
|
||||
wide and tight boxes, in three CJK fonts. Every separation metric measured was
|
||||
negative. This is the 2026-07-18 千问-style wall, except it survived the
|
||||
geometry fix: the mark is small + slanted + half-shared-tail, and no synthetic
|
||||
template separates it on this front-end. The residual levers are a structural
|
||||
two-line verification stage or a learned patch classifier -- both outside the
|
||||
cheap playbook. `yuanbao_alpha.png` + its `MARK_OPTS` recipe stay in
|
||||
`render_vendor_silhouettes.py` as the documented starting point if that lever
|
||||
is ever built. Fit-trap found en route (now guarded): `_fit_one`'s tiny-gw NCC
|
||||
inflation -- a sub-30px template scores spuriously high on smooth tophat
|
||||
responses, so the auto-fit picked a degenerate 0.026 width fraction; the
|
||||
numbers above come from a gw-floored re-fit.
|
||||
* **cat-logo -- probe READY, parked on evidence.** The cohort (USCC
|
||||
91110108562144110X) is 19 frames but only **2 unique carriers** (byte-unique) --
|
||||
the xinghui rule (nothing registered off ~one frame) applies. The mark is an
|
||||
outline cat-head + bold "AI生成", bottom-right, ~0.25 of the width, very bold.
|
||||
A drawn synthetic silhouette (`draw_catlogo` in `render_vendor_silhouettes.py`;
|
||||
a solid filled head scored 0.35, the outline form 0.50 -- iterated against a
|
||||
real tophat response) separates: mark 0.50 vs a diverse clean arm max 0.333
|
||||
(n=29 probe). Registration is a gate pick (~0.42) the moment more unique
|
||||
carriers arrive; recall across diverse cat-logo generations is unmeasurable at
|
||||
n=2. Fit-trap that applied here too: the mark is bigger than Doubao's box
|
||||
(0.25 of width), so the inherited box clipped it exactly like qwen's.
|
||||
|
||||
Do NOT restart the sweeps to "check". Their artifacts are on disk and listed under
|
||||
"Completed full runs" below; re-running costs hours and answers nothing new. The fast way
|
||||
to confirm the whole surface still works after a change is
|
||||
|
||||
@@ -6,8 +6,15 @@ tracked asset (see the repo CLAUDE.md data-safety rule). Seeing real samples to
|
||||
the glyphs, weight and layout is fine; the committed template stays synthetic.
|
||||
|
||||
Covered here:
|
||||
qwen "千问AI生成" -- Alibaba Tongyi Qianwen, bottom-right, 3-lobed logo + text
|
||||
xinghui "星绘AI生成" -- ByteDance 星绘, bottom-right, 4-point sparkle + text
|
||||
qwen "千问AI生成" -- Alibaba Tongyi Qianwen, bottom-right, 3-lobed logo + text
|
||||
xinghui "星绘AI生成" -- ByteDance 星绘, bottom-right, 4-point sparkle + text
|
||||
yuanbao "元宝\nAI生成" -- Tencent Yuanbao, bottom-right, two-line italic block
|
||||
(MEASURED NEGATIVE 2026-07-21, parked: the slanted two-line template does
|
||||
not separate the cohort from clean corners on either front-end; the recipe
|
||||
+ MARK_OPTS stay as the starting point if a structural/learned lever is
|
||||
built -- full record in docs/verification-plan.md)
|
||||
kling "可灵AI 3.0" -- Kuaishou Kling, bottom-right, spiral logo + text
|
||||
(REGISTERED 2026-07-21, kling_engine.py)
|
||||
|
||||
The leading LOGO is deliberately NOT rendered. It is the part that varies most between
|
||||
releases and is hardest to reproduce synthetically, while the CJK run is stable and is
|
||||
@@ -91,6 +98,7 @@ from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import numpy as np
|
||||
from PIL import Image, ImageDraw, ImageFont
|
||||
@@ -103,36 +111,137 @@ _FONT = "/System/Library/Fonts/STHeiti Medium.ttc"
|
||||
MARKS = {
|
||||
"qwen_alpha.png": "千问AI生成",
|
||||
"xinghui_alpha.png": "星绘AI生成",
|
||||
# Yuanbao's stamp is a TWO-LINE block (元宝 over AI生成), left-aligned, tightly
|
||||
# stacked and ITALIC-SLANTED (measured on the 2026-07-21 cohort sheet + real tophat
|
||||
# responses); a rare one-line variant exists but the stacked block is dominant.
|
||||
"yuanbao_alpha.png": "元宝\nAI生成",
|
||||
# Kling (可灵) stamps a thin light-gray one-line "可灵AI 3.0" bottom-right (an
|
||||
# "Omni" suffix variant and a latin "KlingAI 3.0" variant also exist; the CJK
|
||||
# run without the suffix is the common core). The leading spiral logo is NOT
|
||||
# rendered (logos vary; the text run discriminates).
|
||||
"kling_alpha.png": "可灵AI 3.0",
|
||||
# The "cat-logo" cohort (USCC 91110108562144110X) stamps an outline cat-head +
|
||||
# bold "AI生成", bottom-right. PARKED 2026-07-21: the cohort is 19 copies of
|
||||
# only 2 unique carriers -- nothing to calibrate recall against (the xinghui
|
||||
# rule). The probe is ready: this silhouette scores 0.50 on the mark vs 0.333
|
||||
# max on a diverse clean arm, so registration is a gate pick (0.42) the moment
|
||||
# more unique carriers arrive.
|
||||
"catlogo_alpha.png": "CATLOGO", # sentinel: drawn by draw_catlogo(), not font-rendered
|
||||
}
|
||||
|
||||
# Per-mark post-processing for the multi-line / slanted stamps (see render()).
|
||||
MARK_OPTS: dict[str, dict[str, Any]] = {
|
||||
# Fitted against real tophat responses on the Yuanbao cohort (2026-07-21): a
|
||||
# right-aligned, gapped, unslanted render plateaued at ~0.34 NCC; left-align +
|
||||
# tight gap + stroke dilation + shear -0.75 reaches 0.65-0.70 on the same frames,
|
||||
# at/above the real-vs-real ceiling (~0.6).
|
||||
"yuanbao_alpha.png": {"gap_frac": 0.05, "dilate": 2, "shear": -0.75},
|
||||
}
|
||||
|
||||
|
||||
def render(text: str, width: int = 335) -> np.ndarray:
|
||||
def render(text: str, width: int = 335, opts: dict[str, Any] | None = None) -> np.ndarray:
|
||||
"""Binary glyph silhouette (255 = glyph), sized to the doubao asset's convention.
|
||||
|
||||
Matching doubao's 335px asset width keeps the `alpha_*_frac` numbers transferable,
|
||||
since these marks are the same house style and scale.
|
||||
since these marks are the same house style and scale. A "\n" in ``text`` renders a
|
||||
multi-line block: lines drawn left-aligned at one shared font size with a tight
|
||||
gap, then optional stroke dilation and an italic shear (see MARK_OPTS).
|
||||
"""
|
||||
opts = opts or {}
|
||||
gap_frac = float(opts.get("gap_frac", 0.15))
|
||||
dilate = int(opts.get("dilate", 0))
|
||||
shear_k = float(opts.get("shear", 0.0))
|
||||
probe = Image.new("L", (10, 10))
|
||||
d0 = ImageDraw.Draw(probe)
|
||||
lines = text.split("\n")
|
||||
size = 8
|
||||
while size < 200: # grow until the run fills the target width
|
||||
while size < 200: # grow until the LONGEST line fills the target width
|
||||
f = ImageFont.truetype(_FONT, size)
|
||||
if d0.textbbox((0, 0), text, font=f)[2] >= width * 0.98:
|
||||
if max(d0.textbbox((0, 0), ln, font=f)[2] for ln in lines) >= width * 0.98:
|
||||
break
|
||||
size += 1
|
||||
font = ImageFont.truetype(_FONT, size)
|
||||
boxes = [d0.textbbox((0, 0), ln, font=font) for ln in lines]
|
||||
line_h = max(bb[3] - bb[1] for bb in boxes)
|
||||
gap = max(1, int(line_h * gap_frac))
|
||||
w = max(bb[2] - bb[0] for bb in boxes)
|
||||
h = line_h * len(lines) + gap * (len(lines) - 1)
|
||||
pad = max(2, int(line_h * 0.12))
|
||||
im = Image.new("L", (w + 2 * pad, h + 2 * pad), 0)
|
||||
draw = ImageDraw.Draw(im)
|
||||
y = pad
|
||||
for ln, bb in zip(lines, boxes, strict=True):
|
||||
draw.text((pad - bb[0], y - bb[1]), ln, font=font, fill=255)
|
||||
y += line_h + gap
|
||||
sil = np.array(im)
|
||||
if dilate or shear_k:
|
||||
import cv2
|
||||
|
||||
if dilate:
|
||||
sil = cv2.dilate(sil, np.ones((dilate, dilate), np.uint8))
|
||||
if shear_k:
|
||||
hh, ww = sil.shape
|
||||
sil = cv2.warpAffine(sil, np.float32([[1, shear_k, 0], [0, 1, 0]]), (ww + int(abs(shear_k) * hh), hh))
|
||||
return sil
|
||||
|
||||
|
||||
def draw_catlogo(width: int = 335) -> np.ndarray:
|
||||
"""The cat-logo mark: an outline cat-head (integrated pointy ears, two dot eyes)
|
||||
+ a bold "AI生成" run, drawn synthetically from the measured layout (cat ~1.08x
|
||||
the glyph height, stroke ~9%, gap ~35%). Proportions were iterated against a real
|
||||
tophat response (2026-07-21): a solid filled head scored 0.35, this outline form
|
||||
0.50 -- the parked probe, see MARKS."""
|
||||
probe = Image.new("L", (10, 10))
|
||||
d0 = ImageDraw.Draw(probe)
|
||||
text = "AI生成"
|
||||
size = 8
|
||||
while size < 200:
|
||||
f = ImageFont.truetype(_FONT, size)
|
||||
if d0.textbbox((0, 0), text, font=f)[2] >= width * 0.60:
|
||||
break
|
||||
size += 1
|
||||
font = ImageFont.truetype(_FONT, size)
|
||||
bb = d0.textbbox((0, 0), text, font=font)
|
||||
w, h = bb[2] - bb[0], bb[3] - bb[1]
|
||||
tw, th = bb[2] - bb[0], bb[3] - bb[1]
|
||||
cs = int(th * 1.08)
|
||||
stroke = max(2, int(th * 0.09))
|
||||
gap = int(th * 0.35)
|
||||
|
||||
def head(s: int) -> Image.Image:
|
||||
im = Image.new("L", (s, s), 0)
|
||||
d = ImageDraw.Draw(im)
|
||||
f = float(s)
|
||||
pts = [
|
||||
(0.12 * f, 0.95 * f),
|
||||
(0.10 * f, 0.45 * f),
|
||||
(0.12 * f, 0.30 * f),
|
||||
(0.20 * f, 0.05 * f), # left ear tip
|
||||
(0.40 * f, 0.24 * f), # left ear valley
|
||||
(0.60 * f, 0.24 * f), # right ear valley
|
||||
(0.80 * f, 0.05 * f), # right ear tip
|
||||
(0.88 * f, 0.30 * f),
|
||||
(0.90 * f, 0.45 * f),
|
||||
(0.88 * f, 0.95 * f),
|
||||
]
|
||||
d.line([*pts, pts[0]], fill=255, width=stroke, joint="curve")
|
||||
r = max(1.5, stroke * 0.7)
|
||||
d.ellipse([0.35 * f - r, 0.60 * f - r, 0.35 * f + r, 0.60 * f + r], fill=255)
|
||||
d.ellipse([0.65 * f - r, 0.60 * f - r, 0.65 * f + r, 0.60 * f + r], fill=255)
|
||||
return im
|
||||
|
||||
w = cs + gap + tw
|
||||
h = max(th, cs)
|
||||
pad = max(2, int(h * 0.12))
|
||||
im = Image.new("L", (w + 2 * pad, h + 2 * pad), 0)
|
||||
ImageDraw.Draw(im).text((pad - bb[0], pad - bb[1]), text, font=font, fill=255)
|
||||
im.paste(head(cs), (pad, pad + (h - cs) // 2))
|
||||
ImageDraw.Draw(im).text((pad + cs + gap - bb[0], pad + (h - th) // 2 - bb[1]), text, font=font, fill=255)
|
||||
return np.array(im)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
try:
|
||||
for name, text in MARKS.items():
|
||||
sil = render(text)
|
||||
sil = draw_catlogo() if text == "CATLOGO" else render(text, opts=MARK_OPTS.get(name))
|
||||
Image.fromarray(sil).save(_ASSETS / name)
|
||||
print(f"wrote {_ASSETS / name} ({sil.shape[1]}x{sil.shape[0]}) text={text!r}")
|
||||
except OSError as e:
|
||||
|
||||
@@ -229,7 +229,7 @@ _FIT_SCALES = tuple(round(0.4 * (1.03**i), 4) for i in range(80)) # 0.40 .. ~4.
|
||||
_SHIPPED_LADDER = (0.8, 1.0, 1.25)
|
||||
|
||||
|
||||
def _fit_one(args: tuple[str, str]) -> dict[str, Any] | None:
|
||||
def _fit_one(args: tuple[str, str, dict[str, Any]]) -> dict[str, Any] | None:
|
||||
"""Best match over the WIDE ladder, reported as a mark width in pixels.
|
||||
|
||||
Also measures the template ASPECT at the winning width: the mark's true height is
|
||||
@@ -238,14 +238,14 @@ def _fit_one(args: tuple[str, str]) -> dict[str, Any] | None:
|
||||
(that inflated the clean p99 from 0.30 to 0.58 on the 2026-07-18 attempt) and not
|
||||
inherited from doubao.
|
||||
"""
|
||||
path_str, asset = args
|
||||
path_str, asset, overrides = args
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
from remove_ai_watermarks._text_mark_engine import TextMarkEngine
|
||||
from remove_ai_watermarks.image_io import imread
|
||||
|
||||
cfg = build_config(asset, "fit", "width")
|
||||
cfg = build_config(asset, "fit", "width", overrides)
|
||||
eng = TextMarkEngine(cfg)
|
||||
img = imread(path_str)
|
||||
if img is None:
|
||||
@@ -269,11 +269,15 @@ def _fit_one(args: tuple[str, str]) -> dict[str, Any] | None:
|
||||
if v > best:
|
||||
best, best_gw, best_tl = v, gw, (int(tl[0]), int(tl[1]))
|
||||
# Aspect fit at the winning width: sweep gh/gw and keep the argmax. Range covers
|
||||
# everything between samsung's 0.12 and jimeng's 0.29 house styles, plus slack.
|
||||
# everything between samsung's 0.12 and jimeng's 0.29 house styles, plus the
|
||||
# two-line stacked marks (Yuanbao ~0.45), plus slack.
|
||||
best_aspect = 0.0
|
||||
if best_gw > 0:
|
||||
best_gh_score = -1.0
|
||||
for ratio in np.arange(0.12, 0.42, 0.01):
|
||||
# Upper bound raised 0.42 -> 0.62 for two-line marks (Yuanbao's stacked block
|
||||
# has silhouette aspect ~0.45; the old range's 0.12 floor was its own trap --
|
||||
# the fit "won" by squashing the template to a one-line strip).
|
||||
for ratio in np.arange(0.12, 0.62, 0.01):
|
||||
gh = max(4, int(best_gw * float(ratio)))
|
||||
if gh >= resp.shape[0]:
|
||||
continue
|
||||
@@ -299,17 +303,27 @@ def _fit_one(args: tuple[str, str]) -> dict[str, Any] | None:
|
||||
}
|
||||
|
||||
|
||||
def fit_geometry(paths: list[str], asset: str, workers: int, floor: float = 0.50, paths_name: str = "cohort") -> None:
|
||||
def fit_geometry(
|
||||
paths: list[str],
|
||||
asset: str,
|
||||
workers: int,
|
||||
floor: float = 0.50,
|
||||
paths_name: str = "cohort",
|
||||
overrides: dict[str, Any] | None = None,
|
||||
) -> None:
|
||||
"""Which basis and fraction does this vendor's mark actually scale with?
|
||||
|
||||
Only frames matching above ``floor`` are used: below it the winning size is the
|
||||
ladder's best fit to background texture, not a measurement of the mark.
|
||||
``overrides`` adjusts the LOCATE box for the fit (a two-line mark like Yuanbao's
|
||||
is taller than Doubao's inherited box -- scoring it in the inherited box clips
|
||||
the template to zero overlap).
|
||||
"""
|
||||
import numpy as np
|
||||
|
||||
rows: list[dict[str, Any]] = []
|
||||
with ProcessPoolExecutor(max_workers=workers) as ex:
|
||||
for f in as_completed([ex.submit(_fit_one, (p, asset)) for p in paths]):
|
||||
for f in as_completed([ex.submit(_fit_one, (p, asset, overrides or {})) for p in paths]):
|
||||
try:
|
||||
r = f.result()
|
||||
except Exception: # noqa: S112 -- one bad file must not kill the fit
|
||||
@@ -565,7 +579,7 @@ def main() -> None:
|
||||
pos_paths, neg_paths = load_sets(a.cohort)
|
||||
if a.fit_geometry:
|
||||
print(f"cohort {a.cohort}: {len(pos_paths)} candidates")
|
||||
fit_geometry(pos_paths, a.asset, a.workers, paths_name=name)
|
||||
fit_geometry(pos_paths, a.asset, a.workers, paths_name=name, overrides=overrides)
|
||||
return
|
||||
|
||||
if a.crossfire:
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 2.8 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 3.3 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 6.1 KiB |
@@ -447,6 +447,7 @@ _VISIBLE_MARK_PLATFORM = {
|
||||
"doubao": "ByteDance Doubao (visible 豆包AI生成 mark detected)",
|
||||
"jimeng": "ByteDance Jimeng / Dreamina (visible 即梦AI mark detected)",
|
||||
"qwen": "Alibaba Tongyi Qianwen (visible 千问AI生成 mark detected)",
|
||||
"kling": "Kuaishou Kling (visible 可灵AI 3.0 mark detected)",
|
||||
"samsung": "Samsung Galaxy AI (visible 'Contenuti generati dall'AI' mark detected)",
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,149 @@
|
||||
"""Kling (可灵, Kuaishou) visible watermark detector/localizer.
|
||||
|
||||
Kling stamps its generations with a thin, light-gray "可灵AI 3.0" text strip in the
|
||||
bottom-right corner, preceded by the vendor's spiral logo (not part of the detection
|
||||
silhouette -- logos vary between releases, the text run is what discriminates).
|
||||
Known variants: an "Omni" suffix release, a latin "KlingAI 3.0" release, and a
|
||||
version-less "可灵AI" -- the silhouette targets the common "可灵AI 3.0" core, so the
|
||||
suffix variants are only caught when the core run is bold enough (measured below).
|
||||
|
||||
Detection matches the bundled glyph silhouette against the corner; removal is the
|
||||
shared **localize -> fill** (the glyph-bbox :meth:`footprint_mask` feeds
|
||||
``region_eraser``), NOT reverse-alpha. This module supplies only Kling's tuned
|
||||
:class:`TextMarkConfig` (``assets/kling_alpha.png`` -- a font-rendered synthetic
|
||||
silhouette from ``scripts/render_vendor_silhouettes.py``, never cut from an
|
||||
upload). It also feeds ``identify`` as the medium-confidence ``visible_kling``
|
||||
signal via the registry.
|
||||
|
||||
EVERY tuned number below was measured on the vendor cohort (30 TC260 carriers whose
|
||||
producer USCC 91110108335469089C names the entity, 2026-07-21; harness
|
||||
``scripts/vendor_mark_calibrate.py``), NOT inherited from Doubao:
|
||||
|
||||
* The mark scales with the SHORT side at ~0.12 of it (mark_w/short measured
|
||||
0.118-0.122 across portrait AND landscape carriers -- unimodal, so the shipped
|
||||
3-rung ladder covers it) and sits ~0.03 off the right/bottom edges; the locate
|
||||
box fractions below are fitted from the measured absolute mark rects.
|
||||
* ``alpha_height_frac`` comes from the silhouette aspect (0.239) at the fitted
|
||||
width, matching the aspect the fit converged on (0.25).
|
||||
* Gate 0.35, one step above the clean arm's max: on the cohort-vs-clean run
|
||||
(cohort-contamination-guarded, 286 hand-labelled clean frames) the clean arm
|
||||
scored p99 0.304 / max 0.320, and every cohort frame >= 0.35 carries a visible
|
||||
可灵AI 3.0 mark (9 of ~19 eyeballed visible marks fire = ~47% recall of visible
|
||||
marks; the misses are the faint "Omni"-suffix release, the latin "KlingAI"
|
||||
release and the version-less "可灵AI", which score 0.17-0.25 and cannot be
|
||||
reached without engulfing the clean arm).
|
||||
* STRICT ONLY (``provenance_ncc_factor`` 1.0): the sub-gate band holds real Kling
|
||||
variants AND the clean arm's top (clean p90 0.220 vs variant marks at 0.17-0.25
|
||||
-- they overlap), so a provenance-relaxed arm cannot separate them. No
|
||||
provenance relaxation exists for this mark.
|
||||
* No rival margin: at the shipped gate the template fires on 1 of 400
|
||||
Doubao-marked frames (0.2%, a 豆包 frame sitting INSIDE the Kling cohort, still
|
||||
below the gate), 0 of 298 Jimeng-marked frames and 0 of 286 hand-labelled clean
|
||||
frames, and a 0.10 rival margin costs zero genuine Kling detections -- so it is
|
||||
simply unnecessary (same conclusion shape as Qwen).
|
||||
"""
|
||||
# The module-level _alpha_template / _glyph_silhouette / _template_match_score below
|
||||
# are thin test-facing shims (imported by tests/), so pyright's src-only pass sees them
|
||||
# as unused; the use is cross-module.
|
||||
# pyright: reportUnusedFunction=false
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import TYPE_CHECKING, Any
|
||||
|
||||
from remove_ai_watermarks import _text_mark_engine
|
||||
from remove_ai_watermarks._text_mark_engine import TextMarkConfig, TextMarkDetection, TextMarkEngine
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from pathlib import Path
|
||||
|
||||
from numpy.typing import NDArray
|
||||
|
||||
# Locate geometry as a fraction of the image SHORT side (measured basis -- see
|
||||
# scale_base). The box is fitted to the measured mark rects: the mark's right
|
||||
# margin is ~0.034 of the short side and its bottom margin ~0.027; width/height
|
||||
# cover the mark plus NCC slack.
|
||||
WM_WIDTH_FRAC = 0.19
|
||||
WM_HEIGHT_FRAC = 0.05
|
||||
MARGIN_RIGHT_FRAC = 0.03
|
||||
MARGIN_BOTTOM_FRAC = 0.023
|
||||
|
||||
# Glyph appearance: a light, low-saturation gray rendered brighter than the local
|
||||
# background (white top-hat), same overlay class as Doubao -- inherited, and
|
||||
# harmless because the tophat front-end turns these gates into weights.
|
||||
MAX_SATURATION = 55
|
||||
LOGO_MIN_LUMA = 150
|
||||
TOPHAT_DELTA = 12
|
||||
|
||||
DETECT_MIN_COVERAGE = 0.04 # unused by the tophat front-end (kept for config parity)
|
||||
# Calibrated 2026-07-21 on the vendor cohort vs 286 hand-labelled clean frames
|
||||
# (cohort-contamination-guarded): clean p99 0.304 / max 0.320, and every cohort
|
||||
# frame scoring >= 0.35 carries a visible 可灵AI 3.0 mark. 0.35 was picked over
|
||||
# 0.33 (also zero clean fires) for margin against unseen clean content at a cost
|
||||
# of zero measured cohort detections.
|
||||
DETECT_NCC_THRESHOLD = 0.35
|
||||
|
||||
# Detection-silhouette geometry (fraction of the short side), fitted on the
|
||||
# cohort: the mark's width (0.12, unimodal) and the silhouette aspect (0.239).
|
||||
_ALPHA_WIDTH_FRAC = 0.12
|
||||
_ALPHA_HEIGHT_FRAC = 0.0287
|
||||
|
||||
_CONFIG = TextMarkConfig(
|
||||
name="Kling",
|
||||
asset_name="kling_alpha.png",
|
||||
corner="br",
|
||||
margin_floor=4,
|
||||
width_frac=WM_WIDTH_FRAC,
|
||||
height_frac=WM_HEIGHT_FRAC,
|
||||
margin_x_frac=MARGIN_RIGHT_FRAC,
|
||||
margin_bottom_frac=MARGIN_BOTTOM_FRAC,
|
||||
max_saturation=MAX_SATURATION,
|
||||
logo_min_luma=LOGO_MIN_LUMA,
|
||||
tophat_delta=TOPHAT_DELTA,
|
||||
morph_open_size=5,
|
||||
detect_min_coverage=DETECT_MIN_COVERAGE,
|
||||
detect_ncc_threshold=DETECT_NCC_THRESHOLD,
|
||||
detect_frontend="tophat",
|
||||
scale_basis="short", # measured: mark_w/short 0.118-0.122 across orientations
|
||||
alpha_width_frac=_ALPHA_WIDTH_FRAC,
|
||||
alpha_height_frac=_ALPHA_HEIGHT_FRAC,
|
||||
min_gw=8,
|
||||
# STRICT ONLY: the sub-gate band (real Kling variants at 0.17-0.25) overlaps
|
||||
# the clean arm's top (p90 0.220), so provenance relaxation is disabled
|
||||
# outright (factor 1.0 = never relaxed).
|
||||
provenance_ncc_factor=1.0,
|
||||
)
|
||||
|
||||
KlingDetection = TextMarkDetection
|
||||
|
||||
|
||||
def _alpha_template() -> NDArray[Any] | None:
|
||||
"""The bundled Kling alpha template (float [0,1]), or None."""
|
||||
return _text_mark_engine.load_alpha_template(_CONFIG.asset_name)
|
||||
|
||||
|
||||
def _glyph_silhouette() -> NDArray[Any] | None:
|
||||
"""Binary "可灵AI 3.0" silhouette (255 = glyph) from the alpha map, or None."""
|
||||
return _text_mark_engine.glyph_silhouette(_CONFIG.asset_name)
|
||||
|
||||
|
||||
def _template_match_score(box_mask: NDArray[Any], scale_base: int) -> float:
|
||||
"""TM_CCOEFF_NORMED of the Kling glyph silhouette against ``box_mask``."""
|
||||
return _text_mark_engine.template_match_score(box_mask, scale_base, _CONFIG)
|
||||
|
||||
|
||||
class KlingEngine(TextMarkEngine):
|
||||
"""Detect/localize the visible Kling "可灵AI 3.0" watermark (locate -> mask; mask feeds the fill)."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__(_CONFIG)
|
||||
|
||||
|
||||
def load_image_bgr(path: str | Path) -> NDArray[Any]:
|
||||
"""Read an image as BGR ndarray (helper for scripts/tests)."""
|
||||
from remove_ai_watermarks import image_io
|
||||
|
||||
img = image_io.imread(path)
|
||||
if img is None:
|
||||
raise FileNotFoundError(f"Failed to read image: {path}")
|
||||
return img
|
||||
@@ -21,6 +21,7 @@ Entries:
|
||||
- ``doubao`` -- ByteDance Doubao "豆包AI生成" text strip, bottom-right.
|
||||
- ``jimeng`` -- ByteDance Jimeng / Dreamina "★ 即梦AI" wordmark, bottom-right.
|
||||
- ``qwen`` -- Alibaba Tongyi Qianwen "千问AI生成" text strip, bottom-right.
|
||||
- ``kling`` -- Kuaishou Kling "可灵AI 3.0" text strip, bottom-right.
|
||||
- ``samsung`` -- Samsung Galaxy AI "Contenuti generati dall'AI" strip, bottom-left.
|
||||
- ``jimeng_pill`` -- Jimeng-basic "AI生成" pill, top-left (capture-less).
|
||||
"""
|
||||
@@ -85,6 +86,7 @@ _PRODUCT_OF: dict[str, str] = {
|
||||
"jimeng": "jimeng",
|
||||
"jimeng_pill": "jimeng", # same product as the Jimeng wordmark
|
||||
"qwen": "qwen",
|
||||
"kling": "kling",
|
||||
"samsung": "samsung",
|
||||
}
|
||||
|
||||
@@ -359,6 +361,10 @@ def _engine(key: str) -> Any:
|
||||
from remove_ai_watermarks.qwen_engine import QwenEngine
|
||||
|
||||
_engines[key] = QwenEngine()
|
||||
elif key == "kling":
|
||||
from remove_ai_watermarks.kling_engine import KlingEngine
|
||||
|
||||
_engines[key] = KlingEngine()
|
||||
elif key == "samsung":
|
||||
from remove_ai_watermarks.samsung_engine import SamsungEngine
|
||||
|
||||
@@ -509,6 +515,7 @@ _REGISTRY: tuple[KnownMark, ...] = (
|
||||
_text_mark("doubao", "Doubao 豆包AI生成 text", "bottom-right"),
|
||||
_text_mark("jimeng", "Jimeng 即梦AI wordmark", "bottom-right"),
|
||||
_text_mark("qwen", "Qwen 千问AI生成 text", "bottom-right"),
|
||||
_text_mark("kling", "Kling 可灵AI 3.0 text", "bottom-right"),
|
||||
_text_mark("samsung", "Samsung Galaxy AI text", "bottom-left"),
|
||||
KnownMark("jimeng_pill", "Jimeng AI生成 pill", "top-left", True, _pill_detect, _pill_mask, _pill_features),
|
||||
)
|
||||
@@ -590,7 +597,7 @@ def _keep_pill(keys: set[str], *, provenance: frozenset[str], footprint_flat: bo
|
||||
Doubao detection; a Qwen image likewise (another vendor's bottom-right mark naming
|
||||
its own product), so a confident Qwen detection suppresses the pill the same way.
|
||||
No confirmation at all -> never remove (blocks false fires on non-Jimeng content)."""
|
||||
if "doubao" in keys or "qwen" in keys:
|
||||
if "doubao" in keys or "qwen" in keys or "kling" in keys:
|
||||
return False
|
||||
if "jimeng" in keys:
|
||||
return True
|
||||
|
||||
@@ -0,0 +1,149 @@
|
||||
"""Tests for the Kling (可灵AI 3.0) visible-watermark engine (localize -> fill).
|
||||
|
||||
Every tuned constant in ``kling_engine`` was measured on the 30-frame vendor
|
||||
cohort (2026-07-21, ``scripts/vendor_mark_calibrate.py``); these tests pin the
|
||||
load-bearing ones so a later "cleanup" cannot silently re-inherit Doubao's
|
||||
geometry or relax the measured strict-only gate.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
import pytest
|
||||
|
||||
from remove_ai_watermarks import watermark_registry as registry
|
||||
from remove_ai_watermarks.kling_engine import (
|
||||
_ALPHA_HEIGHT_FRAC,
|
||||
_ALPHA_WIDTH_FRAC,
|
||||
KlingEngine,
|
||||
_alpha_template,
|
||||
_glyph_silhouette,
|
||||
)
|
||||
|
||||
_MARK_FRAC = 0.12 # measured mark width, fraction of the short side (unimodal)
|
||||
_MARGIN = 0.03 # measured right/bottom margin of the real mark
|
||||
|
||||
|
||||
def _compose(w: int, h: int, mode: float = _MARK_FRAC, bg: float = 100.0):
|
||||
"""Composite the Kling silhouette at the measured size onto a flat bg."""
|
||||
img = np.full((h, w, 3), bg, np.float32)
|
||||
at = _alpha_template()
|
||||
short = min(w, h)
|
||||
gw = int(mode * short)
|
||||
gh = max(4, int(mode * (_ALPHA_HEIGHT_FRAC / _ALPHA_WIDTH_FRAC) * short))
|
||||
margin = int(_MARGIN * short)
|
||||
ax = w - margin - gw
|
||||
ay = h - margin - gh
|
||||
amap = np.zeros((h, w), np.float32)
|
||||
amap[ay : ay + gh, ax : ax + gw] = cv2.resize(at, (gw, gh))
|
||||
a3 = amap[:, :, None]
|
||||
wm = (a3 * 255.0 + (1 - a3) * img).clip(0, 255).astype(np.uint8)
|
||||
return wm, amap > 0.2
|
||||
|
||||
|
||||
class TestLocate:
|
||||
def test_box_anchored_bottom_right(self):
|
||||
eng = KlingEngine()
|
||||
img = np.zeros((2048, 2048, 3), np.uint8)
|
||||
loc = eng.locate(img)
|
||||
assert 2048 - (loc.x + loc.w) == pytest.approx(2048 * 0.03, rel=0.15)
|
||||
assert 2048 - (loc.y + loc.h) == pytest.approx(2048 * 0.023, rel=0.15)
|
||||
|
||||
def test_box_scales_with_short_side_not_width(self):
|
||||
# scale_basis="short" (measured: mark_w/short 0.118-0.122 across orientations).
|
||||
eng = KlingEngine()
|
||||
landscape = eng.locate(np.zeros((640, 1280, 3), np.uint8))
|
||||
wider = eng.locate(np.zeros((640, 2560, 3), np.uint8))
|
||||
assert wider.w == landscape.w # same short side -> same box
|
||||
bigger = eng.locate(np.zeros((1280, 1920, 3), np.uint8)) # 2x the short side
|
||||
assert bigger.w == pytest.approx(landscape.w * 2, rel=0.05)
|
||||
|
||||
|
||||
class TestConfig:
|
||||
def test_shared_ladder_default(self):
|
||||
# The mark is unimodal at 0.12 of the short side, so Kling keeps the shared
|
||||
# 3-rung ladder (Qwen's per-mark ladder is the measured exception, not a norm).
|
||||
assert KlingEngine().config.ladder == (0.8, 1.0, 1.25)
|
||||
|
||||
def test_strict_only_no_provenance_relaxation(self):
|
||||
# The sub-gate band (real Kling variants at 0.17-0.25) overlaps the clean
|
||||
# arm's top (p90 0.220), so a relaxed arm cannot separate: factor pinned 1.0.
|
||||
assert KlingEngine().config.provenance_ncc_factor == 1.0
|
||||
|
||||
def test_gate_above_clean_arm_max(self):
|
||||
# Clean arm scored p99 0.304 / max 0.320 on 286 hand-labelled frames; the
|
||||
# gate must sit above that with margin.
|
||||
assert KlingEngine().config.detect_ncc_threshold > 0.32
|
||||
|
||||
def test_registry_row(self):
|
||||
mark = registry.get_mark("kling")
|
||||
assert mark.location == "bottom-right"
|
||||
assert "可灵AI" in mark.label
|
||||
assert mark.in_auto
|
||||
|
||||
def test_confident_kling_detection_suppresses_the_jimeng_pill(self):
|
||||
# A Kling image is TC260 too but is not Jimeng-basic: like Doubao and Qwen,
|
||||
# a confident Kling detection must veto the pill (``_keep_pill``).
|
||||
from remove_ai_watermarks.watermark_registry import _keep_pill
|
||||
|
||||
assert not _keep_pill({"kling"}, provenance=frozenset({"jimeng"}), footprint_flat=1.0)
|
||||
|
||||
|
||||
class TestDetect:
|
||||
def test_clean_gradient_not_detected(self):
|
||||
eng = KlingEngine()
|
||||
ramp = np.tile(np.linspace(0, 255, 1024, dtype=np.uint8), (1024, 1))
|
||||
img = cv2.cvtColor(ramp, cv2.COLOR_GRAY2BGR)
|
||||
assert not eng.detect(img).detected
|
||||
|
||||
def test_solid_blob_corner_not_detected(self):
|
||||
eng = KlingEngine()
|
||||
img = np.zeros((1024, 1024, 3), np.uint8)
|
||||
x, y, bw, bh = eng.locate(img).bbox
|
||||
img[y + bh // 4 : y + bh * 3 // 4, x : x + bw // 2] = 200
|
||||
assert not eng.detect(img).detected
|
||||
|
||||
def test_silhouette_loads(self):
|
||||
sil = _glyph_silhouette()
|
||||
assert sil is not None
|
||||
assert set(np.unique(sil)).issubset({0, 255})
|
||||
|
||||
def test_composed_mark_detected(self):
|
||||
# The registration's core claim: a mark at the measured size scores over the
|
||||
# gate. The floor is deliberately far above the gate: the synthetic mark is
|
||||
# clean, so it scores high when the geometry is right.
|
||||
wm, _ = _compose(853, 640)
|
||||
det = KlingEngine().detect(wm)
|
||||
assert det.detected
|
||||
assert det.confidence >= 0.80
|
||||
|
||||
def test_small_image_guarded(self):
|
||||
wm, _ = _compose(853, 640)
|
||||
eng = KlingEngine()
|
||||
assert eng.detect(wm).detected
|
||||
assert not eng.detect(cv2.resize(wm, (150, 112))).detected
|
||||
|
||||
|
||||
class TestFootprintMaskAndRemoval:
|
||||
def test_removes_composed_mark(self):
|
||||
wm, mark = _compose(853, 640)
|
||||
assert float(np.abs(wm.astype(np.float32)[mark] - 100.0).mean()) > 15 # mark visible
|
||||
assert KlingEngine().detect(wm).detected
|
||||
out, region = registry.get_mark("kling").remove(wm, backend="cv2")
|
||||
assert region is not None
|
||||
assert not KlingEngine().detect(out).detected
|
||||
h, w = wm.shape[:2]
|
||||
assert np.array_equal(out[: h // 2, : w // 2], wm[: h // 2, : w // 2]) # far region exact
|
||||
|
||||
def test_footprint_mask_in_bottom_right(self):
|
||||
wm, _ = _compose(853, 640)
|
||||
mask = KlingEngine().footprint_mask(wm)
|
||||
assert mask is not None
|
||||
ys, xs = np.where(mask > 0)
|
||||
assert ys.mean() > wm.shape[0] / 2
|
||||
assert xs.mean() > wm.shape[1] / 2
|
||||
|
||||
def test_clean_frame_produces_no_mask(self):
|
||||
clean = cv2.GaussianBlur(np.full((640, 853, 3), 120, np.uint8), (5, 5), 0)
|
||||
assert KlingEngine().footprint_mask(clean, force=False) is None
|
||||
@@ -14,7 +14,7 @@ DOUBAO_SAMPLE = Path(__file__).resolve().parents[1] / "data" / "samples" / "doub
|
||||
|
||||
class TestCatalog:
|
||||
def test_keys(self):
|
||||
assert reg.mark_keys() == ["gemini", "doubao", "jimeng", "qwen", "samsung", "jimeng_pill"]
|
||||
assert reg.mark_keys() == ["gemini", "doubao", "jimeng", "qwen", "kling", "samsung", "jimeng_pill"]
|
||||
|
||||
def test_all_in_auto(self):
|
||||
assert all(m.in_auto for m in reg.known_marks())
|
||||
@@ -43,7 +43,7 @@ class TestScan:
|
||||
def test_detect_marks_scans_all(self):
|
||||
img = np.zeros((256, 256, 3), np.uint8)
|
||||
keys = {d.key for d in reg.detect_marks(img)}
|
||||
assert keys == {"gemini", "doubao", "jimeng", "qwen", "samsung", "jimeng_pill"}
|
||||
assert keys == {"gemini", "doubao", "jimeng", "qwen", "kling", "samsung", "jimeng_pill"}
|
||||
|
||||
def test_blank_image_no_auto_mark(self):
|
||||
dets = reg.detect_marks(np.zeros((256, 256, 3), np.uint8), include_explicit=False)
|
||||
|
||||
Reference in New Issue
Block a user