Close representation-change campaign: both candidates falsified

Pre-registered gates (AI-test >=1772, fresh OI <=50, Kodak 0, FLUX >=288,
EvalGEN >=95) rejected DINOv2-L ridge head (1427/1847, 162/3000, 1/24, 29/300,
5/100) and the paired-reconstruction FFT tower (14/1847, 12/3000, 0/24, 0/300,
0/100; contract learned at 0.949 pairwise separation but no transfer to
production generators). Campaign verdict: general classifier beyond local
scale as pursued; product stays on documented remainder.

pre-commit: 1) maintain.sh - exit 1, known uv-secure lightning advisory with no upstream fix; core checks green minutes ago (ruff, format, pyright, 1665 tests); 2) /simplify - docs-only single pass, no findings; 3) docs sync - verdict mirrored in data/research/REPRESENTATION-CAMPAIGN.md, no stale refs; 4) CLAUDE.md - no changes
This commit is contained in:
Victor Kuznetsov
2026-08-27 11:23:25 -07:00
parent 85b18af804
commit 7aff09e163
+21
View File
@@ -777,6 +777,27 @@ attribution the honest rule stays argmax until an external held-out meta cell
exists. The margin, not the head, was always the limiter on the small class.
Artifacts: `meta-margin-2026-08-27/report.json`.
## Representation-change campaign closed, 2026-08-27
A pre-registered two-candidate campaign (plan in `data/research/REPRESENTATION-CAMPAIGN.md`,
gates declared before any run: AI-test at least 1,772/1,905, fresh Open
Images at most 50/3,000, Kodak 0/24, FLUX at least 288/300, EvalGEN at least
95/100, no tuning) tested the two remaining representation families.
Candidate 1, a DINOv2-L ridge head (224 px CLS embeddings, nested-CV penalty),
failed all five gates simultaneously: AI-test 1,427/1,847, fresh 162/3,000,
Kodak 1/24, FLUX 29/300, EvalGEN 5/100 — the frozen DINOv2 space separates AI
from photos far worse than fine-tuned CLIP-L everywhere. Candidate 2, a
patch-level FFT tower on the DDA paired-reconstruction contract (1,024 fresh
sd-vae-ft-mse pairs, real 0.549 versus reconstruction 0.457 medians, 0.949
pairwise separation), learned its contract perfectly and still collapsed on
transfer: AI-test 14/1,847, FLUX 0/300, EvalGEN 0/100, because production
generators do not carry this VAE's reconstruction signature. The declared
falsification criterion triggered: neither family is the missing lever, and
the general metadata-free AI-versus-human classifier at the product contract
is beyond local scale as pursued. The product keeps the documented-remainder
contract; Model 1 stays the accepted research operating point with its
measured bounds. Reopening requires a materially new lever.
A taxonomy continuation then changed the training mix itself instead of the
veto: two arms re-ran the expanded quarter-hard recipe (ordinary AI replay,
EvalGEN pair positives, ordinary photos) with one or two of the eight negative