mirror of
https://github.com/CyberSecurityUP/NeuroSploit.git
synced 2026-09-29 12:31:43 +02:00
Every finding in the Arena engagement carried votes reading 1/1 while the run was configured for a 3-model vote. The panel was candidates.take(n), so a run with one configured model produced a "multi-model adversarial validation" that was one model agreeing with itself — the single most important thing the engagement revealed about the harness. Filling the panel by repeating that model would not fix it: its errors are correlated with themselves, and three confident repetitions of one mistake are indistinguishable from a consensus. Anthropic checking Anthropic is not independent; a second vendor is. build_panel() now takes at most one model per provider from the configured candidates, then fills any shortfall from backends this machine can actually reach (an installed subscription CLI, or a provider whose API key is in the environment). When nothing else is available the panel stays small and the yes/total the caller prints tells the truth about it, rather than being padded to look like a quorum. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>