Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detection, option-letter selection bias (chi2), few-shot curve noise, LLM-judge position/verbosity/rubric-echo bias and hidden-instruction payload detection, paired McNemar + Wilson + deterministic bootstrap for score comparisons, WORKED mitigations (permutation majority ensemble, blind content normalization), documented 0-100 severity formula, hash-chained per-target history ledger with trend deltas. Findings cite a STATIC 17-id exploit catalogue with explicit computable flags — invisible exploit classes are declared, never fabricated. 100% stdlib python3. NO network, NO telemetry. Defense/auditing only.

Install

openclaw skills install @orionshaowswmw/benchmark-robustness-auditor