T09 · Insecure Skill Coding Practices
- Location
scripts/modelpilot_report.py:12- Finding
Incomplete Benchmark Data Can Produce a False Replacement-Ready Decision
- Content
View full analysis
Vulnerability Details
File Location:
scripts/modelpilot_report.py, lines 12–26
Vulnerability Type: Insufficient integrity and completeness validation
Risk Level: MediumVulnerable Code
python def decision_for_model(records: list[dict[str, Any]], rounds_required: int = 2) -> tuple[str, str]: rounds = {record.get("round") for record in records} if len(rounds) < rounds_required: return "candidate_only", "Only one benchmark round is complete." failures = [record for record in records if not record.get("success")] format_failures = [record for record in records if not record.get("format_pass")] think_leaks = [record for record in records if record.get("think_leak")] if failures: return "not_recommended", f"{len(failures)} prompt runs failed." if think_leaks: return "not_recommended", f"{len(think_leaks)} outputs show possible thinking leakage." if format_failures: return "observe", f"{len(format_failures)} outputs failed the expected format check." return "replace_ready", "Two rounds passed mechanical checks. Human semantic review is still required."Technical Analysis
The replacement decision treats the presence of any two distinct
roundvalues as proof that two complete benchmark rounds occurred. It does not validate:- That the round identifiers are the expected values, such as rounds 1 and 2.
- That every configured prompt was executed in each round.
- That each round contains the same prompt identifiers.
- That records are unique rather than duplicated.
- That record coverage agrees with
prompt_countandrounds_requested. - That model, prompt, round, and result fields have valid types and values.
Consequently, as few as two successful records with different round labels can cause the function to return
replace_ready. This contradicts the Skill's documented requirement to run the same fixed prompt set in two complete, independent rounds.Because ` ...[truncated 1486 chars]
- Remediation
View remediation
Remediation Suggestions
- Define and enforce a strict schema for benchmark input, including field types and allowed values.
- Require exact expected round identifiers, such as every integer from 1 through
rounds_required. - Determine the expected prompt-ID set and verify that every model has exactly one record for every prompt in every required round.
- Reject duplicate
(model, round, prompt_id)records. - Validate record coverage against trusted
prompt_countandrounds_requestedvalues. - Reject missing, null, non-Boolean, or incorrectly typed result fields instead of interpreting them through Python truthiness.
- Fail closed with an
invalid_resultsorcandidate_onlydecision whenever input completeness cannot be established. - Add tests covering truncated rounds, duplicate records, unexpected round values, inconsistent prompt sets, missing fields, and manipulated metadata.
A hardened decision routine should only return
replace_readyafter proving that all required prompts completed successfully in every required round.
