T09 · Insecure Skill Coding Practices
- Location
scripts/verify_intelligence.py:69- Finding
Citation Verification Can Be Bypassed with Unrelated or Untrusted Citation Tokens
- Content
View full analysis
Vulnerability Details
File Location:
scripts/verify_intelligence.py, lines 69-72 and 124-169
Vulnerability Type: Insufficient semantic validation / integrity-control bypass
Risk Level: MediumComplete Code Snippet
python OFFICIAL_ANCHORS = ( "http://", "https://", )python def citation_window(line: str, next_line: str | None) -> str: return line + "\n" + (next_line or "") def has_any(window: str, anchors: tuple[str, ...]) -> bool: return any(a in window for a in anchors) def classify(window: str) -> str: """Return 'mcp' | 'official' | 'media' | 'forbidden' | 'none'.""" if has_any(window, FORBIDDEN_ANCHORS): return "forbidden" if has_any(window, MCP_TOOL_ANCHORS): return "mcp" # Tushare cited without the aigroup-market-mcp prefix still traces to # the same installed tool — classify as mcp, not media. if "Tushare" in window or "tushare" in window: return "mcp" if has_any(window, OFFICIAL_ANCHORS): return "official" if has_any(window, MEDIA_ANCHORS): return "media" return "none" def scan(text: str, strict_mcp: bool = False) -> list[tuple[int, str, str, str, str]]: """Return list of (line_no, number, unit, snippet, reason) failing the gate.""" lines = text.splitlines() failures: list[tuple[int, str, str, str, str]] = [] for i, line in enumerate(lines): next_line = lines[i + 1] if i + 1 < len(lines) else None window = citation_window(line, next_line) for m in HARD_NUMBER.finditer(line): verdict = classify(window) if verdict == "none": failures.append( (i + 1, m.group(1), m.group(2), line.strip()[:120], "no citation") ) elif verdict == "forbidden": failures.append( (i ...[truncated 3321 chars]- Remediation
View remediation
Remediation Suggestions
- Remove generic
http://andhttps://strings fromOFFICIAL_ANCHORS. - Parse URLs with a standard URL parser and allowlist exact authoritative hostnames, including controlled subdomain handling. Reject user-info tricks, suffix confusion, redirects to untrusted domains, malformed hosts, and non-HTTPS links where HTTPS is available.
- Require each hard number to have a citation in the same sentence, structured footnote, or explicit claim-to-source identifier rather than accepting any token on the next line.
- Validate footnote references bidirectionally: the claim must reference an existing footnote, and that footnote must contain an allowed source associated with the claim.
- In strict mode, accept only validated MCP provenance records or URLs whose normalized host is on the official-source allowlist.
- Do not classify a line solely through broad substring matching. Parse citation syntax and preserve the source type, URL, publication title, filing date, and claim association as structured data.
- Add negative tests covering arbitrary URLs, unrelated next-line links, invalid URLs, deceptive subdomains, multiple unrelated numbers on one line, and citation tokens embedded in ordinary prose.
- Add a regression test asserting that the illustrative attack input fails in both default and strict modes.
- Remove generic
