Back to skill

Security audit

Gaokao Volunteer

Security checks for vulnerabilities and agentic risk

Overview

This Gaokao planning skill mostly matches its stated purpose, but its generated HTML reports can include unescaped imported data, so reports built from untrusted sources could display or run unwanted content.

Review this skill before installing if you plan to open or share generated HTML reports. Use official admissions data where possible, avoid feeding it untrusted JSON or scraped content, and treat the report as advisory because it stores visible score/rank details and currently lacks HTML sanitization.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/plan_generator.py:65
Finding

Stored HTML and Script Injection in Generated Reports

Content
View full analysis
= 0.8 else ("prob-mid" if prob >= 0.5 else "prob-low") risk_icon = {"冲": "🚀", "稳": "✅", "保": "🛡️"}.get(risk, "") hist_text = str(hist_ranks[-1]) if hist_ranks else "-" rows.append(f""" {rank} {uni} {group} {majs} {hist_text} {prob_text} {risk_icon} {risk} {note} """) ``` Profile fields are also inserted without HTML escaping: ```python report = report.replace("{{TITLE}}", f"2026年高考志愿填报方案") report = report.replace( "{{SUBTITLE}}", f"{province_name} | {subject} | {score}分 | 位次{rank}" ) report = report.replace("{{GENERATED_TIME}}", now) ``` Warnings are converted into HTML list items without escaping: ```python warnings_html = "" if warnings: for w in warnings: warnings_html += f'
  • {w}
  • \n' else: warnings_html = '
  • 未检测到明显风险,请结合个人情况复核
  • ' report = report.replace("{{WARNINGS}}", warnings_html) ``` ### Technical Analysis `plan_generator.py` treats profile data, admissions records, major names, notes, risk values, and warning messages as trusted HTML. These values ...[truncated 2730 chars]
    Remediation
    View remediation
    str: return escape(str(value), quote=True) ``` Apply this function to all profile, admissions, plan, and warning fields before inserting them into the template: ```python uni = html_text(item.get("university_name", "-")) group = html_text(item.get("major_group_name", "-")) majs = html_text(", ".join(map(str, item.get("majors_in_group", []))) or "-") note = html_text(item.get("note", "")) ``` 2. Do not insert untrusted values into attribute names or class attributes. Map `risk_level` to a fixed allowlist: ```python allowed_risks = {"冲", "稳", "保"} risk = item.get("risk_level", "") if risk not in allowed_risks: risk = "-" ``` 3. Escape profile and warning values independently: ```python subtitle = " | ".join([ html_text(province_name), html_text(subject), html_text(score), html_text(rank), ]) warnings_html = "\n".join( f"
  • {html_text(w)}
  • " for w in warnings ) ``` 4. Introduce strict JSON-schema validation before report generation. Enforce expected field types, maximum lengths, numeric ranges, and enumerated values. 5. Prefer a template engine with automatic HTML escaping enabled rather than manual string replacement. 6. Add a restrictive Content Security Policy to the generated document, adjusted to the report's actual requirements: ```html ``` 7. Add regression tests containing payloads in every externally controlled field. Verify that characters such as `<`, `>`, `"`, `'`, and `&` appear encoded in the resulting HTML and cannot create new elements or attributes. ]]>
    Vulnerability Patterns
    • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
    • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
    • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
    • Behavioral ASTexec() Call, eval() Call, Dynamic Import
    • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
    Findings (14)

    Direct Prompt Extraction

    High
    Category
    System Prompt Leakage
    Confidence
    85% confidence
    Finding

    Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

    Content

    Scanner excerpt · scripts/plan_generator.py (reported line 26)May include surrounding context.

    python
    def get_province_rule(rules: dict, province: str) -> dict:
        """获取省份填报规则"""
        return rules.get(province.lower(), rules.get(province, {}))
    
    
    def build_risk_distribution_html(strategy: dict) -> str:
    

    Missing User Warnings

    Medium
    Category
    Not specified by scanner
    Confidence
    88% confidence
    Finding

    The README advertises automatic web searching and HTML report generation without warning that the skill may contact external sources or create local output files. This reduces transparency and informed consent, and can expose users to privacy, trust, or local-content risks if queried data is sent externally or generated HTML is opened without understanding its origin.

    Content

    No source excerpt is available for this finding.

    Vague Triggers

    Medium
    Category
    Not specified by scanner
    Confidence
    92% confidence
    Finding

    The trigger phrases are broad enough to match common educational queries such as score lookup, university recommendations, or general gaokao advice, which can cause the skill to activate unexpectedly. In a skill that performs web searches and generates structured outputs, overbroad activation increases the chance of unintended data access, misleading automation, or user confusion about what tool is responding.

    Content

    No source excerpt is available for this finding.

    Undeclared Tool Scope

    Medium
    Category
    MCP Least Privilege
    Confidence
    90% confidence
    Finding

    The skill instructs the agent to execute local scripts, generate HTML output, and access reference files, which implies code execution plus file read/write behavior, yet the manifest declares no explicit tool scope or permission boundary. Without an allowlist, a host agent may grant broader-than-necessary capabilities, increasing the blast radius if the skill is misused or if prompt injection later influences file paths or outputs.

    Content

    No source excerpt is available for this finding.

    Vague Triggers

    Medium
    Category
    Not specified by scanner
    Confidence
    90% confidence
    Finding

    The trigger list contains broad phrases such as 'gaokao', '院校推荐', and '专业推荐', which can cause the skill to activate in conversations that are only loosely related to college application planning. Unintended activation can expose unnecessary tool use, increase privacy risk around student data, and route users into a workflow that performs searches or file operations they did not intend.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    96% confidence
    Finding

    This code file contains user-facing natural-language text that forces a specific language/locale for documentation and command help. Under the policy, language constraints should either offer user choice or be clearly documented as region-specific and justified; that justification is not present here.

    Content

    No source excerpt is available for this finding.

    Dynamic import via __import__()

    Medium
    Category
    Dangerous Code Execution
    Confidence
    75% confidence
    Finding

    Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.

    Content

    Scanner excerpt · scripts/risk_classifier.py (reported line 180)May include surrounding context.

    python
    Path(args.output).parent.mkdir(parents=True, exist_ok=True)
            with open(args.output, "w", encoding="utf-8") as f:
                json.dump(result, f, ensure_ascii=False, indent=2)
            print(f"[risk_classifier] 方案已保存到 {args.output}", file=__import__("sys").stderr)
    
        print(json.dumps(result, ensure_ascii=False, indent=2))
    

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    93% confidence
    Finding

    The module docstring, usage text, and argument help are presented only in Chinese, which imposes a specific language on users. Under the policy, forcing a language without user opt-in or a documented justification is a natural-language policy violation.

    Content

    No source excerpt is available for this finding.

    Vague Triggers

    Low
    Category
    Not specified by scanner
    Confidence
    82% confidence
    Finding

    The "When to Use" section lists positive examples but does not clarify exclusion conditions, such as non-Gaokao college admissions, general career counseling, or overseas applications. Without negative examples or scope limits, the activation condition remains somewhat ambiguous.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Low
    Category
    Not specified by scanner
    Confidence
    88% confidence
    Finding

    This JSON manifest/data file contains natural-language descriptions and labels in Chinese, but it does not indicate that Chinese is an intentional locale-specific setting or offer any language choice. Under the language/locale policy rule, forcing a specific language without user opt-in can be a policy concern even in config/data files.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Low
    Category
    Not specified by scanner
    Confidence
    86% confidence
    Finding

    This JSON uses only Chinese natural-language descriptions and notes, including the top-level description and advisory text, with no indication that users may choose another language or locale. Under the policy rule, forcing a specific language without user opt-in is a natural-language locale violation unless the constraint is explicitly documented and justified.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Low
    Category
    Not specified by scanner
    Confidence
    84% confidence
    Finding

    This JSON manifest/data file uses Chinese-only descriptive text such as the title and note, with no indication that the skill is region-specific or that users can opt into this locale. Under the natural-language policy rule, forcing a specific language without user choice can be a locale-policy issue.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Low
    Category
    Not specified by scanner
    Confidence
    95% confidence
    Finding

    The default HTML template hard-codes lang="zh-CN", and the script’s user-facing strings throughout the file are Chinese-only. This creates a language/locale constraint without any visible opt-in or explanation that the skill is intentionally region-specific.

    Content

    No source excerpt is available for this finding.

    Intent-Code Divergence

    Low
    Category
    Not specified by scanner
    Confidence
    98% confidence
    Finding

    The inline comment says higher admission probabilities are placed later and that safer choices should be toward the back, but the code uses reverse=True, which sorts higher probabilities first. This is a direct contradiction between documentation and implementation, even though it is low impact.

    Content

    No source excerpt is available for this finding.

    Static analysis

    No suspicious patterns detected.