Back to skill

Security audit

gaokao-english-vocabulary

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent Gaokao vocabulary webpage generator, but its template should be fixed before using untrusted vocabulary data or publishing the page.

Installers should understand that this skill is mainly a local template and data generator. It appears purpose-aligned and does not seek sensitive access, but generated pages should only use trusted vocabulary data until count validation and HTML escaping are added, especially before hosting the page on a real site.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
references/template_structure.md:335
Finding

Stored HTML and JavaScript Injection Through an Unvalidated Phrase Count

Content
View full analysis
⚠️ ' + escHtml(w.t) + '' : ''; var phoneticHtml = w.p ? ' [' + escHtml(w.p) + ']' : ''; var meaningHtml = ''; var masteryField = w.v || w.i || 'normal'; // ... return '
' + '
' + '
' + escHtml(w.w) + '' + escHtml(w.s) + '' + phoneticHtml + '
' + '
' + meaningHtml + '
' + '
' + '' + getMasteryLabel(masteryField) + '' + '考查 ...[truncated 3340 chars]
Remediation
View remediation
10000: raise ValueError("The count is outside the permitted range") return value ``` Use the validated value for both words and phrases: ```python count = validate_count(entry.get('c', 0)) ``` 2. **Encode every value placed into HTML.** Even after validation, apply output encoding as a defense-in-depth measure: ```javascript 'Exam count: ' + escHtml(String(maxC)) + '' ``` 3. **Prefer safe DOM APIs over HTML-string concatenation.** Create elements with `document.createElement()` and assign untrusted values through `textContent`. This prevents the browser from interpreting values as markup. 4. **Validate the complete input schema.** Enforce expected types and bounds for `w`, `s`, `m`, `c`, `l`, `v`, `i`, `p`, `e`, and `t`. Reject invalid records instead of silently applying defaults. 5. **Add regression tests.** Include phrase counts containing HTML tags, event handlers, strings, booleans, negative numbers, floating-point values, and excessively large integers. Verify that invalid records are rejected and that no input value can create executable DOM markup. 6. **Consider a restrictive Content Security Policy.** If the generated page is served over HTTP, use a policy that disallows inline event handlers and unauthorized network destinations. This is defense in depth and does not replace validation and safe rendering. ]]>
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (4)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The supplied code is related to Gaokao English vocabulary grading, so it partially aligns in domain and data structure. However, its actual function is narrowly a backend-style data transformation script: it converts JSONL vocabulary records into a JavaScript data file and reports counts/distributions. It does not create webpages, styling, interactivity, search/filter functionality, or a complete study tool as declared. Thus the description materially overstates and misrepresents what this code chunk actually does.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The embedded HTML template sets lang="zh-CN" and presents interface labels, status messages, and placeholders entirely in Chinese. Because the file does not offer localization options or explain that the template is intentionally region-specific, it appears to enforce a specific language/locale without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The skill description, trigger keywords, and later UI examples are presented as Chinese-specific by default, which implies the generated webpage and interaction text will be in Chinese. The file does not mention any user opt-in or option to generate the tool in another language, which can conflict with a language-choice policy.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.