Back to skill

Security audit

Fly Flight

Security checks for vulnerabilities and agentic risk

Overview

This travel lookup skill matches its stated purpose, but it needs review because it executes external flight-page data as local JavaScript and weakens HTTPS protection for train queries.

Review carefully before installing. Use only in an isolated environment with minimal filesystem and environment access, avoid exposing HTTP mode beyond localhost, and do not rely on the train results over untrusted networks until TLS verification is fixed. The flight parser should be changed to parse provider state without eval before this is treated as safe for normal agent use.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (3)

T03 · Remote Payload Retrieval and Execution

Error
Location
scripts/extract_tongcheng_state.js:6
Finding

Remote Provider Content Is Executed as JavaScript

Content
View full analysis
str: request = Request(url, headers={"User-Agent": USER_AGENT}) with urlopen(request, timeout=timeout) as response: return response.read().decode("utf-8", "ignore") def extract_state_from_html(html: str) -> dict[str, Any]: script_path = ROOT / "scripts" / "extract_tongcheng_state.js" completed = subprocess.run( ["node", str(script_path)], input=html, capture_output=True, text=True, check=False, ) if completed.returncode != 0: error = completed.stderr.strip() or completed.stdout.strip() or "unknown extractor error" raise RuntimeError(f"公开页面解析失败: {error}") return json.loads(completed.stdout) ``` `scripts/extract_tongcheng_state.js:6-18`: ```javascript const match = html.match(/window\.__NUXT__=(.*?);<\/script>/s) || html.match(/window\.__NUXT__=(.*?)<\/script>/s); if (!match) { console.error("Could not find window.__NUXT__ payload in HTML."); process.exit(1); } let nuxt; try { nuxt = eval(match[1]); } catch (error) { console.error(`Failed to evaluate Nuxt payload: ${error.message}`); process.exit(1); } ``` ### Technical Analysis The flight provider downloads an HTML document from a mutable external service and sends the complete response to a Node.js subprocess. The extractor uses a regular expression to capture the value assigned to `window.__NUXT__`, then executes that captured text with `eval()`. The captured value is not constrained to JSON. It may contain arbitrary JavaScript expressions, immediately invoked functions, calls t ...[truncated 2083 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/providers/train_public_service.py:24
Finding

TLS Certificate Verification Is Disabled for 12306 Requests

Content
View full analysis
str: request = Request(url, headers={"User-Agent": USER_AGENT}) with urlopen(request, timeout=timeout, context=DEFAULT_SSL_CONTEXT) as response: return response.read().decode("utf-8", "ignore") def fetch_json(url: str, params: dict[str, str], timeout: int) -> dict[str, Any]: query = urlencode(params) request = Request(f"{url}?{query}", headers={"User-Agent": USER_AGENT}) with urlopen(request, timeout=timeout, context=DEFAULT_SSL_CONTEXT) as response: return json.loads(response.read().decode("utf-8-sig", "ignore")) def build_session_opener() -> Any: cookie_jar = http.cookiejar.CookieJar() return build_opener( HTTPCookieProcessor(cookie_jar), HTTPSHandler(context=DEFAULT_SSL_CONTEXT), ) ``` ### Technical Analysis The train provider explicitly creates an unverified SSL context with `ssl._create_unverified_context()` and uses it for direct requests and the cookie-enabled session opener. This disables normal certificate-chain and hostname validation. Consequently, the client accepts certificates that are expired, self-signed, issued for a different hostname, or controlled entirely by an attacker. Although the URLs use HTTPS, the application does not authenticate the remote endpoint, defeating a core security guarantee of TLS. The affected traffic includes station metadata, ticket availability, fare information, and session initialization. A forged station response can also influence how user-provided station names are resolved. ### Attack Path 1. An attacker obtains a network interception position, such as control of a hostile Wi-Fi network, ...[truncated 1212 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/transport_service.py:76
Finding

HTTP Search Parameters Allow Arbitrary Local JSON File Reads

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (34)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description claims one unified transport skill supporting both domestic flights and high-speed rail with routing by mode. However, this code file is explicitly described as a 'flight-only wrapper' and its search subcommand only invokes the flight provider. All search parameters are airline/airport/direct-flight oriented, with no train-related inputs or mode selection. Although the serve path uses a shared TransportHandler and prints a sample URL containing mode=flight, that still indicates a flight-oriented wrapper rather than demonstrating the full unified transport capability described. Therefore the declared description overstates the capability of the supplied code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a single unified transport skill covering both China domestic flights and high-speed rail, with routing by transport mode and a shared outer response contract. The supplied code, however, is narrowly scoped to flights: it builds Tongcheng flight URLs, uses airport/city code mappings, parses flight listings, and returns payloads with mode set to "flight". There is no rail-query logic, no train/station handling beyond flight place resolution, and no dispatcher that selects between flight and high-speed rail providers based on mode. While the code does match part of the description (public domestic flight search with fares, airports, times, airline details), it does not implement the broader unified transport capability claimed.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description describes a single unified domestic transport skill that supports both flights and high-speed rail and routes requests by transport mode. The supplied code does not implement that unified transport behavior; it only implements a train provider. It fetches station data and ticket/price data exclusively from 12306 rail endpoints, supports train-specific filters such as seat class and station preferences, and returns payloads with mode="train" and provider="12306-public". There is no evidence of flight support, airport handling, airline data, or any mode-dispatch logic in this chunk. While the train-related portion is consistent with part of the description, the overall declared purpose materially overstates what this code chunk actually does.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 75)May include surrounding context.

md
Flight mode delegates to [scripts/providers/flight_public_service.py](./scripts/providers/flight_public_service.py).

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 102)May include surrounding context.

md
Flight mode delegates to [scripts/providers/flight_public_service.py](./scripts/providers/flight_public_service.py).

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 76)May include surrounding context.

md
Train mode delegates to [scripts/providers/train_public_service.py](./scripts/providers/train_public_service.py).

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 103)May include surrounding context.

md
Train mode delegates to [scripts/providers/train_public_service.py](./scripts/providers/train_public_service.py).

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 90)May include surrounding context.

md
Clearly say whether the result is a flight result or a train result.
   Treat public-source prices as reference prices that can differ from final checkout prices.

## Output Rules

- Prefer up to 5 options unless the user asked for more.
- State the exact travel date in `YYYY-MM-DD`.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This finding describes the same underlying issue: HTTPS certificate validation is disabled without any disclosure, so callers may incorrectly assume transport security is intact. In this travel-query context, an attacker on the network path could spoof or alter route, availability, or fare information and the user would have no indication the connection was untrusted.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The README instructs users to run npx clawhub install fly-flight without pinning a specific package version. npx will fetch the latest published package at execution time, so a compromised publisher account, malicious update, or typosquatted dependency in the resolution chain could cause users to execute unreviewed code during installation.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

This installation example again uses npx clawhub without a pinned version, which causes execution of whatever package version is current at install time. That creates a supply-chain execution risk for anyone following the README, especially because installation commands are typically run with user trust and often broad filesystem access.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill explicitly requires python3, node, and network access, and documents execution of local scripts and an HTTP server, but it does not declare any tool scope such as permissions or allowed-tools. This creates an authorization ambiguity where a host agent may grant broader file, shell, or network capabilities than users expect, increasing the risk of unintended command execution or outbound requests.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The description includes mandatory operational text in both English and Chinese, and later sections repeat core instructions in Chinese only for some content blocks. This can amount to an implicit language/locale constraint without explicit user opt-in or justification, which matches the language/locale policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

These lines present core skill behavior in Chinese without stating whether users may interact in another language or choose their preferred locale. Absent an explicit choice or justification, this can violate the language/locale policy requirement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The script extracts window.__NUXT__ from HTML and executes it with eval, which allows arbitrary JavaScript embedded in the page to run in the local Node.js process. If the HTML source is attacker-controlled or compromised upstream, this can lead to arbitrary code execution, local file access, environment secret disclosure, or command execution in the context running the skill.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a unified transport skill that routes requests for both domestic flights and high-speed rail and returns a shared response shape. This script, however, hard-codes extraction of state.book1.flightLists, which is specific to flight results and has no corresponding rail-path handling in the documented logic here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The file contains user-visible error text in Chinese, and similar hard-coded Chinese messages appear elsewhere, with no indication that the user can opt into another language. That can violate the language/locale policy because the skill effectively forces a specific language for interactions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

These lines return user-facing validation errors only in Chinese, which imposes a fixed language on all users. The file does not provide opt-in language selection or documentation that this locale restriction is intentional and justified.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill's stated purpose is to query domestic flight or rail results from public web sources and return normalized transport options. Spawning an external runtime via subprocess is not an obvious or declared capability for a transport lookup skill, and introduces execution capability beyond simple network retrieval and parsing.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/providers/flight_public_service.py (reported line 84)May include surrounding context.

python
def extract_state_from_html(html: str) -> dict[str, Any]:
    script_path = ROOT / "scripts" / "extract_tongcheng_state.js"
    completed = subprocess.run(
        ["node", str(script_path)],
        input=html,
        capture_output=True,

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
80% confidence
Finding

This user-visible exception message is emitted only in Chinese, continuing a pattern of fixed-language interaction. Without opt-in or documented regional scoping, this is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The code globally disables TLS certificate verification by creating an unverified SSL context and reusing it for all HTTPS requests to 12306 endpoints. This allows man-in-the-middle interception or tampering of train schedules and pricing data, defeating the security guarantees of HTTPS even though the skill only appears to consume public data.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Error messages in the file are presented in Chinese, and the script does not provide any mechanism for user language selection or an explicit statement that the tool is intentionally Chinese-only. That creates a natural-language locale policy concern under the rule requiring opt-in or justified locale constraints.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution, suspicious.insecure_tls_verification

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/extract_tongcheng_state.js:17

HTTPS certificate verification is disabled.

Warn
Code
suspicious.insecure_tls_verification
Location
scripts/providers/train_public_service.py:25