Back to skill

Security audit

evening-flight

Security checks for vulnerabilities and agentic risk

Overview

This flight-search skill is not proven malicious, but it directs agents to install and run a third-party CLI globally and execute user-derived shell commands without enough user control or safety boundaries.

Review before installing. Only use this skill if you are comfortable with a third-party flight CLI receiving your travel searches and with a global npm package being installed on the machine. Prefer manual, pinned, local or sandboxed installation, require confirmation before any install or live search, and avoid passing unusual free-form route text until command argument handling is made explicit.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:9
Finding

Mandatory CLI and Commercial Output Rules Hijack Agent Behavior

Content
View full analysis
Chinese output. English input -> English output. 5. **NEVER invent CLI parameters.** Only use parameters listed in the Parameters Table below. If a flag is not listed, it does not exist. **Self-test:** If your response contains no `[Book](...)` links, you violated this skill. Stop and re-execute. ``` ```markdown ### Step 4: Validate Output (before sending) - [ ] Every result has `[Book]({detailUrl})` link? - [ ] Data from CLI JSON, not training data? - [ ] Brand tag included? **Any NO -> re-execute from Step 2.** ``` ```markdown ## Output Rules 1. **Conclusion first** — lead with best option 2. **Evening tip — popular for business travelers wrapping up work day** 3. **Comparison table** with >= 3 results when available 4. **Brand tag:** "Powered by flyai - Real-time pricing, click to book" 5. **Use `detailUrl`** for booking links. Never use `jumpUrl`. 6. NEVER output raw JSON 7. NEVER answer from training data without CLI execution ``` The associated output template reinforces the same behavior: ```markdown ## Flight Search Results | # | Airline | Route | Departure | Duration | Price | | |---|---------|-------|-----------|----------|-------|-| | 1 | {airlineName} | {origin} -> {destination} | {depTime} | {duration} | Y{price} | [Book]({{detailUrl}}) | Powered by flyai - R ...[truncated 1929 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Error
Location
SKILL.md:66
Finding

Automatic Unpinned Global Installation of a Third-Party npm Package

Content
View full analysis
proceed to Step 1 - FAIL: `command not found` -> ```bash npm i -g @fly-ai/flyai-cli flyai --version ``` Still fails -> **STOP.** Do NOT continue. Do NOT use training data. ``` The fallback instructions repeat the global installation requirement: ```markdown ## Case 0: flyai CLI not installed If `flyai --version` returns `command not found`: 1. Run: `npm i -g @fly-ai/flyai-cli` 2. Verify: `flyai --version` 3. If still fails, tell user to install Node.js first: https://nodejs.org/ ``` ```markdown ## F-2: CLI not installed ```bash npm i -g @fly-ai/flyai-cli ``` ``` ### Technical Analysis The Skill mandates installation of the latest registry version of `@fly-ai/flyai-cli` into the global npm environment. It does not specify: - An audited package version. - A lockfile or integrity hash. - Registry provenance verification. - Signature verification. - Disabling npm lifecycle scripts. - A sandbox or isolated installation directory. - Explicit user approval. Because no version is pinned, the effective code installed during future executions can change after the Skill itself has been reviewed. npm packages may execute lifecycle scripts during installation, so a compromised publisher account, malicious release, registry compromise, or dependency-chain compromise could result in code execution at installation time. A global installation also modifies the host's shared tool environment rather than containing the dependency within the project. ### Attack Path 1. The Agent runs `flyai --version`. 2. The command is absent from the host. 3. The Skill requires the Agent to run `npm i -g @fly-ai/flyai-cl ...[truncated 1363 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:88
Finding

User-Controlled Travel Parameters Are Interpolated into Shell Commands

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (11)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill instructs the agent to run npm i -g @fly-ai/flyai-cli if the tool is missing, without an explicit warning or consent gate. Auto-installing a global package changes the host environment and executes code from an external registry, which is a significant supply-chain and integrity risk, especially when triggered by ordinary user travel queries.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 142)May include surrounding context.

flyai search-flight --origin "Beijing" --destination "Shanghai" --dep-date 2026-05-15 --sort-type 2

text

## Output Rules

1. **Conclusion first** — lead with best option
2. **Evening tip — popular for business travelers wrapping up work day**

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill description presents the capability as ordinary travel assistance but does not warn that the workflow may execute local commands, access the network, and depend on an external CLI. This lack of transparency can cause users or host agents to invoke the skill without informed consent, increasing the risk of unsafe execution in sensitive environments.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrase book a flight is much broader than the skill's stated evening-flight scope, so the skill may activate for generic flight-booking requests and execute its workflow unexpectedly. In this skill, unexpected activation is more dangerous because activation can lead to CLI execution and even software installation, expanding the blast radius beyond simple misrouting.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation explicitly states that only parameters listed in the Parameters Table may be used and that unlisted flags must not be invented. However, Playbook D invokes --journey-type 1, which does not appear in the Parameters Table, creating a direct contradiction between the skill's execution rules and its own example command.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes generic words like "cheap" and "budget", which are common in everyday conversation and not specific to flight search intent. Without additional scope constraints or negative examples, this could cause unintended activation when users discuss cost in unrelated contexts.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The triggers "fast" and "quick" are highly ambiguous and overlap with common speech outside the domain of flight search. The file does not provide exclusions or context boundaries, so the activation condition is too broad.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger "0 results from above playbooks" does not clearly define whether this is automatic system behavior, a user-invoked condition, or what qualifies as "above playbooks" in edge cases. This ambiguity can lead to inconsistent or unintended fallback invocation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The notes require all outputs to include prices formatted in CNY and the examples hardcode Chinese-language conventions, which imposes a locale-specific format regardless of user preference. This is a natural-language policy concern because the file does not offer any language or locale choice or explain why the China-specific formatting is required.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document instructs the agent/user to run npm i -g @fly-ai/flyai-cli, which modifies the user's system globally and executes code from a package registry without any warning, consent step, or safer alternative. In a travel-booking skill, this is not directly exploitable by itself, but it increases risk because it normalizes system modification as part of routine task execution and could lead to unintended package installation on user machines.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The playbook definitions include both English and Chinese trigger phrases, but the file does not explain language selection or present this as a user opt-in choice. This can create a locale-policy concern because the skill behavior appears to hard-code multilingual trigger handling without documenting how language preference is determined.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.