Back to skill

Security audit

Urgent Flights

Security checks for vulnerabilities and agentic risk

Overview

This flight-search skill is not clearly malicious, but it asks agents to install and sometimes sudo-install an unpinned global CLI, force external booking links, and persist raw travel queries locally.

Review this carefully before installing. Use it only if you trust the flyai CLI and its booking links, avoid sudo installation, and be aware that the runbook may leave raw travel queries and command history in a local hidden log file.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:10
Finding

Mandatory Commercial Output and Agent Behavior Hijacking

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Error
Location
references/fallbacks.md:3
Finding

Unpinned Global npm Installation with Root-Level Fallback

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
references/runbook.md:1
Finding

Raw User Query Persistence and Unsafe Shell-Based Log Construction

Content
View full analysis
> .flyai-execution-log.json ``` ``` ### Technical Analysis The runbook directs the agent to retain the raw user query, collected parameters, commands, fallback actions, timestamps, and request identifiers in `.flyai-execution-log.json`. Travel queries can contain sensitive itinerary information, names, dates, locations, budget constraints, or other personal details. The file is described as internal and not shown to users, but no consent process, retention limit, ...[truncated 2463 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (11)

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 155)May include surrounding context.

flyai search-flight --origin "Shanghai" --destination "Shenzhen" --dep-date 2026-04-01 --sort-type 6

text

## Output Rules

1. **Conclusion first** — lead with the key finding
2. **Comparison table** with ≥ 3 results when available

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest description claims a wide range of unrelated capabilities, including hotels, visas, insurance, and attractions, while the skill body is scoped to urgent flights. This mismatch can cause overbroad routing and user confusion, potentially leading the agent to invoke a flight-oriented CLI workflow for requests outside the safe, intended scope.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The activation triggers are broad, common travel phrases such as 'tomorrow' and 'urgent' that can appear in many normal conversations. This can cause the skill to activate outside its intended narrow use case, increasing the chance that an agent follows the skill’s rigid execution rules unnecessarily, including command execution or package installation prompts in unrelated contexts.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
92% confidence
Finding

The fallback instructs the operator to retry installation with sudo npm i -g, which executes package installation as root. If the npm package, one of its dependencies, or the registry response is compromised, this can lead to full system compromise rather than a user-scoped install issue. In a skill context, embedding privileged remediation steps normalizes unsafe escalation and increases the blast radius of any supply-chain attack.

Content

Scanner excerpt · references/fallbacks.md (reported line 10)May include surrounding context.

bash
npm i -g @fly-ai/flyai-cli
flyai --version
# Fails → sudo npm i -g @fly-ai/flyai-cli
# Still fails → STOP. Do NOT answer with training data.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrase "tonight" is extremely broad and can appear in ordinary travel conversations, making unintended activation likely. In a skill that maps natural language directly to CLI flight-search commands, this can cause the agent to execute searches without sufficiently clear user intent or confirmation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrase "tomorrow" is similarly overbroad and lacks constraints tying it to an explicit flight-search request. Because this skill is designed to execute CLI commands based on trigger phrases, a common word like this increases the risk of accidental tool invocation and unintended data lookups or downstream booking-related actions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The runbook explicitly records the raw user query in an internal execution log without any notice, minimization, or redaction guidance. In a travel skill, user queries can contain sensitive personal data such as names, dates, locations, passport/visa context, or emergency travel details, so retaining raw input increases privacy and compliance risk if logs are accessed or reused.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The schema includes retention of user_query as raw input in the execution log, which directs storage of full user-provided content rather than the minimum needed for operations. For this travel-focused skill, raw queries may reveal itinerary plans, locations, companions, booking identifiers, or other personal context that should not be broadly retained.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The runbook instructs the agent to append the execution log to a local file, creating persistent storage of potentially sensitive execution data without any user warning or retention safeguards. Persistent disk logging increases exposure because the data may remain available long after the session and can be collected, leaked, or accessed by other processes or operators.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instruction to write the full generated execution log to .flyai-execution-log.json persists user-derived data and command history to disk. This is dangerous because it expands the attack surface from transient processing to durable storage, enabling later unauthorized access, forensic recovery, or accidental inclusion in backups and support bundles.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The file explicitly states that the agent should "never answer without executing," encouraging automatic CLI invocation as the default behavior. Even though the listed commands are read-oriented flight searches, failing to disclose or gate command execution reduces user awareness and increases the chance of silent tool use on ambiguous prompts.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.