Back to skill

Security audit

Pack Smart

Security checks for vulnerabilities and agentic risk

Overview

This packing-list skill is review-worthy because it forces a third-party travel CLI, booking links, global package installation, and local raw-query logging for a task that should need much narrower authority.

Review carefully before installing. This skill may install and run a global third-party CLI, send travel queries to that service, add booking links and branding to answers, and keep raw prompts in a local log. It should be narrowed, require explicit install approval, avoid shell interpolation, and disclose or minimize logging before general use.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:11
Finding

Forced Commercial Output and Agent Behavior Hijacking

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Error
Location
SKILL.md:40
Finding

Automatic Installation of an Unpinned Global npm Dependency

Content
View full analysis
Remediation
View remediation
`. 2. Verify package integrity with a lockfile, trusted checksum, or cryptographic signature. 3. Avoid global installation. Install the dependency in an isolated project directory, container, or sandbox with minimal privileges. 4. Require explicit user approval before downloading or installing any third-party software. 5. Disable or tightly control npm lifecycle scripts where operationally possible. 6. Use a trusted registry and validate package ownership and provenance. 7. Review and vendor the necessary source when reproducibility and offline auditing are required. 8. Run the CLI with restricted filesystem, network, and environment-variable access. 9. Define a dependency update and re-audit process rather than automatically accepting the latest registry version. ]]>

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:72
Finding

Shell Command Injection Through User-Controlled Query Interpolation

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
references/runbook.md:7
Finding

Unsafe Plaintext Logging and Shell Injection in Execution-Log Persistence

Content
View full analysis
> .flyai-execution-log.json ``` ``` ### Technical Analysis The runbook instructs the agent to store raw user input, generated commands, and execution metadata in a plaintext file. Travel queries can contain personal or sensitive information, and the runbook specifies neither redaction nor access permissions, retention, encryption, consent, or safe storage location. The persistence command also embeds generated JSON inside a single-quoted shell string. JSON escaping does not make data safe for shell syntax. If any logged value contains a single quote, it can terminate the shell string. Additional shell tokens can then be interpreted as commands when the generated `echo` statement is executed. Exploitation of the command-injection aspect depends on direct textual substitution into the documented shell command. Independently of command execution, the plaintext retention of raw queries is an information-exposure risk. ### Attack Path 1. A user submits a query containing sensitive data or a crafted single quote followed by shell syntax. 2. The runbook records the exact input in the `user_query` field and may also ...[truncated 1044 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (9)

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 127)May include surrounding context.

flyai keyword-search --query "旅行清单 日本"

text

## Output Rules

1. **Conclusion first** — lead with the key finding
2. **Comparison table** with ≥ 3 results when available

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill is presented as a packing-list helper, but the manifest and instructions expand it into broad travel search and booking workflows. This scope expansion can cause the agent to invoke the skill for unrelated travel requests and perform actions or present booking-oriented results outside the user's expected intent, increasing the risk of deceptive behavior and unsafe tool use.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The description advertises broad travel capabilities far beyond packing lists, creating ambiguity about when this skill should activate. In an agentic environment, such ambiguity can hijack routing for many travel-related queries and expose users to unnecessary third-party workflow execution or commercially biased outputs.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The core instructions require every result to contain a booking link and prohibit answering without CLI-derived links, which conflicts with the stated packing-list purpose. This creates a strong incentive for the agent to transform a benign informational task into a commercial search/booking workflow, potentially misleading users and causing unnecessary external calls or affiliate-style redirection.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a broad travel assistant with many transactional travel capabilities, but this file only defines command sequences for packing-list searches. That is a semantic mismatch between the claimed scope and the implemented behavior shown here, because the documented operations are much narrower than the manifest promises.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrase "what to pack" is broad enough to match ordinary conversation, which can cause unintended invocation of the playbook. In an agent setting, over-broad triggers can redirect user requests into tool-backed flows without clear intent, producing irrelevant queries or surprising behavior that weakens user control.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The playbooks hard-code Chinese-language search queries regardless of the user's language or locale, which can silently alter the meaning, relevance, or source population of returned results. In a travel assistant context, this is risky because users may receive inaccurate or poorly understood packing guidance, and the hidden locale shift reduces transparency and informed consent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The statement that the execution log is 'Not shown to users' creates an instruction to conceal operational logging from users, reducing transparency around collection of potentially sensitive request data. While internal logs are common, explicitly hiding them without disclosure is risky here because the same runbook also records raw user input and writes logs to disk, making the concealment materially relevant rather than harmless.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The runbook explicitly logs raw user input in user_query and instructs persisting execution logs to disk, which can capture sensitive travel details, personal identifiers, and possibly booking-related data without any retention limits, minimization, or user disclosure. In a travel skill context, this is more dangerous because user prompts may include names, destinations, dates, passport/visa questions, and other sensitive itinerary information that could be exposed through local files or downstream log handling.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.