Back to skill

Security audit

Gpt

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does what it says, but its cost estimator has a real command-injection bug that could run local commands if unsafe arguments are passed to it.

Review before installing. This skill is useful for generating OpenAI-style request payloads, but do not use its cost command with untrusted or agent-supplied numeric values until the AWK injection is fixed. Also treat generated embedding, batch, and fine-tuning files as content intended for possible external API upload, and avoid including secrets, private records, or confidential datasets unless you have approved that use.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/script.sh:329
Finding

Arbitrary Command Execution Through AWK Program Injection

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (13)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The description partially matches the code's core theme, since the script builds chat and embedding payloads and includes fine-tuning/cost utilities. However, the declared purpose says only 'Generate GPT API request payloads' and mentions use for chat completions, embeddings, fine-tuning data, or estimating API costs. The code materially goes beyond that by implementing a broader command-line assistant with batch request generation, standalone file format conversion between CSV and JSONL, and validation/reporting for fine-tuning datasets. These are meaningful undeclared capabilities rather than mere implementation details, so this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · SKILL.md (reported line 29)May include surrounding context.

bash
bash scripts/script.sh chat --system "You are a translator" --user "Translate: hello"
bash scripts/script.sh chat --system "Helper" --user "Hi" --model gpt-4o --temperature 0.7

embed

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

Estimate API costs based on token count and model pricing.

bash
bash scripts/script.sh cost --tokens 5000 --model gpt-4o
bash scripts/script.sh cost --file input.txt --model gpt-4-turbo

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · SKILL.md (reported line 56)May include surrounding context.

bash
bash scripts/script.sh cost --tokens 5000 --model gpt-4o
bash scripts/script.sh cost --file input.txt --model gpt-4-turbo

batch

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

Generate multiple API request payloads from a list of inputs.

bash
bash scripts/script.sh batch --file prompts.txt --model gpt-4o --output batch_requests.jsonl

convert

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · scripts/script.sh (reported line 101)May include surrounding context.

sh
Examples:
  bash script.sh chat --system "Translator" --user "Hello"
  bash script.sh cost --tokens 5000 --model gpt-4o
  bash script.sh convert --input data.csv --to jsonl
EOF
}

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/script.sh (reported line 469)May include surrounding context.

sh
;;

    csv)
      # JSONL to CSV - extract keys from first line as headers
      local first_line
      first_line=$(head -1 "$input")

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a skill for generating GPT API request payloads and estimating API costs. However, the documented and implemented commands include general-purpose fine-tuning dataset preparation/validation and JSONL↔CSV conversion, which are broader local data-processing capabilities rather than merely building request payloads.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The embed command reads arbitrary local file contents and packages them into an API request payload intended for an external provider, without any warning or confirmation about sensitive-data handling. In an agent skill context, this increases the risk that secrets, proprietary text, or personal data from local files are unintentionally prepared for transmission off-host.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The finetune command converts local CSV content into training-format JSONL suitable for upload to an external model provider, but gives no warning that the data may contain sensitive records and is likely meant for external use. This can enable accidental disclosure of confidential datasets in environments where agents process local files on behalf of users.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The batch command reads prompts from a local file and mass-generates request bodies for chat completions without warning that the file contents are being prepared for external API submission. In an agent setting, bulk preparation amplifies the impact of accidental inclusion of confidential material because many records may be exposed at once.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The cmd_convert function performs bidirectional file format conversion between CSV and JSONL. That is a general data-conversion utility and is not clearly implied by a skill described as generating GPT API request payloads for chat completions, embeddings, fine-tuning data, or estimating API costs.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.