T09 · Insecure Skill Coding Practices
- Location
scripts/script.sh:329- Finding
Arbitrary Command Execution Through AWK Program Injection
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill mostly does what it says, but its cost estimator has a real command-injection bug that could run local commands if unsafe arguments are passed to it.
Review before installing. This skill is useful for generating OpenAI-style request payloads, but do not use its cost command with untrusted or agent-supplied numeric values until the AWK injection is fixed. Also treat generated embedding, batch, and fine-tuning files as content intended for possible external API upload, and avoid including secrets, private records, or confidential datasets unless you have approved that use.
scripts/script.sh:329Arbitrary Command Execution Through AWK Program Injection
The description partially matches the code's core theme, since the script builds chat and embedding payloads and includes fine-tuning/cost utilities. However, the declared purpose says only 'Generate GPT API request payloads' and mentions use for chat completions, embeddings, fine-tuning data, or estimating API costs. The code materially goes beyond that by implementing a broader command-line assistant with batch request generation, standalone file format conversion between CSV and JSONL, and validation/reporting for fine-tuning datasets. These are meaningful undeclared capabilities rather than mere implementation details, so this is a description-behavior mismatch.
Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.
bash scripts/script.sh chat --system "You are a translator" --user "Translate: hello"
bash scripts/script.sh chat --system "Helper" --user "Hi" --model gpt-4o --temperature 0.7
Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.
Estimate API costs based on token count and model pricing.
bash scripts/script.sh cost --tokens 5000 --model gpt-4o
bash scripts/script.sh cost --file input.txt --model gpt-4-turbo
Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.
bash scripts/script.sh cost --tokens 5000 --model gpt-4o
bash scripts/script.sh cost --file input.txt --model gpt-4-turbo
Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.
Generate multiple API request payloads from a list of inputs.
bash scripts/script.sh batch --file prompts.txt --model gpt-4o --output batch_requests.jsonl
Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.
Examples:
bash script.sh chat --system "Translator" --user "Hello"
bash script.sh cost --tokens 5000 --model gpt-4o
bash script.sh convert --input data.csv --to jsonl
EOF
}
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
;;
csv)
# JSONL to CSV - extract keys from first line as headers
local first_line
first_line=$(head -1 "$input")
Without declared permissions the skill's intent is opaque and cannot be validated.
The manifest describes a skill for generating GPT API request payloads and estimating API costs. However, the documented and implemented commands include general-purpose fine-tuning dataset preparation/validation and JSONL↔CSV conversion, which are broader local data-processing capabilities rather than merely building request payloads.
The embed command reads arbitrary local file contents and packages them into an API request payload intended for an external provider, without any warning or confirmation about sensitive-data handling. In an agent skill context, this increases the risk that secrets, proprietary text, or personal data from local files are unintentionally prepared for transmission off-host.
The finetune command converts local CSV content into training-format JSONL suitable for upload to an external model provider, but gives no warning that the data may contain sensitive records and is likely meant for external use. This can enable accidental disclosure of confidential datasets in environments where agents process local files on behalf of users.
The batch command reads prompts from a local file and mass-generates request bodies for chat completions without warning that the file contents are being prepared for external API submission. In an agent setting, bulk preparation amplifies the impact of accidental inclusion of confidential material because many records may be exposed at once.
The cmd_convert function performs bidirectional file format conversion between CSV and JSONL. That is a general data-conversion utility and is not clearly implied by a skill described as generating GPT API request payloads for chat completions, embeddings, fine-tuning data, or estimating API costs.
No suspicious patterns detected.