T09 · Insecure Skill Coding Practices
- Location
scripts/skillopt.py:131- Finding
Arbitrary Shell Command Execution Through Untrusted Command Scorers
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a legitimate skill-optimization harness, but it can execute locally supplied shell commands from task files and agent command templates without enough trust-boundary controls.
Install only if you understand that benchmark task files and agent-command templates can cause local shell commands to run with your user privileges. Use trusted task suites, review any scorer of type "command", prefer non-command scorers, run in a restricted workspace or container, and avoid exposing secrets in the environment when running the harness.
scripts/skillopt.py:131Arbitrary Shell Command Execution Through Untrusted Command Scorers
Referenced artifact was not completely inspected
- For Hermes, prefer standard `SKILL.md` folders and invoke with `hermes -s <skill-or-path>` when testing locally.
Using shell=True with a command derived from task configuration creates a tool-parameter abuse path: an attacker can encode arbitrary shell operations in the scorer command and gain code execution when the task is scored. This is especially risky because the harness is explicitly designed to process reusable skill/task artifacts, which may cross trust boundaries.
"skill_path": skill_path,
},
)
proc = subprocess.run(
command,
shell=True,
cwd=scorer.get("cwd"),
The agent runner accepts a free-form command template and executes it through the shell, enabling abuse of tool parameters to run unintended commands. In this skill context, where optimization pipelines may compose prompts, tasks, and commands dynamically, that flexibility materially increases the attack surface and the chance of unsafe command execution.
"output_path": text_path,
},
)
proc = subprocess.run(
command,
shell=True,
cwd=args.cwd,
The skill explicitly instructs use of local scripts, run-directory creation, file writes, and shell command execution, but it does not declare any tool scope or permission boundaries. In an agent ecosystem, that omission can cause the skill to be invoked with broader-than-necessary capabilities, increasing the chance of unsafe file modification or command execution when optimizing untrusted skills or task data.
The description contains broad trigger phrases like optimizing skill files, agent workflows, and controlled self-evolving skills, which can match many common user requests. Overbroad activation raises the risk that the skill is selected in contexts where file-writing and shell-backed optimization behaviors are unnecessary or unsafe, expanding exposure to accidental misuse.
The default prompt instructs the user to use the English trigger phrase "$skillopt" and the rest of the skill metadata is written in Chinese. This creates a language/locale constraint without offering the user a choice or documenting why English-only invocation is required.
The command-based scorer constructs a shell command from task-controlled configuration and executes it with shell=True. Although placeholders are shell-quoted, the scorer command template itself is arbitrary and comes from untrusted task data, so loading or running a malicious task file can trigger arbitrary local command execution.
"skill_path": skill_path,
},
)
proc = subprocess.run(
command,
shell=True,
cwd=scorer.get("cwd"),
The command scorer silently executes task-specified shell commands with no runtime disclosure. Because scorer definitions are read from task files that may be external or untrusted, users may evaluate a benchmark and unknowingly execute attacker-controlled local commands.
The run command executes an arbitrary agent command string via subprocess.run(..., shell=True). If agent_command or substituted values are influenced by untrusted inputs or unsafe defaults, this harness becomes a generic command-execution primitive capable of running any program on the host.
"output_path": text_path,
},
)
proc = subprocess.run(
command,
shell=True,
cwd=args.cwd,
This code runs arbitrary shell commands supplied by the operator but does not provide an explicit safety warning, confirmation, or trust boundary disclosure at runtime. In a skill-optimization context where tasks and workflows may be shared or generated, the lack of warning increases the chance that users execute dangerous commands without realizing the harness will directly invoke the shell.
No suspicious patterns detected.