Back to skill

Security audit

SkillOpt

Security checks for vulnerabilities and agentic risk

Overview

This skill is a legitimate skill-optimization harness, but it can execute locally supplied shell commands from task files and agent command templates without enough trust-boundary controls.

Install only if you understand that benchmark task files and agent-command templates can cause local shell commands to run with your user privileges. Use trusted task suites, review any scorer of type "command", prefer non-command scorers, run in a restricted workspace or container, and avoid exposing secrets in the environment when running the harness.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/skillopt.py:131
Finding

Arbitrary Shell Command Execution Through Untrusted Command Scorers

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (10)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 160)May include surrounding context.

md
- For Hermes, prefer standard `SKILL.md` folders and invoke with `hermes -s <skill-or-path>` when testing locally.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

Using shell=True with a command derived from task configuration creates a tool-parameter abuse path: an attacker can encode arbitrary shell operations in the scorer command and gain code execution when the task is scored. This is especially risky because the harness is explicitly designed to process reusable skill/task artifacts, which may cross trust boundaries.

Content

Scanner excerpt · scripts/skillopt.py (reported line 146)May include surrounding context.

python
"skill_path": skill_path,
            },
        )
        proc = subprocess.run(
            command,
            shell=True,
            cwd=scorer.get("cwd"),

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
98% confidence
Finding

The agent runner accepts a free-form command template and executes it through the shell, enabling abuse of tool parameters to run unintended commands. In this skill context, where optimization pipelines may compose prompts, tasks, and commands dynamically, that flexibility materially increases the attack surface and the chance of unsafe command execution.

Content

Scanner excerpt · scripts/skillopt.py (reported line 219)May include surrounding context.

python
"output_path": text_path,
            },
        )
        proc = subprocess.run(
            command,
            shell=True,
            cwd=args.cwd,

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill explicitly instructs use of local scripts, run-directory creation, file writes, and shell command execution, but it does not declare any tool scope or permission boundaries. In an agent ecosystem, that omission can cause the skill to be invoked with broader-than-necessary capabilities, increasing the chance of unsafe file modification or command execution when optimizing untrusted skills or task data.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The description contains broad trigger phrases like optimizing skill files, agent workflows, and controlled self-evolving skills, which can match many common user requests. Overbroad activation raises the risk that the skill is selected in contexts where file-writing and shell-backed optimization behaviors are unnecessary or unsafe, expanding exposure to accidental misuse.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The default prompt instructs the user to use the English trigger phrase "$skillopt" and the rest of the skill metadata is written in Chinese. This creates a language/locale constraint without offering the user a choice or documenting why English-only invocation is required.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

The command-based scorer constructs a shell command from task-controlled configuration and executes it with shell=True. Although placeholders are shell-quoted, the scorer command template itself is arbitrary and comes from untrusted task data, so loading or running a malicious task file can trigger arbitrary local command execution.

Content

Scanner excerpt · scripts/skillopt.py (reported line 146)May include surrounding context.

python
"skill_path": skill_path,
            },
        )
        proc = subprocess.run(
            command,
            shell=True,
            cwd=scorer.get("cwd"),

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The command scorer silently executes task-specified shell commands with no runtime disclosure. Because scorer definitions are read from task files that may be external or untrusted, users may evaluate a benchmark and unknowingly execute attacker-controlled local commands.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

The run command executes an arbitrary agent command string via subprocess.run(..., shell=True). If agent_command or substituted values are influenced by untrusted inputs or unsafe defaults, this harness becomes a generic command-execution primitive capable of running any program on the host.

Content

Scanner excerpt · scripts/skillopt.py (reported line 219)May include surrounding context.

python
"output_path": text_path,
            },
        )
        proc = subprocess.run(
            command,
            shell=True,
            cwd=args.cwd,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code runs arbitrary shell commands supplied by the operator but does not provide an explicit safety warning, confirmation, or trust boundary disclosure at runtime. In a skill-optimization context where tasks and workflows may be shared or generated, the lack of warning increases the chance that users execute dangerous commands without realizing the harness will directly invoke the shell.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.