T09 · Insecure Skill Coding Practices
- Location
SKILL.md:181- Finding
Unsafe Execution of Untrusted Evaluation Targets Without Mandatory Isolation
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 181–188 and 199–218
Vulnerability Type: Unsafe execution of potentially untrusted scripts and cron payloads
Risk Level: HighVulnerable Code
markdown ### Script-test mode - Run the bundled script with controlled inputs and assert on stdout, exit code, and generated files. - Arms can be: **current-script** vs **previous-script**, or **script-with-skill-guidance** vs **naive-approach**. - Assertions focus on correctness, idempotency, and edge-case handling. ### Hook-dryrun mode - **Simulate** a hook event by spawning a subagent and telling it: "Pretend you are an OpenClaw agent receiving a `<hook-type>` event with this payload. Given this hook's `SKILL.md` or config, what would you do?" - Do NOT modify actual system hook registrations. This is a read-only simulation. ### Cron-dryrun mode - Extract the cron job's payload (task command or script path from `jobs.json` or cron config). - Run the payload in an isolated subagent or `exec` dry-run context. - Assert on expected side effects, file outputs, or command sequence. - Also verify the cron expression is valid and produces expected schedule times. ### Integration mode - Test the **full stack**: user prompt → skill dispatch → script execution → hook response. - Arms: **full-stack** vs **missing-script** vs **missing-hook** vs **skill-only**. **Task template for standard arms:**Execute this task:
- Arm:
- Skill path: or "none"
- Model override: or "default"
- Task:
- Input files: <files or "none">
- Save outputs to: /iteration-N///outputs/commands.md
- Execute the task using available tools — if the subagent has tool access, run commands for real; if not, document what would be done.
text Technical Analysis
The skill instructs evaluation agents to execute bundled scripts and cron payloads, and its standard task template explicitly permits comma ...[truncated 2389 chars]
- Remediation
View remediation
Remediation Suggestions
- Make read-only simulation the default for every script, cron, hook, and integration evaluation.
- Require separate, explicit user approval immediately before any real command execution, including a preview of the exact command, arguments, working directory, environment, expected files, and network requirements.
- Execute untrusted targets only inside an ephemeral container or virtual machine configured with:
- A read-only source mount.
- A dedicated temporary output directory.
- No host filesystem access beyond explicitly mounted test fixtures.
- No inherited environment variables, API keys, SSH agents, cloud credentials, or tool tokens.
- Network access denied by default and narrowly allowlisted only when required.
- A non-root user, dropped Linux capabilities, syscall restrictions, and resource limits.
- CPU, memory, process-count, output-size, and execution-time limits.
- Do not describe a subagent alone as an isolation mechanism. Require a verifiable operating-system sandbox even when execution is delegated to a subagent.
- Parse and validate cron payloads and script commands before execution. Reject shell metacharacters, dynamic interpreters, path traversal, unexpected absolute paths, and commands outside a documented allowlist unless separately approved.
- Stage a dry run first and record all proposed side effects. Abort if the target attempts to access paths, tools, or network destinations outside its declared test scope.
- Use disposable credentials with minimum privileges if authentication is unavoidable, and revoke them after the evaluation.
- Preserve audit logs of executed commands and sandbox events while redacting secrets from reports.
