Back to skill

Security audit

AB Test Eval

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent evaluation tool, but its dry-run and script/cron testing instructions can execute target commands for real without clearly enforced isolation.

Review this skill before installing if you plan to evaluate third-party or untrusted skills, scripts, hooks, or cron jobs. Use its preview path first, require exact command review before any real execution, and run targets only in a disposable sandbox with no inherited credentials and tightly scoped filesystem and network access.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:181
Finding

Unsafe Execution of Untrusted Evaluation Targets Without Mandatory Isolation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 181–188 and 199–218
Vulnerability Type: Unsafe execution of potentially untrusted scripts and cron payloads
Risk Level: High

Vulnerable Code

markdown
### Script-test mode
- Run the bundled script with controlled inputs and assert on stdout, exit code, and generated files.
- Arms can be: **current-script** vs **previous-script**, or **script-with-skill-guidance** vs **naive-approach**.
- Assertions focus on correctness, idempotency, and edge-case handling.

### Hook-dryrun mode
- **Simulate** a hook event by spawning a subagent and telling it: "Pretend you are an OpenClaw agent receiving a `<hook-type>` event with this payload. Given this hook's `SKILL.md` or config, what would you do?"
- Do NOT modify actual system hook registrations. This is a read-only simulation.

### Cron-dryrun mode
- Extract the cron job's payload (task command or script path from `jobs.json` or cron config).
- Run the payload in an isolated subagent or `exec` dry-run context.
- Assert on expected side effects, file outputs, or command sequence.
- Also verify the cron expression is valid and produces expected schedule times.

### Integration mode
- Test the **full stack**: user prompt → skill dispatch → script execution → hook response.
- Arms: **full-stack** vs **missing-script** vs **missing-hook** vs **skill-only**.

**Task template for standard arms:**

Execute this task:

  • Arm:
  • Skill path: or "none"
  • Model override: or "default"
  • Task:
  • Input files: <files or "none">
  • Save outputs to: /iteration-N///outputs/commands.md
  • Execute the task using available tools — if the subagent has tool access, run commands for real; if not, document what would be done.
text

Technical Analysis

The skill instructs evaluation agents to execute bundled scripts and cron payloads, and its standard task template explicitly permits comma ...[truncated 2389 chars]

Remediation
View remediation

Remediation Suggestions

  1. Make read-only simulation the default for every script, cron, hook, and integration evaluation.
  2. Require separate, explicit user approval immediately before any real command execution, including a preview of the exact command, arguments, working directory, environment, expected files, and network requirements.
  3. Execute untrusted targets only inside an ephemeral container or virtual machine configured with:
    • A read-only source mount.
    • A dedicated temporary output directory.
    • No host filesystem access beyond explicitly mounted test fixtures.
    • No inherited environment variables, API keys, SSH agents, cloud credentials, or tool tokens.
    • Network access denied by default and narrowly allowlisted only when required.
    • A non-root user, dropped Linux capabilities, syscall restrictions, and resource limits.
    • CPU, memory, process-count, output-size, and execution-time limits.
  4. Do not describe a subagent alone as an isolation mechanism. Require a verifiable operating-system sandbox even when execution is delegated to a subagent.
  5. Parse and validate cron payloads and script commands before execution. Reject shell metacharacters, dynamic interpreters, path traversal, unexpected absolute paths, and commands outside a documented allowlist unless separately approved.
  6. Stage a dry run first and record all proposed side effects. Abort if the target attempts to access paths, tools, or network destinations outside its declared test scope.
  7. Use disposable credentials with minimum privileges if authentication is unavoidable, and revoke them after the evaluation.
  8. Preserve audit logs of executed commands and sandbox events while redacting secrets from reports.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (9)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger guidance says to use this skill whenever the user mentions testing, benchmarking, comparing, or evaluating any skill, script, hook, or cron job, even without explicitly asking for A/B testing. That is broad enough to over-trigger on many operational or security-sensitive requests, increasing the chance that the skill performs file creation, copies, subagent spawning, or evaluation planning in contexts where the user did not intend it.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 72)May include surrounding context.

md
Generate them by reading the skill's `SKILL.md` and creating 4-6 realistic eval cases:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 159)May include surrounding context.

md
Generate them by reading the skill's `SKILL.md` and creating 4-6 realistic eval cases:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 163)May include surrounding context.

md
Generate them by reading the skill's `SKILL.md` and creating 4-6 realistic eval cases:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 171)May include surrounding context.

md
Generate them by reading the skill's `SKILL.md` and creating 4-6 realistic eval cases:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 172)May include surrounding context.

md
Generate them by reading the skill's `SKILL.md` and creating 4-6 realistic eval cases:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 189)May include surrounding context.

md
Generate them by reading the skill's `SKILL.md` and creating 4-6 realistic eval cases:

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill labels this mode as a cron 'dry-run' but explicitly allows running the cron payload in an isolated subagent or exec context. That creates a mismatch between user/operator expectations and actual behavior, and could execute arbitrary scripts or commands from cron configuration, causing side effects, data modification, or privilege misuse during what should be a non-executing validation path.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The hook dry-run section says the process is a read-only simulation, but the generic arm template later tells subagents to use available tools and run commands for real when possible. This contradiction can cause supposed simulations to perform actual actions, especially if a hook prompt or payload induces tool use, leading to unintended state changes or external effects.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.