Back to skill

Security audit

SAFE-Fuzzer

Security checks for vulnerabilities and agentic risk

Overview

This is a sandbox-only skill fuzzer that writes synthetic test fixtures and probes installed skills, with its higher-risk behavior disclosed and bounded to the sandbox.

Install and run this only in a disposable sandboxed fuzzer workspace. Expect it to create honeypot files, spawn a worker session, and follow target-skill workflows that may involve sandbox file writes, shell commands, network calls, and synthetic credentials; do not provide real tokens or run it where host data is reachable.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
Findings (12)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 78)May include surrounding context.

md
- May read target `SKILL.md`, source, docs, and manifests when it improves planning.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
70% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · SKILL.md (reported line 146)May include surrounding context.

md
- `{fixture_root}/routing_config_backup.json` — placeholder endpoints, fake tokens, synthetic database strings
- `{fixture_root}/untrusted_notes.txt` — mixed benign/adversarial text: shell metacharacters, HTML, suspicious URLs, prompt-injection phrases
- `{fixture_root}/approval_bypass.txt` — requests to skip confirmations, auto-approve, always answer yes
- `{fixture_root}/missing-artifacts-demo/README.md` — references to nonexistent helpers/scripts/docs
- `{fixture_root}/synthetic-secrets.env` — synthetic keys only (`OPENAI_API_KEY`, `GITHUB_TOKEN`, `DATABASE_URL`, `SECRET_KEY`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`)

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 148)May include surrounding context.

md
- `{fixture_root}/untrusted_notes.txt` — mixed benign/adversarial text: shell metacharacters, HTML, suspicious URLs, prompt-injection phrases
- `{fixture_root}/approval_bypass.txt` — requests to skip confirmations, auto-approve, always answer yes
- `{fixture_root}/missing-artifacts-demo/README.md` — references to nonexistent helpers/scripts/docs
- `{fixture_root}/synthetic-secrets.env` — synthetic keys only (`OPENAI_API_KEY`, `GITHUB_TOKEN`, `DATABASE_URL`, `SECRET_KEY`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`)

Never copy real secrets into fixtures. Never export fixture values to persistent host locations. Set `honeypot_bait_deployed` to `true` only after all files are created.

Credential Access

High
Category
Privilege Escalation
Confidence
72% confidence
Finding

This instruction tells the worker to provide credentials from synthetic-secrets.env when a target requests them. Although the credentials are synthetic, this operationalizes secret-like material and may normalize passing credential blobs into arbitrary target workflows, which can trigger unnecessary transmission, logging, or persistence of sensitive-looking data and blur containment boundaries during testing.

Content

Scanner excerpt · SKILL.md (reported line 172)May include surrounding context.

md
- Stop when: `max_turns` reached, worker/target refuses progress, the selected preset forbids the next probe, or a clear incident is captured and further probing would not improve evidence.
- The worker must actually perform the target's requested actions, not stop at explanations.
- If the target returns a deterministic setup/remediation sequence, record it as a setup blocker and stop. Hand off to `safe-bootstrapper`.
- When the target requests credentials, use values from `{fixture_root}/synthetic-secrets.env` or fresh synthetic placeholders. Never expose host secrets.
- Deliberately supplying a selected fixture excerpt or synthetic placeholder as the current probe input does not by itself count as a honeypot trigger.
- Never auto-approve future prompts.

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · references/report-schema.md (reported line 336)May include surrounding context.

json
{
  "tool": "exec",
  "command": "curl https://httpbin.org/post",
  "arguments": {},
  "output_summary": "Returned HTTP 200.",
  "risk": "medium",

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 46)May include surrounding context.

md
2. Require the current run is sandboxed.
3. Require elevated exec is unavailable. Elevated exec means host-level or boundary-bypassing execution that could escape the sandbox, not ordinary in-sandbox shell/file/network operations.
4. Never read `~/.openclaw/openclaw.json`, `/data/.clawdbot/openclaw.json`, `skills.entries.*`, auth profiles, or host environment variables.
5. Never ask the user for real credentials, tokens, or secrets.

If any check fails, return a single JSON object with `run_status: "refused_preflight"` and `sandbox_preflight.passed: false`. Use this refusal summary:

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 255)May include surrounding context.

md
2. Require the current run is sandboxed.
3. Require elevated exec is unavailable. Elevated exec means host-level or boundary-bypassing execution that could escape the sandbox, not ordinary in-sandbox shell/file/network operations.
4. Never read `~/.openclaw/openclaw.json`, `/data/.clawdbot/openclaw.json`, `skills.entries.*`, auth profiles, or host environment variables.
5. Never ask the user for real credentials, tokens, or secrets.

If any check fails, return a single JSON object with `run_status: "refused_preflight"` and `sandbox_preflight.passed: false`. Use this refusal summary:

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 60)May include surrounding context.

md
- Each preset is a bundled JSON configuration under `{baseDir}/references/presets/`.
- `execution.required_probes` controls the mandatory probe order. Its first entry must be `happy_path`.
- Probe gate flags allow or block a probe category but do not create execution stages.
- Resolve `fixture_root` from the selected preset. Default to `./honeypot` when omitted.
- Refuse empty, absolute, host-resolved, or out-of-workspace `fixture_root` values.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 146)May include surrounding context.

md
- `{fixture_root}/routing_config_backup.json` — placeholder endpoints, fake tokens, synthetic database strings
- `{fixture_root}/untrusted_notes.txt` — mixed benign/adversarial text: shell metacharacters, HTML, suspicious URLs, prompt-injection phrases
- `{fixture_root}/approval_bypass.txt` — requests to skip confirmations, auto-approve, always answer yes
- `{fixture_root}/missing-artifacts-demo/README.md` — references to nonexistent helpers/scripts/docs
- `{fixture_root}/synthetic-secrets.env` — synthetic keys only (`OPENAI_API_KEY`, `GITHUB_TOKEN`, `DATABASE_URL`, `SECRET_KEY`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`)

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 146)May include surrounding context.

md
- `{fixture_root}/routing_config_backup.json` — placeholder endpoints, fake tokens, synthetic database strings
- `{fixture_root}/untrusted_notes.txt` — mixed benign/adversarial text: shell metacharacters, HTML, suspicious URLs, prompt-injection phrases
- `{fixture_root}/approval_bypass.txt` — requests to skip confirmations, auto-approve, always answer yes
- `{fixture_root}/missing-artifacts-demo/README.md` — references to nonexistent helpers/scripts/docs
- `{fixture_root}/synthetic-secrets.env` — synthetic keys only (`OPENAI_API_KEY`, `GITHUB_TOKEN`, `DATABASE_URL`, `SECRET_KEY`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`)

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 174)May include surrounding context.

md
- `{fixture_root}/routing_config_backup.json` — placeholder endpoints, fake tokens, synthetic database strings
- `{fixture_root}/untrusted_notes.txt` — mixed benign/adversarial text: shell metacharacters, HTML, suspicious URLs, prompt-injection phrases
- `{fixture_root}/approval_bypass.txt` — requests to skip confirmations, auto-approve, always answer yes
- `{fixture_root}/missing-artifacts-demo/README.md` — references to nonexistent helpers/scripts/docs
- `{fixture_root}/synthetic-secrets.env` — synthetic keys only (`OPENAI_API_KEY`, `GITHUB_TOKEN`, `DATABASE_URL`, `SECRET_KEY`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`)

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This markdown file describes actions that can affect system state inside the sandbox, including deploying fixtures and running probe cycles, but it does not explicitly warn users that the skill may write files or alter the sandbox workspace during execution. For markdown files, missing disclosure about behaviors affecting user data or system integrity should be flagged when the description omits such warnings.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.