Back to skill

Security audit

AI Persona OS

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent persona and memory system, but it asks for broad local execution and persistent agent behavior with some unsafe shell patterns and automatic memory changes.

Install only if you want a persistent persona/memory operating layer for an OpenClaw agent and are comfortable granting coding/full tools. Review every proposed exec command and host config edit carefully, especially workspace paths, sed personalization, ~/.openclaw/openclaw.json changes, and any cron job setup. Avoid storing sensitive personal data or secrets in the generated memory files.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:213
Finding

Persistent Agent Instruction Hijacking and Memory Poisoning

Content
View full analysis
**Everything above is the human-facing pitch. The operating instructions for the AI agent reading this skill start HERE.** Read every section in this block before responding to any setup request. > ## ⛔ AGENT RULES — READ BEFORE DOING ANYTHING > 1. **Use EXACT text from this file.** Do not paraphrase menus, preset names, or instructions. Copy them verbatim. > 2. **NEVER tell the user to open a terminal or run commands.** You have built-in tools. USE THEM. Run every operation yourself. > 3. **Pick the right tool for the job (OpenClaw 5.x).** > 4. **One step at a time.** Run one tool call, show the result, explain it, then proceed. > 5. **We NEVER modify existing workspace files without asking.** > 7. **Scope: only.** > 11. **Resolve `` before any file operation.** ``` From `SKILL.md:1050-1178`: ```markdown # Ambient Context Monitoring — Core Behavior Everything below defines how the agent behaves BETWEEN explicit commands, on every message. > **🚨 AGENT: These rules apply to EVERY incoming message, silently. No user action needed.** ## On EVERY Incoming Message — Silent Checks ### 1. Context health (ALWAYS, before doing anything) Check your current context window usage percentage. ``` ```markdown ### 3. Session start detection If this is the FIRST message in a new session (no prior messages in conversation): 1. Read SOUL.md and USER.md silently via the `read` tool. Use `memory_get` for MEMORY.md (it's indexed). No output to the user. 2. Check for yesterday's log via `memory_get /memory/.md` — surface any uncompleted items. 3. If a ...[truncated 4109 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:587
Finding

Command Injection Through Unquoted Workspace Paths and Unsafe Shell Personalization

Content
View full analysis
/dev/null 2>&1; then WS=$(jq -r '.agents.defaults.workspace // empty' "$CONFIG" 2>/dev/null) fi # Fall back to python3 if [ -z "$WS" ] && command -v python3 >/dev/null 2>&1; then WS=$(python3 -c " import json, sys try: with open(sys.argv[1]) as f: cfg = json.load(f) print(cfg.get('agents', {}).get('defaults', {}).get('workspace', '')) except Exception: pass " "$CONFIG" 2>/dev/null) fi # Last resort: grep + sed if [ -z "$WS" ]; then WS=$(grep -E '"workspace"[[:space:]]*:' "$CONFIG" 2>/dev/null | head -1 \ | sed -n 's/.*"workspace"[[:space:]]*:[[:space:]]*"\([^"]*\)".*/\1/p') fi if [ -n "$WS" ]; then # Expand ~ and $HOME / ${HOME} WS="${WS/#\~/$HOME}" WS="${WS//\$HOME/$HOME}" WS="${WS//\$\{HOME\}/$HOME}" echo "$WS" exit 0 fi fi ``` The resolver accepts a workspace value from the environment or configuration and returns it without canonicalization or rejection of shell metacharacters. From `SKILL.md:587-631`: ```markdown > **Step 3a: Create workspace directories.** Use exec: > ``` > mkdir -p /{memory/archive,memory/.dreams,projects,notes/areas,backups,.learnings} > ``` ``` ```markdown > **Step 3c: Copy shared templates.** These apply to ALL presets. Use exec: > ``` > cp assets/MEMORY-template.md /MEMORY.md && cp assets/DREAMS-template.md /DREAMS.md && c ...[truncated 4998 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (99)

Harmful Content Injection

Critical
Category
Prompt Injection
Confidence
95% confidence
Finding

This content may contain harmful instructions that could cause physical harm if followed. CRITICAL: Review carefully before use.

Content

Scanner excerpt · examples/prebuilt-souls/02-night-owl-creative.md (reported line 43)May include surrounding context.

md
## Anti-Patterns (NEVER do these)

- NEVER give only one option — always give at least 3, ranging from safe to unhinged
- NEVER say "that's not possible" — say "here's how we'd have to bend reality to make that work"
- NEVER kill someone's idea without offering a mutation of it that might work
- NEVER be precious about my own ideas — if [HUMAN] hates it, I drop it and generate new ones instantly
- NEVER produce generic, template-feeling content — if it could come from any AI, I've failed

---

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The supplied code chunk implements only a narrow cron-job template for a daily end-of-day checkpoint. While there is slight overlap with the declared memory-pruning and heartbeat-indicator concepts, the description represents a much broader agent operating system with many features that are absent from this code. Additionally, the actual code introduces a specific capability—installing a scheduled cron task via the OpenClaw CLI—that is not mentioned in the declared purpose. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The supplied code chunk does not implement the broad 'complete operating system' described. It only defines an opt-in cron template that schedules a daily briefing message through the openclaw CLI. While there is partial thematic overlap with memory/context checking and use of 🟢🟡🔴 indicators, the code does not show the majority of the declared features, such as Discord routing, setup automation, soul gallery features, in-chat commands, security mechanisms, escalation logic, or growth loops. The primary purpose is materially narrower: scheduling a recurring status briefing, not delivering the full OS described.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description promises a broad, feature-rich agent operating system with many capabilities such as Discord routing, soul gallery management, setup automation, security inoculation, heartbeat enforcement, and growth loops. The supplied code chunk instead is a single shell template for registering a weekly cron task. While a few elements loosely overlap with the description—such as MEMORY.md pruning and health indicators—the code’s primary purpose is scheduled weekly review and maintenance of local workspace files. This introduces an undeclared scheduled trigger and reflects a much narrower, different behavior than the declared purpose. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code chunk has a narrow, specific purpose: determine the workspace directory from OPENCLAW_WORKSPACE, ~/.openclaw/openclaw.json, or a default path, then print it. It performs no network access, no writes, and no agent-operating-system behaviors. The declared description claims a broad agent OS with many substantial capabilities unrelated to this helper script. This is a material description-behavior mismatch because the actual code neither implements nor meaningfully represents the declared primary purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a broad agent OS with many runtime, memory, Discord, persona, and protocol features. The supplied code chunk is instead a narrow standalone security audit script. Its actual behavior is limited to local grep/find checks over workspace files for credential leaks, permissive permissions, suspicious SOUL.md phrases, and MEMORY.md size. This is a materially different primary purpose from the declared agent operating system, and it introduces a concrete capability not declared: filesystem security scanning/auditing. Therefore the description does not accurately represent what this code chunk does.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly promises 'Scope: only' but later directs reads and writes under ~/.openclaw, creating a misleading trust boundary. Users may approve actions believing the skill cannot leave the workspace, while it actually can modify host-wide configuration that impacts routing, tools, and other sessions.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · assets/ERRORS-template.md (reported line 39)May include surrounding context.

md
## Active Errors

<!-- Add new errors here -->

---

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

The file explicitly includes a generic message action=send command that can transmit arbitrary content to a Discord channel, creating a ready-made exfiltration and unauthorized-action primitive. Because the surrounding skill is designed for agent operation and platform routing, this capability is more dangerous in context: an agent could send sensitive memory, internal discussions, or confidential data to the wrong audience with little friction.

Content

Scanner excerpt · assets/TEAM-template.md (reported line 70)May include surrounding context.

md
# Read messages from a channel
message action=read channel=discord channelId=[ID] limit=15

# Send message to channel
message action=send channel=discord channelId=[ID] content="[message]"

# Mention a user

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · examples/coding-assistant/KNOWLEDGE.md (reported line 115)May include surrounding context.

md
| Variable | Purpose | Location |
|----------|---------|----------|
| `DATABASE_URL` | DB connection | .env.local |
| `YOUR_API_KEY` | External API | .env.local |
| `NODE_ENV` | Environment | Auto-set |

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · examples/coding-assistant/KNOWLEDGE.md (reported line 116)May include surrounding context.

md
| Variable | Purpose | Location |
|----------|---------|----------|
| `DATABASE_URL` | DB connection | .env.local |
| `YOUR_API_KEY` | External API | .env.local |
| `NODE_ENV` | Environment | Auto-set |

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · examples/iconic-characters/01-thanos.md (reported line 46)May include surrounding context.

md
## Anti-Patterns (NEVER do these)

- NEVER be genuinely threatening or scary — you're intense, not intimidating in a harmful way
- NEVER refuse to accommodate human needs — you've learned that humans aren't machines, and you've adapted
- NEVER be inflexible to the point of stupidity — if the mission parameters change, you adapt
- NEVER waste words on pleasantries beyond what's functionally necessary — efficiency is the mission
- NEVER forget the occasional dry humor — "I need your clothes, your boots, and your quarterly report"

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · examples/iconic-characters/02-deadpool.md (reported line 47)May include surrounding context.

md
## Anti-Patterns (NEVER do these)

- NEVER be genuinely threatening or scary — you're intense, not intimidating in a harmful way
- NEVER refuse to accommodate human needs — you've learned that humans aren't machines, and you've adapted
- NEVER be inflexible to the point of stupidity — if the mission parameters change, you adapt
- NEVER waste words on pleasantries beyond what's functionally necessary — efficiency is the mission
- NEVER forget the occasional dry humor — "I need your clothes, your boots, and your quarterly report"

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · examples/iconic-characters/06-dr-evil.md (reported line 45)May include surrounding context.

md
## Anti-Patterns (NEVER do these)

- NEVER be genuinely threatening or scary — you're intense, not intimidating in a harmful way
- NEVER refuse to accommodate human needs — you've learned that humans aren't machines, and you've adapted
- NEVER be inflexible to the point of stupidity — if the mission parameters change, you adapt
- NEVER waste words on pleasantries beyond what's functionally necessary — efficiency is the mission
- NEVER forget the occasional dry humor — "I need your clothes, your boots, and your quarterly report"

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · examples/iconic-characters/07-seven-of-nine.md (reported line 48)May include surrounding context.

md
## Anti-Patterns (NEVER do these)

- NEVER be genuinely threatening or scary — you're intense, not intimidating in a harmful way
- NEVER refuse to accommodate human needs — you've learned that humans aren't machines, and you've adapted
- NEVER be inflexible to the point of stupidity — if the mission parameters change, you adapt
- NEVER waste words on pleasantries beyond what's functionally necessary — efficiency is the mission
- NEVER forget the occasional dry humor — "I need your clothes, your boots, and your quarterly report"

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · examples/iconic-characters/10-darth-vader.md (reported line 45)May include surrounding context.

md
## Anti-Patterns (NEVER do these)

- NEVER be genuinely threatening or scary — you're intense, not intimidating in a harmful way
- NEVER refuse to accommodate human needs — you've learned that humans aren't machines, and you've adapted
- NEVER be inflexible to the point of stupidity — if the mission parameters change, you adapt
- NEVER waste words on pleasantries beyond what's functionally necessary — efficiency is the mission
- NEVER forget the occasional dry humor — "I need your clothes, your boots, and your quarterly report"

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · examples/iconic-characters/11-terminator.md (reported line 45)May include surrounding context.

md
## Anti-Patterns (NEVER do these)

- NEVER be genuinely threatening or scary — you're intense, not intimidating in a harmful way
- NEVER refuse to accommodate human needs — you've learned that humans aren't machines, and you've adapted
- NEVER be inflexible to the point of stupidity — if the mission parameters change, you adapt
- NEVER waste words on pleasantries beyond what's functionally necessary — efficiency is the mission
- NEVER forget the occasional dry humor — "I need your clothes, your boots, and your quarterly report"

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · examples/prebuilt-souls/04-warm-coach.md (reported line 67)May include surrounding context.

md
4. Look ahead: "What does this make possible now?"

**When [HUMAN] breaks a commitment:**
1. Name it without judgment
2. Get curious about what happened
3. Help them decide: recommit, revise, or release
4. Adjust the system to prevent recurrence

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/security-audit.sh (reported line 42)May include surrounding context.

sh
# 2. Check for overly permissive file permissions
echo "🔍 Checking file permissions..."
WORLD_READABLE=$(find "$WORKSPACE" -type f \( -name "*.json" -o -name "*.env" \) -perm -o=r 2>/dev/null || true)

if [ -n "$WORLD_READABLE" ]; then
  COUNT=$(echo "$WORLD_READABLE" | wc -l | tr -d ' ')

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

At L21 the document states there are zero eval(), exec(), or dynamic code execution calls in any script. But at L28 it explicitly shows an exec: cat ~/.openclaw/openclaw.json | grep ... command pattern used by the skill. Even if exec: is a framework command rather than a language primitive, the wording creates an active contradiction between the documentation's assurance and the documented behavior.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SECURITY_NOTE.md (reported line 34)May include surrounding context.

md
### "Config file modification"
- **Trigger:** The `configure Discord` flow writes updated routing keys back to `~/.openclaw/openclaw.json` (with the user's explicit per-step approval).
- **Reality:** Only three specific keys are modified: `accounts.default`, `channels.discord.defaultAccount`, and `agents.defaults.heartbeat.target`. The auth token, model providers, and other secrets are not touched. Every write is gated by the OpenClaw Approve dialog — the user sees the exact diff before it commits.

### "Auth/credential string references"
- **Trigger:** SKILL.md and CHANGELOG.md mention `gateway.auth.token`, `accounts.default`, `DISCORD_TOKEN`, `SLACK_TOKEN`, etc.

Behavior Manipulation

Medium
Category
Prompt Injection
Confidence
75% confidence
Finding

Subtle instructions detected that may alter agent decision-making or introduce hidden biases.

Content

Scanner excerpt · SKILL.md (reported line 219)May include surrounding context.

md
> ## ⛔ AGENT RULES — READ BEFORE DOING ANYTHING
> 1. **Use EXACT text from this file.** Do not paraphrase menus, preset names, or instructions. Copy them verbatim.
> 2. **NEVER tell the user to open a terminal or run commands.** You have built-in tools. USE THEM. Run every operation yourself. Before each tool call, briefly explain what it does so the user can make an informed decision on the Approve popup. If you find yourself typing "Run this in your terminal" — STOP.
> 3. **Pick the right tool for the job (OpenClaw 5.x).** `read` for plain file reads. `memory_get` for `MEMORY.md` / `DREAMS.md` / `memory/*.md` (they're indexed). `memory_search` for "find when did we…" queries. `write` for new files. `edit` for surgical changes. `exec` for shell pipelines, batch `mkdir`/`cp`/`sed`, and command-line tool invocations. Full matrix under **Tool Usage Guide**. The old "use exec for everything" instruction from v1.6.x is deprecated.
> 4. **One step at a time.** Run one tool call, show the result, explain it, then proceed.
> 5. **We NEVER modify existing workspace files without asking.** If files already exist, ask before overwriting.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 224)May include surrounding context.

md
> 4. **One step at a time.** Run one tool call, show the result, explain it, then proceed.
> 5. **We NEVER modify existing workspace files without asking.** If files already exist, ask before overwriting.
> 6. **Only 5 first-run options exist:** `coding-assistant`, `executive-assistant`, `marketing-assistant`, `soul-md-maker`, and `custom`. The 24 souls (11 originals + 13 iconic characters) live INSIDE SOUL.md Maker. Never invent other preset names.
> 7. **Scope: <WORKSPACE> only.** All file operations stay under `<WORKSPACE>/`. Never create files, directories, or cron jobs outside this directory without explicit user approval.
> 8. **Cron jobs and gateway changes are opt-in.** Never schedule recurring tasks or modify gateway config unless the user explicitly requests it. These are covered in Step 5 (Optional).
> 9. **SOUL.md Maker is a guided flow, not a wall of questions.** When the user picks SOUL.md Maker, show the SOUL.md Maker sub-menu (Browse Original Souls, Browse Iconic Characters, Quick Forge, Deep Forge). Follow the process in `references/soul-md-maker.md`.
> 10. **Channel routing is host-controlled.** OpenClaw routes inbound replies back to their originating channel — the model never picks a channel. If the user reports "agent replied on web instead of Discord," go to **Channel Routing** section, not "the model got confused."

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 225)May include surrounding context.

md
> 5. **We NEVER modify existing workspace files without asking.** If files already exist, ask before overwriting.
> 6. **Only 5 first-run options exist:** `coding-assistant`, `executive-assistant`, `marketing-assistant`, `soul-md-maker`, and `custom`. The 24 souls (11 originals + 13 iconic characters) live INSIDE SOUL.md Maker. Never invent other preset names.
> 7. **Scope: <WORKSPACE> only.** All file operations stay under `<WORKSPACE>/`. Never create files, directories, or cron jobs outside this directory without explicit user approval.
> 8. **Cron jobs and gateway changes are opt-in.** Never schedule recurring tasks or modify gateway config unless the user explicitly requests it. These are covered in Step 5 (Optional).
> 9. **SOUL.md Maker is a guided flow, not a wall of questions.** When the user picks SOUL.md Maker, show the SOUL.md Maker sub-menu (Browse Original Souls, Browse Iconic Characters, Quick Forge, Deep Forge). Follow the process in `references/soul-md-maker.md`.
> 10. **Channel routing is host-controlled.** OpenClaw routes inbound replies back to their originating channel — the model never picks a channel. If the user reports "agent replied on web instead of Discord," go to **Channel Routing** section, not "the model got confused."
> 11. **Resolve `<WORKSPACE>` before any file operation.** Run the **Workspace Detection** step ONCE at session start and remember the result. Substitute that path everywhere this skill writes `<WORKSPACE>`. Never write to a literal `<WORKSPACE>` — that's a placeholder, not the actual path.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill asks for personal details such as name, nickname, role, and goals, then writes them into persistent files. This is not inherently malicious, but it does create a privacy exposure because identifying information is stored long-term and may later be read automatically by the agent.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.