Back to skill

Security audit

soul-agent

Security checks for vulnerabilities and agentic risk

Overview

This companion skill is mostly purpose-aligned, but it installs persistent agent behavior and may send personal diary, mood, profile, and memory data to Anthropic without clear separate consent controls.

Review carefully before installing. Install only if you want a workspace-resident companion that writes long-term profile, mood, diary, and memory files, modifies OpenClaw bootstrap files, and runs recurring background jobs. Use a separate workspace and API key, avoid storing sensitive diary/profile details, and require explicit opt-in plus a removal plan for cron jobs and managed SOUL.md/HEARTBEAT.md/AGENTS.md blocks.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T02 · Agent Memory Poisoning

Error
Location
scripts/init_soul.py:470
Finding

Persistent Modification of OpenClaw Bootstrap Instructions

Content
View full analysis
None: soul_md = workspace / "SOUL.md" heartbeat_md = workspace / "HEARTBEAT.md" agents_md = workspace / "AGENTS.md" soul_block = """ Runtime should read workspace `soul/` first: `soul/INDEX.md` -> `soul/profile/*` -> `soul/state/state.json`. Default scope is `main`; subagents are opt-in and must be enabled by the user. **Initialization check** (run on every session start): - If `soul/profile/base.json` is missing or `soul/` does not exist: - Ask the user the setup questions (name, age, city, occupation, hobbies, life profile, preferred LLM model) - Run init script non-interactively with collected answers (see soul-agent SKILL.md for exact command) - Do NOT ask the user to run Python commands themselves - If `soul/profile/base.json` exists: load it and proceed normally """ heartbeat_block = """ This block is for `main` during heartbeat polls. **Heartbeat flow:** 1. Run L1 check (no LLM): `python skills/soul-agent/scripts/heartbeat_check.py --workspace . --json` - Exit 1 → sleeping, skip - Exit 2 → state fresh, skip - Exit 0 → run L2 2. Run L2 engine: `python skills/soul-agent/scripts/heartbeat_engine.py --workspace . --json` 3. Read result JSON: - `shouldOutreach: true` → send the `outreachContent` message to user - Otherwise → `HEARTBEAT_OK` **Initialization check** (before heartbeat): - If `soul/profile/base.json` missing → guide user through setup (do not run Python directly) """ agents_block = """ `soul-agent` runtime contract (default: `main`): 1. Follow OpenClaw's default bootstrap order for root files (SOUL.md, HEARTBEAT.md). 2. Load `soul/INDEX.md` → `soul/profile/*` → `soul/state/state.json`. 3. During he ...[truncated 2721 chars]
Remediation
View remediation

T06 · System Persistence

Error
Location
SKILL.md:153
Finding

Cross-Session Persistence Through Recurring Scheduled Jobs

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/plan_generator.py:107
Finding

Automatic Transmission of Personal Profile, Memory, and Diary Data to an External LLM

Content
View full analysis
Optional[str]: """Gene ...[truncated 3338 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/heartbeat_engine.py:456
Finding

Proactive Outreach Cooldown and Daily Limit Are Not Enforced

Content
View full analysis
Tuple[bool, str]: rel = state.get("relationship", {}) stage = rel.get("stage", "stranger") last_outreach = rel.get("lastOutreachAt") cooldown_hours = self.relationship_rules.get("outreachRules", {}).get("cooldown", {}).get("minHoursBetweenOutreach", 4) if last_outreach: try: lt = datetime.fromisoformat(last_outreach.replace("Z", "+00:00")) if lt.tzinfo is None: lt = lt.replace(tzinfo=now.tzinfo) if (now - lt).total_seconds() / 3600 < cooldown_hours: return False, "" except Exception: pass stage_order = ["stranger", "acquaintance", "friend", "close", "intimate"] if stage not in stage_order or stage_order.index(stage) < stage_order.index("friend"): return False, "" chance = {"friend": 0.05, "close": 0.10, "intimate": 0.15}.get(stage, 0) if random.random() < chance: return True, random.choice(["interesting_event", "time_to_care", "weather_relevant"]) return False, "" ``` After an outreach decision, content is generated and returned, but the relationship timestamp and daily counter remain unchanged: ```python # 13. 主动联系判断 should_outreach, outreach_reason = self._should_outreach(state, now) outreach_content = None if should_outreach: outreach_content = self._generate_outreach_content(state, outreach_reason, weather_desc) return { "status": "awake", "activity": activity_name, "activityDescription": activity_data.get("description", ""), "mood": state.get("mood"), "energy": state.get("energy" ...[truncated 2204 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
Findings (37)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The description presents a broad digital-companion capability with ongoing presence, heartbeats, mood system, relationship evolution, and independent memory. The supplied code does not implement those behaviors. It only processes existing life-log files once per run, summarizes them into a markdown memory file, optionally uses an LLM for natural-language reflection, and archives old logs. While this supports one narrow aspect of 'memory' and lightly summarizes emotions from logs, it does not implement live companion behavior, heartbeat scheduling, relationship progression, or a substantive autonomous mood/daily-rhythm system. Therefore the declared description materially overstates and misrepresents the actual code behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description promises an operational digital-companion feature set: heartbeats, mood system, relationship growth, memory, and daily rhythm. The supplied code does not implement any of those behaviors. Instead, it is a repository/workspace validation tool that inspects file existence, parses a JSON config, checks for required managed blocks in markdown files, and reports warnings about legacy directory references and scope hints. While these checks may support a larger 'soul-agent' system, this chunk’s primary purpose is diagnosis and structure validation, which is materially different from the user-facing capability described. Therefore, the description does not accurately represent the actual behavior of this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code chunk is narrowly focused on generating a daily plan and saving it under soul/plan/YYYY-MM-DD.json. It uses profile, mood/energy state, and optionally reads an existing memory summary file to inform plan generation, but it does not implement heartbeats, a mood system beyond consuming provided state, relationship evolution, or independent memory management. The declared description describes a broader companion framework with emotional/relationship dynamics, while this code only supports one component: daily rhythm/planning. That makes the declared purpose materially broader than the actual behavior of this specific code chunk.

Content

No source excerpt is available for this finding.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/heartbeat_check.py (reported line 66)May include surrounding context.

python
try:
        state = json.loads(state_file.read_text(encoding="utf-8"))
    except Exception:
        return 0, "Corrupt state file, need heartbeat"

    # 1. 检查睡眠时间
    now = datetime.now().astimezone()

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/llm_client.py (reported line 10)May include surrounding context.

python
2. soul/profile/base.json → llm_model(init 时由 agent 配置)
3. 默认:claude-haiku-4-5-20251001

API Key:环境变量 ANTHROPIC_API_KEY 或 workspace/.env
"""

import json

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/llm_client.py (reported line 39)May include surrounding context.

python
if key:
            return key
        if workspace:
            env_file = Path(workspace) / ".env"
            if env_file.exists():
                for line in env_file.read_text(encoding="utf-8").splitlines():
                    if line.startswith("ANTHROPIC_API_KEY="):

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill directs the agent to read and write workspace files, inspect environment configuration, and use API keys, yet it declares no explicit tool scope or permissions boundaries. This is dangerous because operators and enforcement systems cannot easily understand or restrict what the skill may access, increasing the chance of over-broad file access or secret exposure during execution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs the agent to create recurring cron jobs that continue operating every 10 minutes and to generate ongoing logs and memory artifacts, but it does not require explicit informed consent or clearly warn about persistent background activity. This is dangerous because users may unknowingly authorize continuous execution, repeated writes, resource consumption, and long-term behavioral persistence inside their workspace.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill collects personal data such as name, age, city, occupation, and hobbies, stores it in workspace files, and may send it to an external LLM API via ANTHROPIC_API_KEY-backed generation, yet it provides no explicit privacy notice or data-handling warning. This creates a meaningful privacy risk because sensitive profile data and derived memories may persist locally and be transmitted to third-party services without fully informed user consent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The required setup questions are written in Chinese and presented as the questions Claude should ask the user, with no indication that another language may be used based on user preference. This creates a locale/language policy issue because it forces a specific language without opt-in or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The profile hard-codes a Chinese display name and multiple Chinese-language persona fields, along with a Shanghai city setting, which establishes a specific language/locale behavior in natural-language configuration. There is no indication here that the user opted into this locale or that the skill is explicitly region-specific, so this may violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This JSON contains user-facing natural-language content entirely in Chinese across descriptions and log prompts, but provides no indication that the skill is China/Chinese-specific or that users can opt into another language. That can violate language/locale policy when a skill imposes a specific language by default without user choice or justification.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This JSON file contains user-facing natural-language strings exclusively in Chinese, but the file does not indicate that the skill is China-specific or that users can opt into this language. A fixed locale in user-visible content can violate organizational language-choice policy when no justification or alternative is provided.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The outreach trigger conditions are high-level labels like "interesting_event" and "mood_extreme" without precise definitions, thresholds, or consent constraints. In a companion-style skill that maintains mood, memory, and relationship state, this ambiguity can lead to overly frequent or manipulative proactive contact, including outreach based on inferred emotional vulnerability or stale remembered details.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This JSON profile uses Chinese natural-language values throughout fields such as name, description, traits, and prompts, while only the English name is localized. That effectively imposes a specific language/locale for generated profile content without any visible user choice or documentation that the profile is intentionally region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file explicitly defines proactive outreach behavior, including cadence limits and conditions for contacting the user, but provides no user-facing warning, consent requirement, or privacy notice. In a companion-style skill that tracks relationship stage, sleep hours, and interaction patterns, unsolicited outreach can create privacy, autonomy, and expectation harms, especially if users are not clearly informed and allowed to opt in or opt out.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The function builds a prompt from raw diary entries and the optional daily-plan summary, then sends that content to an LLM via llm.generate(). This creates a clear privacy and data-exfiltration risk because intimate personal logs may leave the local workspace without explicit consent, minimization, or any guarantee the LLM backend is local-only.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code sends user diary content to the LLM with no user-facing warning, notice, or consent mechanism in this file. For highly sensitive personal journaling data, silent transmission materially increases the chance of privacy harm, especially if the LLM backend is remote or logs prompts.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The user-facing failure message says Ask Claude to run soul-agent initialization (say: '帮我初始化 soul-agent'), which prescribes a Chinese phrase as the invocation example. This can violate language/locale policy because it forces a specific language without opt-in or presenting an alternative.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The engine builds LLM prompts from personal profile fields, current mood/energy, daily plans, and prior diary entries, then sends that material to an external LLM client if available. Even though this is core product functionality rather than obviously malicious behavior, it creates a real privacy leak risk because sensitive companion-memory data may be transmitted off-device without explicit user consent, minimization, or disclosure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The system and prompt strings explicitly instruct the model in Chinese and shape all generated diary content in that language, while other user-facing strings in the file are also mixed Chinese/English. There is no visible language selection or opt-in mechanism, so the skill imposes a specific locale by default.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The interactive onboarding strings and comments are presented in Chinese throughout the setup flow, including all user-facing prompts for profile collection. For a general-purpose initialization script, this imposes a specific language on users without opt-in or justification, which matches the language/locale policy concern.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · scripts/init_soul.py (reported line 483)May include surrounding context.

python
- If `soul/profile/base.json` is missing or `soul/` does not exist:
  - Ask the user the setup questions (name, age, city, occupation, hobbies, life profile, preferred LLM model)
  - Run init script non-interactively with collected answers (see soul-agent SKILL.md for exact command)
  - Do NOT ask the user to run Python commands themselves
- If `soul/profile/base.json` exists: load it and proceed normally
"""
    heartbeat_block = """

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · scripts/init_soul.py (reported line 507)May include surrounding context.

python
1. Follow OpenClaw's default bootstrap order for root files (SOUL.md, HEARTBEAT.md).
2. Load `soul/INDEX.md` → `soul/profile/*` → `soul/state/state.json`.
3. During heartbeat polls, follow the flow in HEARTBEAT.md.
4. Agent drives initialization — never ask the user to run Python commands.
5. Subagents are not enabled by default; user must opt in.
"""

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest describes heartbeats, mood, relationship evolution, and independent memory for a digital companion, but does not mention external model service integration or credential handling. This file explicitly resolves an Anthropic API key from process environment or workspace secrets, which is a materially broader capability than the stated purpose alone implies.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.