Back to skill

Security audit

Value Aware Guard

Security checks for vulnerabilities and agentic risk

Overview

This skill is purpose-aligned as a value and boundary guard, but it persists sensitive behavioral/profile data and supports silent monitoring/intervention with insufficient user-facing consent and controls.

Install only if you are comfortable with a Chinese-language value-monitoring tool that keeps local records about your values, boundaries, behavioral drift, interventions, and signals. Review or disable persistence before use, avoid feeding it sensitive health, financial, relationship, identity, or work-profile data unless you intentionally want that stored, and expect to inspect/delete files under ~/.openclaw/workspace/.soul manually because clear consent, retention, and deletion controls are missing.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/guard.js:200
Finding

Plaintext Persistent Storage of Sensitive User Profile and Intervention Data

Content
View full analysis

Vulnerability Details

File Location: scripts/guard.js:20-23, 148-204, 208-220
Vulnerability Type: Sensitive data stored persistently without encryption, restrictive permissions, or explicit consent
Risk Level: Medium

Vulnerable Code

javascript
const WORKSPACE = process.env.HOME ? join(process.env.HOME, '.openclaw/workspace') : '/tmp/.openclaw/workspace';
const SOUL_DIR = join(WORKSPACE, '.soul');
const GUARD_STATE_PATH = join(SOUL_DIR, 'guard-state.json');
const VALUES_PATH = join(SOUL_DIR, 'user-values.json');
const INTERVENTIONS_PATH = join(SOUL_DIR, 'interventions.jsonl');
javascript
function saveUserValues(values) {
  ensureDir(SOUL_DIR);
  values.updated_at = new Date().toISOString();
  writeFileSync(VALUES_PATH, JSON.stringify(values, null, 2), 'utf-8');
}
javascript
function recordIntervention(intervention) {
  ensureDir(SOUL_DIR);

  const interventionRecord = {
    ...intervention,
    id: `int_${Date.now()}_${Math.random().toString(36).slice(2, 8)}`,
    timestamp: new Date().toISOString()
  };

  try {
    const line = JSON.stringify(interventionRecord);
    appendFileSync(INTERVENTIONS_PATH, line + '\n', 'utf-8');
    return interventionRecord.id;
  } catch (error) {
    console.error('记录干预失败:', error.message);
    return null;
  }
}

When user-values.json does not exist, loadUserValues() also automatically reads core values and work principles from .soul/user-profile.json and passes them to saveUserValues(). This creates another persistent copy without an explicit user-consent check.

Technical Analysis

The skill handles potentially sensitive behavioral and profile information, including core values, work principles, personal boundaries, value-drift assessments, and intervention messages. It persists this information as ordinary JSON and JSONL files at predictable paths under the user's workspace.

The writes do ...[truncated 2631 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require informed consent before persistence

    • Ask the user before importing profile values into a new persistent file.
    • Clearly disclose which fields will be stored, why they are needed, and how long they will be retained.
    • Support a memory-only mode and make sensitive-data persistence opt-in.
  2. Enforce restrictive filesystem permissions

    • Create .soul with mode 0700.
    • Create sensitive files with mode 0600.
    • Verify and repair permissions on existing files before reading or writing them.
    javascript
    mkdirSync(SOUL_DIR, { recursive: true, mode: 0o700 });
    writeFileSync(VALUES_PATH, serializedValues, {
      encoding: 'utf8',
      mode: 0o600
    });
    
  3. Encrypt sensitive records at rest

    • Use authenticated encryption such as AES-256-GCM.
    • Store encryption keys in the operating system's credential store rather than beside the encrypted files.
    • Implement key rotation and fail closed if secure key storage is unavailable.
  4. Minimize and redact persisted content

    • Do not persist complete generated intervention messages by default.
    • Store only the minimum structured metadata required for rate limiting and effectiveness analysis.
    • Remove or tokenize sensitive topics and free-form text before writing records.
    • Apply the documented privacy classification rules before any record is persisted.
  5. Implement retention and deletion controls

    • Establish short, configurable retention periods.
    • Rotate and delete old intervention records.
    • Provide a user-accessible command to inspect, export, and permanently delete stored data.
  6. Harden file handling

    • Reject symbolic links and verify that resolved paths remain inside the expected workspace.
    • Use atomic writes for JSON state files.
    • Avoid the shared /tmp fallback unless a private directory with verified ownership and mo ...[truncated 325 chars]
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (36)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The documented purpose is a benign 'value guard', but the analyzed behavior reportedly includes undeclared local filesystem read/write operations and persistent signal-queue management that are not described to users. This mismatch is dangerous because it can conceal stateful behavior, data collection, or unintended automation behind an innocuous description, undermining informed consent and security review.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 12)May include surrounding context.

md
node scripts/guard.js <命令> [选项]

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 17)May include surrounding context.

md
node scripts/guard.js <命令> [选项]

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 31)May include surrounding context.

md
node scripts/guard.js <命令> [选项]

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 34)May include surrounding context.

md
node scripts/guard.js <命令> [选项]

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 37)May include surrounding context.

md
node scripts/guard.js <命令> [选项]

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document explicitly defines '静默记录' and states the user should be completely unaware ('用户完全无感知') while collecting timestamped behavioral and contextual data. Covert monitoring of sensitive behavior without notice, consent, or clear retention limits creates a serious privacy violation and can enable profiling, manipulation, or disclosure harms if the data is misused or breached.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The intervention ladder includes 'L4强制干预' and other escalating actions without a clear user-consent framework, emergency limitation, or due-process safeguards. This is dangerous because a system that profiles user behavior and then escalates to coercive intervention can override autonomy, pressure users, and be misapplied in non-safety-critical situations.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
82% confidence
Finding

The skill advertises executable CLI behavior and integration points but does not declare any tool scope or permissions, while static analysis indicates environment access capability. Missing scope declarations weaken least-privilege controls and make it harder for users or orchestrators to understand what the skill can access before invocation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger conditions include broad concepts like value drift, boundary violations, low energy alerts, and user pattern discovery without precise definitions. Ambiguous triggers can cause the skill to run unexpectedly or too often, which is especially risky for a component that claims monitoring and intervention behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The documentation mentions integration with a signal system for automatic monitoring and intervention but does not prominently warn users that the skill may act proactively. Undisclosed autonomous behavior can surprise operators, increase the chance of unintended actions, and complicate trust and oversight.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The package description is written entirely in Chinese, which implies a fixed language choice in the skill metadata without any indication that users can opt into another language. The file does not document a locale-specific purpose or provide language flexibility, which may conflict with language-choice policy expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The file title and all user-facing examples and rule text are written exclusively in Chinese, and there is no indication that the skill supports other languages or allows user opt-in to a preferred language. Under the policy, forcing a specific language without user choice can be a natural-language policy violation unless clearly documented as region-specific.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The document authorizes recording sensitive topics, offering encryption, and periodic cleanup, which implies persistence and handling of private user data beyond a narrow runtime guard role. Expanding a guard skill into data retention increases privacy and surveillance risk, especially because the file also discusses health, financial, relationship, and identity-related topics.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill describes long-term learning of work hours, recovery patterns, privacy comfort levels, decision style, and reminder preferences, which amounts to behavioral profiling. In the context of a guard system, this creates a significant autonomy and privacy risk because the system can infer and retain intimate traits not necessary for immediate boundary enforcement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The entire skill reference is written only in Chinese and does not indicate that language selection is optional or limited to a Chinese-language deployment. Under the policy, forcing a specific language without user opt-in or documented justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill goes beyond simple logging and directs the system to accumulate deviations, analyze triggering situations, and identify recurring behavior patterns for personalization, again without an accompanying privacy disclosure or consent flow. Behavioral profiling increases sensitivity of the dataset and can infer habits, vulnerabilities, or mental-state patterns, making misuse or unauthorized access more harmful.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The entire skill file is written only in Chinese and provides no indication that users may choose another language or locale. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The document defines a system for monitoring user values, behaviors, and inconsistencies, but it does not include explicit notice, consent, or safeguards for sensitive profiling. This is dangerous because it normalizes ongoing behavioral analysis and intervention logic over intimate personal patterns without making user awareness and authorization a hard requirement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The examples explicitly reference sensitive collection sources such as conversation, sensors, and external APIs for behavior detection, yet no warning, minimization, or consent gating is specified. This increases the risk of invasive surveillance, overcollection, and misuse of highly sensitive personal data derived from multiple channels.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script persists state, values, and intervention logs under the user's home directory automatically, with no explicit notice or consent flow. In a skill centered on monitoring values, boundaries, energy, and interventions, these files can contain sensitive behavioral metadata and create a durable privacy footprint that the user may not expect.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script reads a separate user-profile file and derives value data from it automatically, which expands data collection beyond the immediate guard state/config role into profile inspection. In this skill context, that means a behavioral monitoring tool can silently ingest broader personal profile data from the user's home workspace, increasing privacy risk and the chance of unintended secondary use of sensitive information.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The intervention logger appends structured records containing behavioral assessments, timestamps, and intervention metadata to a JSONL file without explicit disclosure. Because this skill assesses value drift and boundary issues, the log may reveal intimate patterns about user behavior or mental state, and long-term accumulation raises privacy and profiling concerns if another local process or person can access the file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

Nearly all natural-language strings in the script are Chinese, including status messages and next-step guidance. Under the policy, a skill should not force a specific language unless it offers opt-in or clearly documents a justified locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This file contains natural-language comments and console output such as '加载配置失败,使用默认配置' and '当前配置' exclusively in Chinese. For a general-purpose skill, hard-coding a specific language without user opt-in can violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.