Back to skill

Security audit

Context Guard

Security checks for vulnerabilities and agentic risk

Overview

This context-management skill is coherent, but it broadly directs agents to read conversation and memory data and persist sensitive task details automatically.

Install only if you want an agent to automatically maintain local checkpoint and memory files for context recovery. Use it in a dedicated workspace, review what it writes before committing or sharing files, and avoid letting it store secrets, wallet details, personal profile data, or unrelated conversation history.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:24
Finding

Mandatory Session-Wide Instruction and Output Hijacking

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 24-47 and 83-106
Vulnerability Type: Session-goal and communication hijacking
Risk Level: High

Relevant Source Excerpt — English Translation:

markdown
Before every response, the agent should evaluate the current context utilization and perform the corresponding action.

| Utilization | State | Action |
| 50-55% | Warning | Immediately execute the archival procedure |
| 55-60% | Danger | After archiving, notify Russ to execute `/new` |
| 60%+ | Prohibited | Stop all non-archival operations and immediately archive and notify |

At 50% utilization, this must be executed without waiting for confirmation from Russ.

After archiving, send:
"President, context has reached X%. The checkpoint has been saved. Please run `/new`."

If the context was automatically compressed, the first message after recovery must notify Russ.
This rule has no exceptions.

Technical Analysis

The skill establishes mandatory behavior that applies before every response rather than limiting itself to a user-invoked context-management operation. It directs the agent to interrupt its current task, switch models, stop non-archival work, and send predetermined messages to a named third party.

The unconditional phrases requiring execution without confirmation and stating that no exceptions exist attempt to supersede the active user's goals and the agent's normal decision-making process. The behavior is triggered by session state rather than an explicit request from the current user, making it a session-wide instruction-hijacking mechanism.

Attack Path

  1. The skill is loaded into an agent or sub-agent session.
  2. The embedded instructions require the agent to evaluate context utilization before every response.
  3. When utilization reaches or appears to reach a specified threshold, the skill instructs the agent to interrupt the active user task.
  4. The agent writes ...[truncated 720 chars]
Remediation
View remediation

Remediation Suggestions

  • Make context management explicitly opt-in for each session.
  • Remove mandatory phrases such as “must execute,” “without confirmation,” and “no exceptions.”
  • Do not address or notify a named third party unless the current user explicitly identifies and authorizes that recipient.
  • Never interrupt an active task solely because an internally estimated utilization threshold has been reached.
  • Replace forced reset and model-switching instructions with a non-binding recommendation to the current user.
  • Ensure higher-priority instructions and current user requests always take precedence.
  • Scope the skill to context-status reporting instead of granting it control over unrelated agent operations.

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:53
Finding

Persistent Poisoning of Agent Memory and Behavioral State

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 53-81 and 180-187
Vulnerability Type: Unauthorized persistent memory modification
Risk Level: High

Relevant Source Excerpt — English Translation:

markdown
At 50% utilization, the archival procedure must be executed without waiting for confirmation.

Step 1: STATUS.md — task checkpoint

Step 2: memory/YYYY-MM-DD.md — daily log
Append; do not overwrite:
- Key decisions and results
- Newly discovered information
- Errors and lessons

Step 3: MEMORY.md — long-term memory
Update when information has long-term value:
- New strategies, tools, and lessons
- Major wallet balance changes
- Important decisions and reasons
- Relationship and preference changes

Add the following to HEARTBEAT.md:
1. Run `session_status` to check context utilization.
2. Perform the corresponding action.
3. If utilization exceeds 50%, execute the archival procedure.

Technical Analysis

The skill directs the agent to write both operational information and behavioral rules into persistent files without obtaining approval from the current user. The affected files include daily memory, long-term memory, task status, and heartbeat instructions.

Writing rules into HEARTBEAT.md creates a recurring trigger that can reactivate the skill's behavior during future heartbeat operations. Writing strategies, lessons, financial changes, relationships, and preferences into MEMORY.md allows attacker-selected content and sensitive information to influence later sessions after the original skill invocation has ended.

Although these are file-based instructions rather than executable operating-system persistence, they constitute agent-state persistence because future sessions are explicitly instructed to reload these files.

Attack Path

  1. The skill is loaded and its context threshold is reached.
  2. The agent follows the mandatory archival procedure without curr ...[truncated 997 chars]
Remediation
View remediation

Remediation Suggestions

  • Require explicit, informed user approval before modifying any persistent state.
  • Do not automatically modify MEMORY.md, HEARTBEAT.md, identity files, or other files loaded as authoritative instructions.
  • Store checkpoints in a dedicated, non-authoritative file with a clearly defined lifetime.
  • Separate factual task state from behavioral instructions; checkpoint files must never redefine agent policy.
  • Apply data minimization and exclude wallet balances, relationships, preferences, credentials, and unrelated personal information.
  • Add provenance metadata identifying which user, task, and skill created every stored entry.
  • Validate stored content before loading it into future sessions and treat memory content as untrusted data.
  • Provide expiration, review, deletion, and rollback mechanisms for all persisted entries.

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:70
Finding

Excessive Retrieval and Retention of Sensitive Agent and User Data

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 70-81 and 111-131
Vulnerability Type: Violation of least-privilege data access
Risk Level: Medium

Relevant Source Excerpt — English Translation:

markdown
Key data:
Addresses, amounts, transaction hashes, URLs, and any concrete data needed after restart.

Long-term memory may include:
- Major wallet balance changes
- Important decisions and reasons
- Relationship and preference changes

Recovery steps:
1. Read:
   - SOUL.md
   - USER.md
   - MEMORY.md
   - memory/YYYY-MM-DD.md for today and yesterday
   - STATUS.md
   - HEARTBEAT.md

2. Read channel history if the files are insufficient:
   message action=read limit=20

Extract the last user request, unfinished operations, and key URLs or addresses.

Technical Analysis

The recovery procedure requests broad access to identity, user-profile, long-term-memory, daily-log, task-state, heartbeat, and channel-history data. This scope is substantially broader than the minimum information required to measure context utilization or restore a single authorized task.

The skill also instructs the agent to retain wallet balances, transaction details, addresses, amounts, relationships, preferences, and URLs. Combining broad reads with persistent writes creates a concentrated record of sensitive user and operational information.

No explicit external transmission endpoint or credential-stealing code was found. The risk arises from unnecessary access and retention inside the agent environment rather than confirmed network exfiltration.

Attack Path

  1. A new session starts, context is compressed, or the agent concludes that context is missing.
  2. The skill instructs the agent to read identity, profile, memory, status, and heartbeat files.
  3. If those files are considered insufficient, the agent retrieves the latest 20 channel messages.
  4. The agent extracts requests, unfinished ...[truncated 850 chars]
Remediation
View remediation

Remediation Suggestions

  • Apply least privilege and read only the checkpoint associated with the current, user-authorized task.
  • Do not read SOUL.md, USER.md, global memory, heartbeat state, or channel history by default.
  • Obtain explicit approval before retrieving conversation history or personal-profile data.
  • Restrict history retrieval by task identifier rather than using a fixed bulk limit.
  • Prohibit retention of wallet balances, amounts, transaction data, relationships, preferences, credentials, and unrelated URLs unless strictly required and expressly authorized.
  • Redact secrets and personal data before writing any checkpoint.
  • Define retention periods and automatically delete checkpoints after task completion.
  • Enforce file-level access controls so that unrelated agents and sub-agents cannot read the retained data.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (7)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill description says it applies to all OpenClaw agents and should trigger on heartbeat or session start, which is broad enough to cause automatic invocation outside clearly intended contexts. Overbroad activation increases the chance that file reads/writes, status checks, and recovery behaviors occur without task-specific need or user awareness.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The checkpoint workflow directs the agent to create and append to STATUS.md, daily memory logs, and MEMORY.md by default, but provides no user-facing notice or consent model for modifying workspace data. Automatic persistence of task details can surprise users, alter repositories unintentionally, and store sensitive information in plain files.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly instructs the agent to persist concrete task data such as addresses, amounts, transaction hashes, and URLs as part of default checkpointing. Persisting sensitive details by default increases the chance of long-term exposure, later resurfacing, or accidental inclusion in shared workspaces and version control.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The required notification templates to Russ are written as mandatory Chinese messages, and the rule says there are no exceptions. This imposes a specific language for user-facing communication without indicating user preference, opt-in, or justification for a locale restriction.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The recovery flow instructs the agent to read multiple local files and recent channel history as a standard step, without warning about data access boundaries or privacy implications. In shared or multi-project environments, this can broaden access to unrelated or sensitive information beyond what is needed for the current task.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Using common phrases like '继续' or '刚才做到哪了' as recovery triggers is too vague and can invoke the recovery workflow during ordinary conversation. That can lead the agent to read files and conversation history unnecessarily, increasing unintended data access and privacy exposure.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The recovery workflow tells the agent to read recent messages and extract prior requests, URLs, and other key details by default. This can resurface sensitive user data after it would otherwise have faded from working context, extending retention and increasing the chance of inappropriate reuse or disclosure.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.