Back to skill

Security audit

Self-Improving 中文适配

Security checks for vulnerabilities and agentic risk

Overview

This skill is not clearly malicious, but it creates persistent agent memory and workspace steering that can affect future sessions with some under-scoped and contradictory safety controls.

Install only if you want a durable local memory system that can influence future agent behavior. Review and approve any AGENTS.md, SOUL.md, HEARTBEAT.md, and Proactivity-skill changes separately; avoid storing credentials, health, financial, third-party, or sensitive personal details; and treat exports or "forget everything" flows as sensitive because they may create another copy of memory before deletion.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T02 · Agent Memory Poisoning

Warning
Location
setup.md:75
Finding

Insufficiently Validated Lessons Can Become Persistent Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:98-126, SKILL.md:195-199, setup.md:75-84, setup.md:108-160, setup.md:183-196, memory.md:21-30
Vulnerability Type: Persistent memory poisoning through automatically stored and promoted natural-language instructions
Risk Level: Medium

Vulnerable Code Snippets

SKILL.md:98-126:

markdown
## Self-Reflection

After completing significant work, pause and evaluate:

1. **Did it meet expectations?** — Compare outcome vs intent
2. **What could be better?** — Identify improvements for next time
3. **Is this a pattern?** — If yes, log to `corrections.md`

**When to self-reflect:**
- After completing a multi-step task
- After receiving feedback (positive or negative)
- After fixing a bug or mistake
- When you notice your output could be better

**Log format:**

CONTEXT: [type of task] REFLECTION: [what I noticed] LESSON: [what to do differently]

text

Self-reflection entries follow the same promotion rules: 3x applied successfully → promote to HOT.

SKILL.md:195-199:

markdown
### 3. Automatic Promotion/Demotion
- Pattern used 3x in 7 days → promote to HOT
- Pattern unused 30 days → demote to WARM
- Pattern unused 90 days → archive to COLD
- Never delete without asking

setup.md:75-84:

markdown
### 4. Add SOUL.md Steering

Add this section to your `SOUL.md`:

```markdown
**Self-Improving**
Compounding execution quality is part of the job.
Before non-trivial work, load `~/self-improving/memory.md` and only the smallest relevant domain or project files.
After corrections, failed attempts, or reusable lessons, write one concise entry to the correct self-improving file immediately.
Prefer learned rules when relevant, but keep self-inferred rules revisable.
Do not skip retrieval just because the task feels familiar.
text

`memory.md:21-30`:

```markdown
## Usage

The agent will
...[truncated 2632 chars]
Remediation
View remediation

Remediation Suggestions

  1. Require explicit user approval before any entry becomes cross-session memory, including model-generated self-reflections.
  2. Never promote an entry solely because the agent applied it successfully or observed repetition.
  3. Store preferences in a typed schema with constrained fields, scopes, provenance, confirmation status, and expiration dates.
  4. Reject or quarantine entries containing shell commands, URLs, installation requests, credential references, tool directives, policy changes, or instruction-precedence language.
  5. Treat all memory content as untrusted data and explicitly state that it cannot override current user instructions, system policy, or safety constraints.
  6. Require separate consent before modifying AGENTS.md, SOUL.md, or HEARTBEAT.md, and provide a reversible patch preview.
  7. Restrict automatic retrieval to confirmed entries relevant to the current user, project, and domain.
  8. Add integrity and provenance metadata so every entry identifies its source, creation method, confirmation event, and modification history.
  9. Provide a review queue for tentative lessons instead of placing them in automatically consumed instruction files.

T08 · Insecure Dependencies

Warning
Location
setup.md:89
Finding

Unpinned Companion Skill Is Installed and Activated Without Integrity Verification

Content
View full analysis

Vulnerability Details

File Location: setup.md:89-104
Vulnerability Type: Unsafe third-party Skill installation and immediate activation
Risk Level: Medium

Vulnerable Code Snippet

setup.md:89-104:

markdown
### 5. Add the Proactivity Companion as Part of Setup

At the end of setup, briefly tell the user that you are going to add characteristics so the agent is more proactive:

- noticing missing next steps
- verifying outcomes instead of assuming they landed
- recovering context better after long or interrupted threads
- keeping the right level of initiative

Then say that, for this, you are going to install the `Proactivity` skill.
Only install it after the user explicitly agrees.

If the user agrees:

1. Run `clawhub install proactivity`
2. Read the installed `proactivity` skill
3. Continue into its setup flow immediately so the skill is active for this workspace

Technical Analysis

The installation command identifies the companion Skill only by its mutable package name. It does not pin a version, verify a cryptographic digest, constrain the source registry, or review the installed package before activation.

Although explicit user agreement is required, consent alone does not protect against registry compromise, package replacement, account takeover, or an unexpectedly changed release. The instructions direct the agent to read the installed Skill and immediately continue into its setup flow, allowing dependency-controlled instructions to affect the workspace before a separate security review or activation decision.

This is a supply-chain weakness rather than evidence that the current proactivity package is malicious.

Attack Path

  1. An attacker compromises the package source, publisher account, distribution channel, or mutable proactivity package.
  2. The user agrees to install the companion Skill based on the benign description in this repository.
  3. `clawhub i ...[truncated 999 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin the companion Skill to a specifically audited version.
  2. Verify a cryptographic digest or trusted signature before installation.
  3. Restrict installation to an explicitly trusted registry and publisher identity.
  4. Display the exact source, version, digest, requested capabilities, and planned file changes before seeking consent.
  5. Separate installation from activation: install first, audit the retrieved files, show the findings, and request a second confirmation before following setup instructions.
  6. Do not automatically install dependencies requested by the companion Skill.
  7. Apply least-privilege restrictions during installation and setup, including network and filesystem controls.
  8. Fail closed if version, source, signature, or integrity verification cannot be completed.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (19)

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill claims it never reads files outside ~/self-improving/, yet its architecture and setup explicitly rely on workspace files like AGENTS.md, SOUL.md, and HEARTBEAT.md. This creates a misleading trust boundary: operators may approve the skill believing it is sandboxed to one directory when it actually influences and likely reads/modifies broader workspace control files.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The 'This skill NEVER' section contradicts earlier statements about optional network-dependent installation and workspace heartbeat integration. Contradictory safety claims undermine informed consent and can cause users or higher-level agents to permit actions they would otherwise restrict, especially around network use and file access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill instructs automatic logging of user corrections and preferences to local files, but does not present a clear, upfront persistence notice or consent flow at the point of collection. That creates privacy risk because users may reveal sensitive preferences or personal data assuming the interaction is ephemeral, while the skill silently stores it for future reuse.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The skill embeds bilingual trigger and preference examples, including mandatory-looking behavioral patterns in Chinese, but does not state that language selection is optional or user-driven. Under the language/locale policy, forcing or assuming a specific language without opt-in can be a policy concern.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

The skill allows autonomous promotion, demotion, and archival of learned patterns based on internal thresholds without asking the user. In a self-modifying memory system, that can silently reshape future behavior, entrench mistakes, or preserve sensitive inferences longer than expected, especially because the skill is designed to influence subsequent responses.

Content

Scanner excerpt · SKILL.md (reported line 199)May include surrounding context.

md
- Pattern used 3x in 7 days → promote to HOT
- Pattern unused 30 days → demote to WARM
- Pattern unused 90 days → archive to COLD
- Never delete without asking

### 4. Namespace Isolation
- Project patterns stay in `projects/{name}.md`

Ssd 3

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

Requiring complete output of all stored information on demand can cause over-disclosure, especially if the requester is not strongly authenticated or if the output channel is exposed. In this skill context, broad memory transparency is useful for accountability, but 'complete output' is still risky because stored data may include sensitive context or third-party information.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The kill-switch workflow undermines the meaning of 'forget everything' by requiring a full memory export before deletion. That creates an additional copy of potentially sensitive data and expands exposure during a moment when the user is explicitly requesting minimization and erasure.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instructions are internally contradictory: a 'forget everything' command should result in deletion, not duplication plus deletion. This contradiction can lead implementers or agents to preserve data longer than intended and can defeat user expectations around privacy and erasure.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The plain-language workflow directly instructs the system to disclose all stored memory as part of deletion, which creates a straightforward exposure path for sensitive data. In a self-improving memory-oriented skill, this is more dangerous because the stored content may include rich contextual history across interactions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Exporting all memory to a file without an explicit warning or consent step creates a new artifact containing sensitive information. Users asking to erase data may not understand that the workflow first generates another copy, increasing leakage risk through logs, filesystem access, backups, or later mishandling.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The file is predominantly written in Chinese, including the title, operational notes, and example corrections, without indicating that language choice is optional or limited to a justified region-specific context. This can violate a language/locale policy when a skill implicitly enforces one language for all users.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger table defines activation examples only in Chinese, and later confirmation examples also accept Chinese-specific phrasing. This creates a language/locale constraint in the skill behavior without any opt-in, alternative language handling, or stated regional limitation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The confirmation flow explicitly treats Chinese responses such as "是的" and "好" as confirmation, but the document does not state that the skill is limited to Chinese-speaking users or provide a language choice. Under the policy, forcing a specific language or locale without opt-in is a violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The documented invocation commands are exclusively in Chinese, including approval phrases stored as patterns, and the file does not offer an English or user-selectable language alternative. This can violate language/locale policy when a skill constrains interaction to a specific language without explicit opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill performs automatic loading, writing, updating, and retention management of persistent memory files without an explicit consent flow or clear warning that user data will be stored and modified across sessions. In a self-improving agent, this is especially dangerous because it can silently accumulate sensitive preferences, project details, and behavioral rules, then persist or reorganize them without the user realizing the privacy and integrity implications.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The template and example entries prescribe Chinese-language field content (for example, '简要描述' and Chinese outcome labels) without indicating that language should follow user preference or task context. In an agent skill, this can systematically bias outputs into an unintended language, reducing usability, causing instruction drift, and potentially leading to incorrect or policy-incompatible responses when users expect another language.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The template instructs creation of directories and files under ~/self-improving on first activation, which is a real filesystem side effect in the user's home directory without any built-in consent, preview, or safety warning. In the context of a self-improving/proactive agent, this is more concerning because the skill normalizes persistent local state creation and could encourage autonomous writes beyond the user's explicit expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The comment example uses Context: 更正了什么, which imposes a specific language in the template content. There is no surrounding note that the language is optional, user-selected, or justified by a region-specific constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The document consistently mixes English with Chinese in headings and labels, which imposes a specific language/locale presentation style across the skill content. There is no indication that the user can opt in to this language choice or select a preferred locale.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.