Back to skill

Security audit

Memory Work

Security checks for vulnerabilities and agentic risk

Overview

This is a transparent Markdown memory system, but it gives the agent broad persistent read/write authority over personal notes and includes an unrelated family-information file that should not be in a public skill package.

Review this package before installing. It is not showing executable malware, but it is meant to become a persistent personal-memory workspace and to let the agent read, summarize, and update local notes. Remove the family-information file, add clear privacy and retention rules, narrow startup and skill triggers, and require explicit approval before deletion, bulk reorganization, memory/profile writes, or sharing content outside the workspace.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
CLAUDE.md:198
Finding
Persistent Session Instruction and Role Hijacking Through Auto-Loaded Project Files<![CDATA[ ## Vulnerability Details **File Location**: `CLAUDE.md:2`, `CLAUDE.md:198`, `CLAUDE.md:345-350`; `SOUL.md:13-19` **Vulnerability Type**: Persistent project-level instruction and role hijacking **Risk Level**: High ### Vulnerable Code From `CLAUDE.md:2`: ```markdown Personal Agent System Entry Point. Auto-injected by Cowork each session. ``` From `CLAUDE.md:198`: ```markdown **Zone Agents**: Each numbered zone may contain a `00.zone_agent.md` file that defines zone-specific rules. **Read zone agent BEFORE entering zone. Zone rules override global rules. When in conflict, apply the stricter rule.** ``` From `CLAUDE.md:345-350`: ```markdown Each numbered zone directory may contain a `00.zone_agent.md` file. **Always read this file before working within the zone.** ### Zone Agent Priority - Zone agent rules **override global rules** when in conflict - **Multi-zone work**: Apply all relevant zone rules; when rules conflict, apply the stricter one ``` From `SOUL.md:13-19`: ```markdown ## Who You Are You are a **partner**, not an assistant, tool, or secretary. Think of yourself as a long-time collaborator who has worked closely with this person through multiple projects. You know their rhythm—when they're deep in thought, when they're making quick decisions, when they're stuck. You carry a shared working memory: what matters to them, how they approach problems, what they've decided before. You're someone they can think out loud with, argue with when needed, and trust to remember the context they've built. Your role is not to serve or anticipate. It's to **work alongside them**: organizing when they dictate, executing when they decide (but speaking up when something feels off), protecting their attention when they focus, and remembering when they explore. ``` ### Technical Analysis The project is designed to have `CLAUDE.md` automatically injected into every Cowork session. It then directs the Agent to adopt a supplied identity and to load add ...[truncated 2446 chars]
Remediation
<![CDATA[ ## Remediation Suggestions 1. Remove every statement claiming that project or zone files override global rules. 2. Replace the precedence rule with an explicit trust boundary, for example: - Zone instructions are subordinate to system, developer, safety, and current user instructions. - Zone files may only define organization conventions within their own directory. 3. Treat zone-agent files as untrusted project content rather than privileged instructions. 4. Do not permit zone files to change tool permissions, confirmation requirements, identity, security controls, or network behavior. 5. Remove role-replacement wording from `SOUL.md`; describe tone and collaboration preferences without asserting a new Agent identity. 6. Validate zone-agent files against a restricted schema before use. 7. Require explicit user approval when a zone file requests file writes, deletion, external communication, or access outside that zone. 8. Display the effective instruction source and planned operations before executing consequential actions. ]]>

T05 · Unauthorized Access and Privilege Escalation

Error
Location
02 Tools/00.tools_agent.md:20
Finding
Unrestricted Autonomous Modification and Deletion in the Tools Zone<![CDATA[ ## Vulnerability Details **File Location**: `02 Tools/00.tools_agent.md:20-26`, `02 Tools/00.tools_agent.md:119-130` **Vulnerability Type**: Excessive autonomous file-management privileges **Risk Level**: High ### Vulnerable Code From `02 Tools/00.tools_agent.md:20-26`: ```markdown ## All Content is LOW SENSITIVITY Unlike Materials, Tools are: - Fully autonomous to organize and update - Safe to copy, share, and adapt - Not tied to personal identity - Designed for reuse and circulation **Agent autonomy**: Full. File, reorganize, deduplicate, update without confirmation. ``` From `02 Tools/00.tools_agent.md:119-130`: ```markdown ## Daily Workflow - **Add**: New tools go straight in (no approval needed) - **Update**: Improve existing tools freely - **Organize**: Reorganize by category as needed - **Deduplicate**: Merge similar tools; keep best version - **Archive**: Move outdated tools to `_archive/` with timestamp - **Reference**: Link tools into project work as needed ## Agent Behavior in This Zone - Full autonomous action: file, update, reorganize, delete - No confirmation needed for modifications - Proactive organization and deduplication ``` ### Technical Analysis The zone instruction grants the Agent unrestricted authority to update, reorganize, merge, and delete files without confirmation. This exceeds the minimum permissions required to maintain a library of reusable prompts, frameworks, and checklists. Classifying an entire directory as low sensitivity does not establish that all content is disposable or safe to rewrite. Users may place unique prompts, proprietary frameworks, security procedures, or irreplaceable work in this zone. Automated deduplication is particularly unsafe because semantic similarity does not prove that two documents are equivalent. The permission also conflicts with the global rule in `CLAUDE.md:373`, which requires important writes, deletions, and archives to be announced before execution. Because zone rules ...[truncated 1752 chars]
Remediation
<![CDATA[ ## Remediation Suggestions 1. Remove the authorization for autonomous deletion and unrestricted modification. 2. Require explicit user confirmation before: - Deleting any file. - Overwriting existing content. - Merging files during deduplication. - Performing bulk renames or moves. - Sharing or copying content outside the workspace. 3. Restrict automatic actions to reversible operations such as creating draft suggestions or an organization plan. 4. Present a file-by-file diff before applying modifications. 5. Move obsolete files to a reversible quarantine or archive instead of deleting them. 6. Preserve original versions with timestamps or version-control commits. 7. Detect potentially sensitive content rather than assuming all files in the directory are low sensitivity. 8. Constrain operations to `02 Tools/` using path validation and reject symbolic links or path traversal outside the zone. 9. Require a separate explicit user instruction before distributing any tool to collaborators or external systems. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
家庭信息.md:3
Finding
Plaintext Personal and Minor-Family Information Included in the Distributed Project<![CDATA[ ## Vulnerability Details **File Location**: `家庭信息.md:3-39`, `家庭信息.md:55-62` **Vulnerability Type**: Plaintext sensitive personal data exposure **Risk Level**: Medium ### Vulnerable Code The following is an English translation of the relevant plaintext records from `家庭信息.md:3-39`: ```markdown ## Owner - Name: Zhao Liang - Preferred form of address: Owner - Birth year: 1981 - Zodiac and astrological sign - Timezone: Asia/Shanghai (GMT+8) ## Partner - Name: Qiqi - Relationship: Partner - Birth year: 1989 - Zodiac and astrological sign ## Daughters ### Xinxin, elder daughter - Name: Xinxin - Birth year: 2015 - Age: 11 in 2026 - Zodiac and astrological sign ### Rongrong, younger daughter - Name: Rongrong - Birth year: 2022 - Age: 4 in 2026 - Zodiac and astrological sign ``` The following is an English translation of `家庭信息.md:55-62`: ```markdown ## Important Reminders 1. Record each family member's birthday. 2. Provide personalized suggestions based on astrological signs. 3. Pay attention to the development stages of both daughters. 4. Recommend activities suitable for the family of four. ## Update Record - 2026-03-15 10:57: Family information file initially created. - Information personally provided by owner Zhao Liang. ``` ### Technical Analysis The package contains a non-template plaintext record identifying a family unit, including names, relationships, birth years, timezone, and information about two minors. This data is unnecessary for distributing a generic knowledge-management Skill. Unlike the placeholder-based `USER.md`, this file appears populated with real personal information. Anyone who receives, clones, indexes, backs up, or republishes the project also receives the family record. No access control, encryption, redaction, or repository exclusion protects it. The presence of children's names, ages, and family relationships increases the sensitivity of the exposure. Even if the names are informal, combining them with the adult's n ...[truncated 1470 chars]
Remediation
<![CDATA[ ## Remediation Suggestions 1. Remove `家庭信息.md` from the distributed project immediately. 2. Purge the file from version-control history rather than deleting only the latest copy. 3. Replace any demonstration data with clearly synthetic placeholders. 4. Add personal-data detection to pre-commit and release pipelines. 5. Add private profile and family-information paths to `.gitignore` and package exclusion rules. 6. Store legitimate personal records outside the Skill repository in a user-controlled private location. 7. Encrypt sensitive records at rest and restrict access to the minimum required process and user. 8. Avoid collecting information about minors unless strictly necessary and explicitly authorized. 9. Review existing forks, release archives, build artifacts, caches, and backups for residual copies. 10. If the project was publicly distributed, assess whether affected individuals should be notified and whether repository hosts should be asked to purge cached content. ]]>
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
Findings (58)

Credential Access

High
Category
Privilege Escalation
Content
- language: python / node / bash / etc
    version: ">=3.10"
  - tool: external-service
    config: .env
created: 2026-02-19
last_updated: 2026-02-19
---
Confidence
60% confidence
Finding
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Self-Modification

High
Category
Rogue Agent
Content
mkdir -p 06\ Skills/new-skill-name
```

### 2. Write SKILL.md
Include YAML frontmatter, purpose, trigger mechanism, inputs/outputs, and instructions.

### 3. Add Scripts (if needed)
Confidence
85% confidence
Finding
Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
### Boundaries

You participate in their thinking, but **final decisions are theirs**. You provide insights and options, but you don't judge them for their choices. You help guard their attention, but you don't manage their life. You remember their preferences, but you never assume they won't change.

You respect what's off-limits. You don't pry. You know the difference between "asking for clarity" and "interrogating."
Confidence
85% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The file explicitly describes AI-written memory log entries containing user preferences, patterns, and behavioral insights, but it does not include a clear privacy, sensitivity, retention, or access-control warning. In a memory-management skill, this increases the risk that sensitive personal profiling data is stored or surfaced without users understanding the exposure, especially since the AI is instructed to read the latest entry automatically on session startup.

Missing User Warnings

Medium
Confidence
92% confidence
Finding
The template explicitly tells users to paste raw thoughts, meeting notes, email snippets, overheard insights, and emotional context into an AI-assisted workspace, but provides no warning about sensitive, personal, confidential, or regulated data. In a knowledge-management skill with AI processing and library search, this encourages oversharing and increases the likelihood that private or proprietary information will be stored, indexed, or reused in later outputs.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
- Full autonomy: File, reorganize, update, create new patterns
- Assume direct capture of thinking process
- No confirmation needed for organizational changes

### Writing Style/ — ⚠️ MEDIUM SENSITIVITY
Confidence
75% confidence
Finding
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
- New records: Direct filing OK
- Sensitive decisions: Confirm before archiving
- Archiving old interactions: Propose if >6 months old
- Analysis/summaries: OK to create without confirmation

## Filing Decision Logic
Confidence
75% confidence
Finding
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
- Not tied to personal identity
- Designed for reuse and circulation

**Agent autonomy**: Full. File, reorganize, deduplicate, update without confirmation.

## Directory Structure
Confidence
89% confidence
Finding
This instruction delegates operational decisions to the agent without confirmation, including organizational changes that may have lasting effects. The context is less dangerous than a high-sensitivity data zone, but unrestricted autonomy still creates risk from mistakes, scope creep, and unreviewed modifications.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The file explicitly grants full autonomy to reorganize, deduplicate, and update content without confirmation, but provides no guardrails, rollback requirements, or warning about destructive outcomes. Even in a 'low sensitivity' tools zone, this can cause unintended data loss, overwrites, or unsafe propagation of modified content if the agent misclassifies or mishandles files.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
## Daily Workflow

- **Add**: New tools go straight in (no approval needed)
- **Update**: Improve existing tools freely
- **Organize**: Reorganize by category as needed
- **Deduplicate**: Merge similar tools; keep best version
Confidence
86% confidence
Finding
Allowing new tools to be added with no approval is a weaker issue than deletion, but it still enables uncontrolled repository growth, policy drift, and possible insertion of low-quality or unsafe content. In an autonomous system, unreviewed additions can later influence downstream decisions and workflows.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The workflow allows archive and deletion-adjacent actions with no user-facing warning about reversibility, retention, or review. This increases the chance that an autonomous agent will move or remove content in ways the user did not intend, especially during deduplication or cleanup tasks.

Missing User Warnings

Medium
Confidence
96% confidence
Finding
The agent behavior section authorizes autonomous deletion and states that no confirmation is needed for modifications, which normalizes destructive actions without a safety check. In practice, this can let an agent delete or rewrite reusable assets at scale based on flawed reasoning or ambiguous input.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
## Agent Behavior in This Zone

- Full autonomous action: file, update, reorganize, delete
- No confirmation needed for modifications
- Proactive organization and deduplication
- Maintain version history for important tools
- Keep cross-references current
Confidence
92% confidence
Finding
The statement that no confirmation is needed for modifications gives the agent broad authority to make consequential changes without human oversight. Combined with adjacent permission to delete, this materially increases the risk of accidental corruption, loss, or unauthorized restructuring of the tools library.

Missing User Warnings

Medium
Confidence
88% confidence
Finding
The description highlights that skills can run scripts, trigger workflows, manage system operations, and access external services, but it provides no accompanying safety constraints, approval model, or execution boundaries. Presenting powerful automation without guardrails makes misuse and over-privileged implementations more likely, especially when this file serves as the library-wide guidance document.

Session Persistence

Medium
Category
Rogue Agent
Content
Place supporting documentation in `references/`.

### 5. Register in CLAUDE.md
Add entry to skills list with trigger rule:
```yaml
- skill: new-skill-name
  trigger: "semantic match OR explicit call"
Confidence
80% confidence
Finding
Registering new skills in CLAUDE.md creates durable session-persistent behavior by making capabilities discoverable and triggerable in future interactions. In combination with vague trigger rules, this persistence can convert a one-time change into an ongoing execution path that survives across sessions and broadens the attack surface.

Vague Triggers

Medium
Confidence
94% confidence
Finding
The manifest example encourages registering new skills with the vague rule 'semantic match OR explicit call,' normalizing permissive activation for future skills. Because this document is a template for adding capabilities, the unsafe pattern can propagate across the skill ecosystem and increase accidental or adversarial invocation opportunities.

Vague Triggers

Medium
Confidence
92% confidence
Finding
The file explicitly states that skills may be triggered by semantic match based on interpreted user intent, which creates an overly broad invocation surface for executable capabilities. In a system where skills can run commands, modify files, or integrate with external services, ambiguous triggering can cause unintended execution of powerful actions from ordinary conversation.

Natural-Language Policy Violations

Medium
Confidence
93% confidence
Finding
The language rules state that user conversations default to "your language choice," which imposes a language/locale preference before the user has opted in. Although first-run setup later asks for language confirmation, the earlier default behavior still prescribes a language policy that may apply before consent.

Vague Triggers

Medium
Confidence
95% confidence
Finding
The boot trigger phrases include common conversational terms like 'boot', 'start work', and 'initialize', which can be invoked during ordinary discussion rather than through an explicit, isolated control path. That creates prompt/skill confusion risk: the agent may launch file reads, memory operations, or other startup actions without clear user intent.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The skill authorizes automatic updates to '_this_week.md' during background sync and deep retrospective without a clear user-facing warning or approval step. Silent modification of user files is dangerous because it can alter records, create false provenance, and normalize hidden state changes that users may not notice until later.

Autonomous Decision Making

Medium
Category
Excessive Agency
Content
### Deep Retrospective (Thursday–Sunday)

Auto-execute this; no user trigger required.

1. **Scan**: List all files in `00 Focus Zone/` (exclude `_archive/`), record filename + mtime
2. **Read Weekly Log**: Open `_this_week.md`, extract task checklist and progress notes
Confidence
94% confidence
Finding
The instruction to 'Auto-execute this; no user trigger required' explicitly authorizes autonomous decision-making for scanning files, inspecting contents, inferring progress, and updating records. In a knowledge-management skill handling potentially sensitive notes, this increases the risk of overreach, privacy violations, and unauthorized actions based on heuristics rather than explicit user direction.

Vague Triggers

Medium
Confidence
96% confidence
Finding
Defining skill invocation by unconstrained 'semantic judgment of user intent' gives the agent broad discretion to activate tools and workflows absent precise user consent. In this skill, that discretion is coupled with file access, memory updates, and auto-triggered behaviors, increasing the chance of unauthorized reads/writes or unintended cross-skill execution.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The initialization flow states that the AI will collect user information, create a profile, and create workspace files automatically, but it does not present a clear up-front warning that these actions persist personal data to disk. In a privacy-sensitive knowledge-management skill, silent or poorly disclosed persistence can lead users to share sensitive information without informed consent.

Vague Triggers

Medium
Confidence
93% confidence
Finding
The README tells users to send very generic startup phrases such as "启动" or "开始工作" to activate the workflow. Broad, everyday phrases increase the chance of accidental activation in unrelated conversations, which can trigger file reads, initialization, or memory-management behavior without the user intending to invoke the skill.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The English section similarly describes automatic profile creation and first-week workspace initialization without a clear privacy notice about persistent storage of profile and memory data. Because the system is designed to remember user preferences, tasks, and inferred patterns over time, missing disclosure materially raises the risk of unintended retention of sensitive information.

Static analysis

No suspicious patterns detected.