Back to skill

Security audit

Use when a session reveals stable user preferences, workflow corrections, or project conventions that should be preserved for future sessions. Also use when user explicitly asks to remember rules, update guidance, or summarize learnings.

Security checks for vulnerabilities and agentic risk

Overview

This skill is not malware, but it can write lasting agent instructions and may skip confirmation when run automatically at session end.

Install only if you want an agent to persist preferences into CLAUDE.md. Review each proposed rule before it is written, avoid automatic end-of-session writes, and use project-local CLAUDE.md unless you intentionally want the rule to affect all future projects.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:26
Finding

Untrusted Session Instructions Can Be Persisted as Agent Memory Without Confirmation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:26-40, 61-67, 76-83; related persistence criteria in references/learning-rules.md:5-8
Vulnerability Type: T02: Agent Memory Poisoning
Risk Level: Medium

Vulnerable Code Snippets

SKILL.md:26-40

markdown
### Step 1: Scan the session

Review the whole session for two categories:

**A. Workflow preferences and corrections**
- User corrections ("don't do that", "stop ...")
- User-confirmed non-obvious approaches ("yes, exactly", "perfect")
- Output format, communication style, and collaboration flow preferences
- Tool usage guidance
- `prompt-refiner` choice preferences (prefers refined vs original, wants compare-before-execute)

**Judging prompt-refiner preference signals:**
- Repeatedly choosing the same version → store as stable preference
- One-off different choice → don't overfit, may be scenario-specific
- Explicit verbal correction ("don't show me the original anymore") → promote immediately to rule

SKILL.md:61-67

markdown
### Step 3: Read current CLAUDE.md

Choose target file by scope:
- **Global rules** → `~/.claude/CLAUDE.md`
- **Project-specific rules** → project `CLAUDE.md`

If a project file is needed and missing, create it.

SKILL.md:76-83

markdown
### Step 5: Confirm changes

Show the user:
- Rules to add
- Rules to update (old → new)
- Target file

Wait for confirmation before writing, unless user says "update directly" or the skill is running from an automatic SessionEnd hook.

references/learning-rules.md:5-8

markdown
### High-value signals (always learn)
- User explicitly corrects your behavior → immediate rule
- User confirms a non-obvious approach → stable preference
- Repeated pattern across 3+ interactions → stable convention
- User states a project rule ("we always use X for Y") → project convention

Technical Analysis

The Skill interpr ...[truncated 2409 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require explicit user confirmation before every persistent write, including writes initiated by SessionEnd hooks. Remove the automatic-confirmation bypass.
  2. Display the exact proposed rule, its source message, destination file, and scope before requesting approval.
  3. Accept learning signals only from direct user-authored messages. Exclude tool output, retrieved content, repository files, quoted text, generated prompts, and assistant paraphrases.
  4. Reject candidate rules that modify safety constraints, instruction precedence, permission boundaries, credential handling, confirmation requirements, or network and tool access.
  5. Default to project-local persistence. Require separate, explicit consent for each write to global ~/.claude/CLAUDE.md.
  6. Record provenance metadata or maintain an auditable change log so users can identify, review, and revert learned rules.
  7. Validate generated rules against a restrictive allowlist of preference categories, such as formatting and non-security-sensitive coding conventions.
  8. Treat persisted preferences as lower-priority, untrusted guidance rather than authoritative instructions that can override system or security policies.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill description is broad enough to trigger on ordinary conversations about remembering preferences, updating guidance, or summarizing learnings, which can cause the agent to activate in contexts the user did not clearly intend. Because this skill reads the whole session and proposes persistence into CLAUDE.md, overbroad invocation increases the chance of capturing sensitive, transient, or poorly-vetted instructions and turning them into durable behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This markdown file includes example 'After' rules written in Chinese, such as '测试优先:先写测试再写实现代码' and '使用纯文本确认,不使用 AskUserQuestion 弹窗', without offering a language choice or explaining that the skill is region-specific. That can encode a specific language/locale preference into learned rules, which conflicts with the policy against forcing a language without user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.