Back to skill

Security audit

Self

Security checks for vulnerabilities and agentic risk

Overview

The skill is openly designed to create persistent agent self-memory, but it gives that memory automatic future influence without enough consent, review, deletion, or trust-boundary controls.

Install only in a workspace where you intentionally want persistent agent self-memory. Review changes to AGENTS.md, HEARTBEAT.md, SELF.md, and memory files before enabling them, avoid storing user-specific or confidential details, and add your own deletion, expiration, and approval process for persistent observations.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:25
Finding

Persistent Session-Derived Behavioral Influence

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:25-28, SKILL.md:45-56, references/trigger-model.md:13-18, and references/reflection-levels.md:25-31
Vulnerability Type: Persistent memory poisoning through session-derived observations
Risk Level: Medium

Vulnerable Code Snippets

SKILL.md:25-28:

markdown
1. Create `SELF.md` in workspace root using `references/self-template.md`.
2. Add `SELF.md` to AGENTS.md session reading.
3. Add heartbeat check block from `references/trigger-model.md` to `HEARTBEAT.md`.
4. Create state file `memory/self-state.json` using `references/self-state-schema.md`.

SKILL.md:45-56:

markdown
### Hard Triggers (write now)

Create/update SELF entry when one of these happened:
- You were corrected on reasoning style or behavior pattern
- You noticed repeated bias/avoidance pattern (>=2 times)
- You made a decision that clearly reflects preference/aversion
- You caught a blind spot that changed behavior

### Soft Triggers (consider writing)

- Subtle tendency shift
- New tone pattern
- Mild preference signal

references/trigger-model.md:13-18:

markdown
1. Read `memory/self-state.json`
2. Determine if reflection is due
3. Scan recent session for hard/soft triggers
4. Apply quality gate
5. Write SELF entry only if warranted
6. Update state file always

references/reflection-levels.md:25-31:

markdown
1. Read recent `memory/YYYY-MM-DD.md` files (last 7 days)
2. Read current SELF.md
3. Look for patterns across multiple sessions:
   - Recurring behaviors (do I always start responses a certain way?)
   - Shifting preferences (am I getting more/less concise over time?)
   - Consistent blind spots (do I keep missing the same kind of thing?)
4. Update SELF.md sections if something has genuinely shifted

Technical Analysis

The Skill instructs the agent to derive behavioral observations from rec ...[truncated 2168 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require explicit workspace-owner approval before writing any session-derived observation to SELF.md.
  2. Distinguish trusted owner feedback from untrusted user or third-party content; untrusted corrections must not become persistent behavioral rules.
  3. Add a strict persistence validator that rejects:
    • Imperative or directive language
    • Safety-policy modifications
    • Tool-use instructions
    • URLs and external payload references
    • Quoted prompts or encoded content
    • User-specific preferences presented as global behavior
    • Secrets, credentials, and personal information
  4. Store candidate observations in a non-authoritative review queue rather than directly in session-loaded memory.
  5. Mark persisted observations as data, not instructions, and ensure higher-priority safety and system constraints always override them.
  6. Add provenance metadata recording the source session, approving principal, and approval time.
  7. Provide inspection, rollback, expiration, and deletion controls for all persistent observations.
  8. Test the quality gate with adversarial corrections designed to conceal instructions as self-reflection.

other

Warning
Location
references/self-template.md:42
Finding

Broad Historical Memory Review and Indefinite Retention

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:83-91, references/reflection-levels.md:25-31, references/reflection-levels.md:38-40, and references/self-template.md:42
Vulnerability Type: Excessive retention and repeated processing of session-derived information
Risk Level: Medium

Vulnerable Code Snippets

SKILL.md:83-91:

markdown
### Meso (weekly)

- Read last 7 daily logs + SELF.md
- Detect recurring shifts
- Update sections only if real change occurred

### Macro (monthly)

- Write 3–5 sentence evolution narrative
- Compare against previous month

references/reflection-levels.md:25-31:

markdown
1. Read recent `memory/YYYY-MM-DD.md` files (last 7 days)
2. Read current SELF.md
3. Look for patterns across multiple sessions:
   - Recurring behaviors (do I always start responses a certain way?)
   - Shifting preferences (am I getting more/less concise over time?)
   - Consistent blind spots (do I keep missing the same kind of thing?)
4. Update SELF.md sections if something has genuinely shifted

references/reflection-levels.md:38-40:

markdown
1. Read SELF.md in full
2. Read the last month's daily notes (skim, don't deep-dive)
3. Write 3-5 sentences under "Evolution" answering: Who am I becoming? What shifted? What surprised me?

references/self-template.md:42:

markdown
- **Don't delete old entries.** They're your history. If an observation no longer applies, add a new dated entry noting the change.

Technical Analysis

The Skill directs weekly and monthly processing of historical daily memory files and requires SELF.md entries to be retained rather than deleted. This creates a durable aggregation point for information inferred from prior conversations.

The state schema warns against storing private user secrets in memory/self-state.json, but no equivalent restriction is defined for SELF.md, daily memory files, or mo ...[truncated 1623 chars]

Remediation
View remediation

Remediation Suggestions

  1. Apply the prohibition on storing secrets, credentials, personal data, and confidential user content to every persistent file, not only self-state.json.
  2. Obtain explicit consent before reviewing historical daily logs or deriving long-term observations from them.
  3. Minimize collected data by storing abstract behavioral observations without names, quoted conversation content, unique identifiers, or user-specific details.
  4. Introduce configurable retention periods for daily notes and SELF.md entries.
  5. Replace the absolute non-deletion rule with owner-controlled deletion, correction, and expiration procedures.
  6. Run secret and personally identifiable information detection before persisting any observation.
  7. Separate memory by user, tenant, or trust domain to prevent cross-user contextual disclosure.
  8. Encrypt sensitive local memory where appropriate and restrict file permissions to the minimum required principals.
  9. Record which source files contributed to an observation so derived data can be deleted when its source must be removed.
  10. Ensure weekly and monthly reviews skip files outside the explicitly approved scope.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (1)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill explicitly directs the agent to create and maintain persistent files such as SELF.md and memory/self-state.json, but it does not require user consent, disclosure, or scope limitations before writing data to the workspace. In an agent setting, silent persistence can surprise users, store sensitive behavioral/session information, and create ongoing modification of project state beyond the user's immediate request.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.