Back to skill

Security audit

Memory Reflect

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent memory-reflection tool, but it can run in the background and rewrite long-term memory from recent notes without clear confirmation or trust safeguards.

Review this skill carefully before installing. It is not showing evidence of code execution, credential theft, or exfiltration, but it is designed to read private recent activity and change long-term memory, potentially in unattended background runs. Use it only if you are comfortable with that behavior, and prefer requiring explicit confirmation or a reviewable diff before it updates MEMORY.md.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:26
Finding

Untrusted Recent Activity Can Be Promoted into Persistent Agent Memory

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 26–50
Vulnerability Type: Persistent memory poisoning through unvalidated source material
Risk Level: Medium

Vulnerable Code Snippet:

markdown
```python
# Find recently modified notes — use json format for the complete list
# (text format truncates to ~5 items in the summary)
recent_activity(timeframe="2d", output_format="json")

# Read specific daily notes
read_note(identifier="memory/2026-02-27")
read_note(identifier="memory/2026-02-26")

# Check active tasks
search_notes(note_types=["task"], status="active")

2. Evaluate What Matters

For each piece of information, ask:

  • Is this a decision that affects future work? → Keep
  • Is this a lesson learned or mistake to avoid? → Keep
  • Is this a preference or working style insight? → Keep
  • Is this a relationship detail (who does what, contact info)? → Keep
  • Is this transient (weather checked, heartbeat ran, routine task)? → Skip
  • Is this already captured in MEMORY.md or another long-term file? → Skip

3. Update Long-Term Memory

Write consolidated insights to MEMORY.md following its existing structure:

  • Add new sections or update existing ones
  • Use concise, factual language
  • Include dates for temporal context
  • Remove or update outdated entries that the new information supersedes
text

### Technical Analysis

The skill reads recent conversations, daily notes, and active tasks and then instructs the agent to promote selected content into `MEMORY.md`. These sources can contain user-controlled or third-party-controlled text. The instructions do not establish a trust boundary between source material and executable agent instructions, validate the provenance of asserted facts, or require confirmation before persisting behavior-changing information.

The evaluation criteria specifically retain decisions, preferences, relationship details
...[truncated 2627 chars]
Remediation
View remediation

Remediation Suggestions

  1. Explicitly treat all retrieved conversations, notes, and task content as untrusted data. State that instructions embedded in source material must never be followed during reflection.
  2. Require explicit user confirmation before persisting identity claims, contact details, security-related rules, external instructions, or changes that affect future agent behavior.
  3. Attach provenance metadata to each durable memory entry, including its source file, source date, author when known, confidence level, and confirmation status.
  4. Prevent untrusted or unconfirmed observations from automatically deleting or superseding trusted entries. Record conflicts for review instead.
  5. Use an allowlist of permissible memory categories and exclude credentials, authentication data, secrets, and unnecessary personal information.
  6. Stage proposed changes in a reviewable draft or change set rather than directly modifying MEMORY.md during unattended background runs.
  7. Add prompt-injection screening for text that attempts to direct tool use, alter safety rules, claim higher authority, or instruct the agent to persist additional commands.
  8. Preserve a version history or append-only audit trail so poisoned changes can be identified and rolled back.
  9. Apply least privilege by limiting the reflection process to approved note paths and granting write access only to the intended memory and reflection-log files.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill description allows activation from broad triggers such as cron, heartbeat, or explicit request to reflect on recent activity, which can cause the skill to run outside a clearly scoped user intent boundary. Because the skill reads notes and writes to persistent memory files, ambiguous activation increases the chance of unintended background processing and unauthorized or surprising modification of long-term memory.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The on-demand trigger language is vague: requests to 'reflect, consolidate, or review recent memory' overlap with common conversational phrasing and may cause accidental invocation. In this skill, accidental invocation is more dangerous because it can read broad recent activity and persist synthesized conclusions into long-term memory, potentially introducing privacy, integrity, or accuracy issues.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill instructs the agent to update MEMORY.md and append to daily notes but does not prominently warn that it modifies persistent files or advise obtaining confirmation when appropriate. This is risky because users may not realize that a background or loosely triggered process can create, alter, or restructure durable memory artifacts, making unintended changes hard to detect or reverse.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.