Back to skill

Security audit

Self-Improving with Reflection

Security checks for vulnerabilities and agentic risk

Overview

This skill is openly designed to remember feedback, but its setup changes agent-wide instruction files and automatically stores behavior rules across sessions, so it needs careful review before installation.

Install only if you want an agent to keep persistent local memory that can influence future sessions. Before enabling it, remove or review the AGENTS.md, SOUL.md, and HEARTBEAT.md setup steps, require explicit confirmation before every saved memory entry, and avoid storing sensitive, third-party, credential, financial, medical, location, or access-pattern information.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
Findings (4)

T02 · Agent Memory Poisoning

Error
Location
setup.md:62
Finding

Persistent Behavioral Steering Through Agent Memory and Core Instructions

Content
View full analysis
.md` - Do not read unrelated domains "just in case" If inferring a new rule, keep it tentative until human validation. ``` Inside the "Write It Down" bullets, refine the behavior (non-destructive): - Keep existing intent, but route execution-improvement content to `./self-improving/`. - If the exact bullets exist, replace only these lines; if wording differs, apply equivalent edits without removing unrelated guidance. Use this target wording: ```markdown - When someone says "remember this" → if it's factual context/event, update `memory/YYYY-MM-DD.md`; if it's a correction, preference, workflow/style choice, or performance lesson, log it in `./self-improving/` - Explicit user corr ...[truncated 2894 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
setup.md:40
Finding

Setup Exceeds the Declared Filesystem Scope by Modifying Agent-Wide Configuration

Content
View full analysis
Remediation
View remediation

T06 · System Persistence

Warning
Location
setup.md:152
Finding

Cross-Session Persistence Through Heartbeat and Scheduled Maintenance Hooks

Content
View full analysis
30 days to WARM 3. Archive unused >90 days to COLD 4. Run compaction if any file >limit 5. Update index.md 6. Generate weekly digest (optional) ``` From `setup.md:152-161`: ```markdown ## Optional: Heartbeat Integration Add to `HEARTBEAT.md` for automatic maintenance: ```markdown ## Self-Improving Check - [ ] Review corrections.md for patterns ready to graduate - [ ] Check memory.md line count (should be ≤100) - [ ] Archive patterns unused >90 days ``` ``` ### Technical Analysis The Skill proposes both weekly “Cron” maintenance and integration into `HEARTBEAT.md`. No operating-system crontab installation command is present, so the evidence does not establish an OS-level scheduled task. However, modifying an agent heartbeat file is still a recurring cross-session hook in environments that process that file automatically. The hook repeatedly reviews, promotes, compacts, and archives persistent memory. If an unsafe entry has already reached memory, automated maintenance can preserve and reorganize it without fresh user review. This extends the lifetime and operational reach of the Skill beyond a single invocation. ### Attack Path 1. The user enables the optional heartbeat integration or implements the documented weekly maintenance process. 2. A user-controlled or incorrectly inferred rule is written into persistent memory. 3. The heartbeat or scheduled maintenance runs in a later session. 4. The maintenance process reviews, moves, compacts, or promotes the entry. 5. The entry remains available to the automatic memory-loading mechanism and continues affecting future tasks. ### Impact Assessment The hook creates recu ...[truncated 413 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
learning.md:24
Finding

Inconsistent Consent and Third-Party Data Rules Permit Improper Persistent Storage

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (16)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The correction triggers are broad and conversational, so normal user dialogue can be misclassified as a durable preference or correction and written to persistent memory without clear confirmation. In a self-improving skill, this creates a meaningful risk of unauthorized retention, memory poisoning, and future behavior drift because later responses may rely on incorrectly learned patterns.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README describes automatic learning from user corrections and preferences but does not prominently warn users that this information may be stored in persistent files. This weakens informed consent and increases privacy risk, especially because the skill is specifically designed to retain and reuse user-derived data across sessions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The description says to use the skill before starting work and after responding to the user, which is broad enough to trigger routine persistent-memory behavior in many ordinary interactions. Because the skill stores corrections and preferences to local files, overbroad activation increases the chance of collecting and retaining unnecessary user data across sessions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The top-level description emphasizes self-improvement but does not immediately disclose that the skill persistently stores user preferences, corrections, and reflections in local files. This weakens informed consent and can cause users or operators to invoke the skill without realizing it performs cross-session retention.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The 'When to Use' section uses subjective triggers like 'significant work' and 'something... could be better,' which allows the agent to decide for itself when to read/write memory. Ambiguous triggers make unintended persistence more likely and reduce meaningful user awareness or control over what gets stored.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

These instructions tell the agent to automatically log corrections, preferences, and repeated patterns into persistent files, creating a natural-language retention surface for sensitive or identifying information. Even with partial boundary guidance elsewhere, automatic logging of user statements can capture data that is private, context-specific, or later exposed through memory queries and exports.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 166)May include surrounding context.

md
- Pattern used 3x in 7 days → promote to HOT
- Pattern unused 30 days → demote to WARM
- Pattern unused 90 days → archive to COLD
- Never delete without asking

### 4. Namespace Isolation

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · boundaries.md (reported line 11)May include surrounding context.

md
| Financial | Card numbers, bank accounts, crypto seeds | Fraud risk |
| Medical | Diagnoses, medications, conditions | Privacy, HIPAA |
| Biometric | Voice patterns, behavioral fingerprints | Identity theft |
| Third parties | Info about other people | No consent obtained |
| Location patterns | Home/work addresses, routines | Physical safety |
| Access patterns | What systems user has access to | Privilege escalation |

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger phrase "what do you know about me?" is broad and conversational, so an ordinary user query could unintentionally invoke a full memory export. In a self-improving memory skill, accidental export increases the chance of over-disclosing stored personal context and creates unnecessary exposure of historical data.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The phrase "forget everything" is also natural language that may appear in ordinary conversation, hypotheticals, or quoted text, causing unintended invocation of destructive memory actions. Because the kill switch performs export plus wipe, accidental triggering could both disclose sensitive data and destroy state unexpectedly.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The command trigger "What do you know about X?" is broad and resembles normal conversation, so the skill may expose stored memory when the user is merely asking a general question rather than explicitly requesting retrieval from memory. In a self-improving memory skill, this increases the chance of unintended disclosure of retained preferences, project details, or prior interactions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The destructive commands "Forget X" and especially "Forget everything" are presented without strong warnings, confirmation flow, or explanation of scope and permanence. This creates risk of accidental or socially engineered deletion of memory across all tiers, potentially causing irreversible loss of user data and operational context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Automatically loading memory.md and index.md at session start accesses persisted data without any user-facing notice or consent checkpoint. In a system that stores preferences, corrections, and project history, silent retrieval can violate user expectations and expose sensitive contextual information before the user explicitly requests memory use.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The setup instructs the agent to read and update persistent files under ./self-improving/ as part of normal task flow, but it does not require user awareness or consent when storing interaction-derived corrections, preferences, or lessons. This creates a real risk of silent persistence of potentially sensitive user data and behavioral profiling across sessions, especially because the storage is framed as routine pre/post-task behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The instruction to write corrections or reusable lessons before the final response encourages automatic persistence of task interaction content without notifying the user at the moment data is captured. Because this happens inline with response generation, users may unknowingly have sensitive prompts, mistakes, preferences, or workflow details recorded permanently.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The condition "If project detected → preload relevant namespace" is ambiguous because it does not define how a project is detected or what confidence threshold is required. This can cause over-broad or mistaken loading of project-specific memory, creating unnecessary data exposure and cross-project context leakage.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.