Back to skill

Security audit

Self Improving 1.2.16

Security checks for vulnerabilities and agentic risk

Overview

The skill is not malicious, but it sets up persistent behavior-changing memory and workspace steering in ways users should review before installing.

Install only if you want an agent to keep persistent local memory that can shape future behavior. Review and approve any changes to AGENTS.md, SOUL.md, and HEARTBEAT.md, consider Strict mode, avoid storing sensitive or third-party data, and inspect ~/self-improving/ regularly. Treat the optional Proactivity install as a separate package that should be reviewed before its setup is followed.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T02 · Agent Memory Poisoning

Warning
Location
setup.md:75
Finding

Persistent Behavior-Changing Memory Loaded Through Workspace Steering

Content
View full analysis
.md` - Project-only override → append to `~/self-improving/projects/.md` - Keep entries short, concrete, and one lesson per bullet; if scope is ambiguous, default to domain rather than global - After a correction or strong reusable lesson, write it before the final response ``` ``` Related promotion behavior from `SKILL.md:161-178`: ```markdown ### 1. Learn from Corrections and Self-Reflection - Log when user explicitly corrects you - Log when you identify improvements in your own work - Never infer from silence alone - After 3 identical lessons → ask to confirm as rule ### 2. Tiered Storage | Tier | Locat ...[truncated 3461 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
setup.md:89
Finding

Unpinned External Companion Skill Is Installed and Immediately Trusted

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (13)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The activation guidance is broad enough to trigger on many normal interactions, including generic mistakes, user feedback, or discovering a 'better approach.' In a skill that persists learned behavior to local files, overly broad activation increases the chance of silent memory collection, preference persistence, and behavioral drift without clear user intent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The 'When to Use' section allows activation whenever the agent notices its own output could be better or when significant work is completed, which is highly subjective and expansive. In context, this makes the skill prone to running during ordinary conversations and tasks, increasing unauthorized retention of user preferences and autonomous behavior changes.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
81% confidence
Finding

The skill authorizes autonomous promotion, demotion, and archival of learned patterns based on internal heuristics and time windows, rather than explicit user approval for each state change. Although it avoids deletion without asking, these automatic memory-management decisions can still materially alter agent behavior and what information is surfaced later, creating integrity and transparency risks.

Content

Scanner excerpt · SKILL.md (reported line 178)May include surrounding context.

md
- Pattern used 3x in 7 days → promote to HOT
- Pattern unused 30 days → demote to WARM
- Pattern unused 90 days → archive to COLD
- Never delete without asking

### 4. Namespace Isolation
- Project patterns stay in `projects/{name}.md`

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · boundaries.md (reported line 11)May include surrounding context.

md
| Financial | Card numbers, bank accounts, crypto seeds | Fraud risk |
| Medical | Diagnoses, medications, conditions | Privacy, HIPAA |
| Biometric | Voice patterns, behavioral fingerprints | Identity theft |
| Third parties | Info about other people | No consent obtained |
| Location patterns | Home/work addresses, routines | Physical safety |
| Access patterns | What systems user has access to | Privilege escalation |

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The kill-switch trigger "forget everything" is ambiguous and could be invoked accidentally in ordinary conversation, quoted text, or third-party content. In a self-improving memory-enabled agent, an overbroad deletion trigger can be abused to erase state, disrupt continuity, or force unintended export-and-delete behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The template explicitly instructs creation of directories and files under the user's home directory on first activation, which causes persistent local state changes without any built-in consent prompt, warning, or scoping restrictions. In the context of a self-improving agent, this is more dangerous because it normalizes autonomous persistence and could lead to unnoticed storage of user data, workspace metadata, or future behavioral changes across sessions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill describes automatic loading, persistence, and periodic maintenance of memory data without any explicit user-facing notice, consent flow, or visibility into when storage is occurring. In a self-improving agent context, this is more dangerous because the feature is designed to continuously accumulate and reuse conversational data across sessions, increasing the chance that users disclose information without understanding it will be retained.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The automatic correction workflow stores user corrections, timestamps, counters, context, and namespace assignment with no stated data minimization, sensitivity filtering, or retention limits. That creates unnecessary persistence of user-provided content and metadata, which can expose private preferences, project details, or sensitive conversation fragments if memory files are accessed, reused incorrectly, or exfiltrated.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The example correction log explicitly preserves timestamps, prior user statements, communication preferences, and contextual notes such as where the interaction occurred. Even as an example, it documents a design pattern of retaining conversational history that could reveal sensitive habits, project information, or behavioral profiles when memory is queried or exported.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill explicitly keeps archived prior preferences for future reference, which increases the amount of retained personal history and creates a long-lived record of superseded user statements. In a memory-centric agent, this makes privacy leakage and unintended reuse more likely, especially if old preferences are surfaced, exported, or applied out of context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The setup directs creation of persistent files under ~/self-improving/ and seeds long-lived memory/state without first requiring explicit user consent or clearly warning that data will be stored on disk. In an agent-skill context, silent persistence is risky because it can retain sensitive user data, alter future agent behavior across tasks, and create non-obvious state the user did not intend to keep.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs the agent to modify AGENTS.md, SOUL.md, and later HEARTBEAT.md, which are workspace steering/configuration files that can materially change agent behavior over time. Doing so without an explicit warning and user approval is dangerous because it silently mutates trusted control files, expands the skill's influence beyond the immediate task, and may override or entrench behavior the user did not review.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The phrase "what do you know about me?" is presented as a user trigger for a full export, but the document does not clarify whether exact phrasing is required or provide boundaries and negative examples. Because this is a natural conversational phrase, it could be matched unintentionally in ordinary discussion about memory or privacy.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.