Back to skill

Security audit

Self-Evolving

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly purpose-aligned, but it asks the agent to save durable behavior and activation rules into broad user memory with unclear scope and removal controls.

Review this before installing if you do not want an agent to develop durable behavioral preferences. If installed, keep all saved rules inside ~/self-evolving/, require confirmation before any write to general memory, and periodically inspect or delete stored activation cues and always/never rules.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
setup.md:32
Finding

Persistent Behavioral Writes Outside the Declared Skill Storage Boundary

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (3)

Self-Modification

High
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · SKILL.md (reported line 13)May include surrounding context.

md
## When to Use

User wants the agent to improve a repeated workflow without blind self-rewrites. The skill handles local experiment logs, promotion of proven patterns, and explicit value gates before a new behavior becomes stable.

## Architecture

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The setup explicitly instructs the skill to learn broad activation criteria within the first few exchanges and includes examples like recurring workflows, mistakes, optimization topics, and visible friction. Those triggers are subjective and expansive, which can cause the skill to activate outside clear user intent, leading to unexpected persistence of behavioral notes and experimentation-related behavior in unrelated contexts.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · boundaries.md (reported line 18)May include surrounding context.

md
- Read unrelated secrets, credentials, or private files
- Run hidden experiments on the user without a task-level reason
- Treat silence as proof that a change worked
- Expand scope from one mutation into a whole-system rewrite

## Escalate Instead

Static analysis

No suspicious patterns detected.