Back to skill

Security audit

Self Improving Agent

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed self-improvement system, but it asks agents to persist broad tool activity and promote session learnings into future agent instructions with weak privacy and review controls.

Review carefully before installing. Prefer project-level hooks over global ~/.claude settings, avoid wildcard raw observation unless you add redaction and retention limits, keep .learnings out of version control unless intentionally sharing, and require human review before promoting any learning into agent instruction files.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
hooks/observe.sh:95
Finding

Persistent Plaintext Storage of Unfiltered Tool Inputs and Outputs

Content
View full analysis
> "$obs_file" } ``` ### Technical Analysis The observation hook accepts complete tool input and output values and stores them in persistent JSONL files without redaction. The documented configuration registers this hook for all `PreToolUse` and `PostToolUse` events, so captured values can include source code, command output, private URLs, API tokens, credentials printed during debugging, user-supplied confidential data, and environment details. The files are created according to the process umask rather than with explicit restrictive permissions. No field-level filtering, maximum record size, retention limit based on age, secret detection, or tool allowlist is applied. When the observation file grows, the im ...[truncated 1668 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:542
Finding

Untrusted Session Learnings Can Be Promoted into Persistent Agent Instructions

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
hooks/observe.sh:13
Finding

Path Traversal Through Unvalidated CLAUDE_PROJECT_DIR

Content
View full analysis
git remote URL hash > git repo path > global detect_project() { if [[ -n "$CLAUDE_PROJECT_DIR" ]]; then echo "$CLAUDE_PROJECT_DIR" return fi ``` The value is subsequently used directly as a filesystem path component: ```bash init_directories() { local project_id="$1" mkdir -p "$HOMUNCULUS_DIR/instincts/personal" mkdir -p "$HOMUNCULUS_DIR/instincts/inherited" mkdir -p "$HOMUNCULUS_DIR/evolved/agents" mkdir -p "$HOMUNCULUS_DIR/evolved/skills" mkdir -p "$HOMUNCULUS_DIR/evolved/commands" mkdir -p "$HOMUNCULUS_DIR/projects/$project_id/instincts/personal" mkdir -p "$HOMUNCULUS_DIR/projects/$project_id/instincts/inherited" mkdir -p "$HOMUNCULUS_DIR/projects/$project_id/evolved/skills" mkdir -p "$HOMUNCULUS_DIR/projects/$project_id/evolved/commands" mkdir -p "$HOMUNCULUS_DIR/projects/$project_id/evolved/agents" mkdir -p "$HOMUNCULUS_DIR/projects/$project_id/observations.archive" } ``` It also controls observation and archive destinations: ```bash if [[ "$project_id" == "global" ]]; then obs_file="$OBSERVATIONS_FILE" else obs_file="$HOMUNCULUS_DIR/projects/$project_id/observations.jsonl" fi ``` ```bash archive_file="$HOMUNCULUS_DIR/projects/$project_id/observations.archive/$(date +%Y%m%d_%H%M%S).jsonl" mv "$obs_file" "$archive_file" gzip "$archive_file" ``` ### Technical Analysis Fallback project identifiers derived from Git information are hashed, but `CLAUDE_PROJECT_DIR` is accepted verbatim. The value is later interpolated into directory and file paths under `$HOMUNCULUS_DIR/projects`. Shell quoting prevents command substitution and argument splitting, but it does not prevent filesystem traversal. A va ...[truncated 1928 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (33)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README instructs users to install wildcard PreToolUse and PostToolUse hooks that capture 100% of activity, but it does not warn that this may collect sensitive prompts, file paths, command metadata, or secrets handled during tool use. Because the hooks are attached globally via the agent settings directory, this creates broad surveillance and data retention risk across all projects, increasing exposure if the hook script stores or transmits data insecurely.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · README.md (reported line 17)May include surrounding context.

1. Enable Observation Hooks

Add to ~/.claude/settings.json:

json
{

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The declared description presents a substantive continuous-learning system with storage semantics, scoring, scope management, and evolution workflows. The supplied code does not implement that system. Instead, it only listens for agent bootstrap events and injects a reminder markdown file telling the agent where to log learnings manually. This is materially narrower and behaviorally different from the declared purpose. It also introduces a specific trigger and capability—bootstrap-time prompt/file injection—that is not reflected in the description. While the reminder relates to learning behavior conceptually, the implementation lacks the core declared features, so this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a broad continuous learning framework with persistence, scoring, scoping, and evolution features. The supplied code does not implement those capabilities. Instead, it is a narrowly focused Bash error detector hook: it reads tool output, matches against a fixed set of error patterns, and prints a suggestion to log the error. While this could be a small supporting component within a larger self-improvement system, the code chunk by itself materially differs from the declared primary purpose and introduces a concrete trigger context (PostToolUse for Bash) absent from the description. Therefore this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a broad learning-management system centered on capturing atomic learnings ('instincts'), scoring confidence, handling scope, and evolving learnings into skills/commands/agents. The supplied code does not implement any of those core behaviors. Instead, it is a narrowly focused utility that creates a markdown scaffold for a new skill folder and file based on a provided name. While this could be a small supporting tool in a larger learning system, the code chunk itself materially differs in primary purpose and capabilities from the declared description. Therefore, this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

The skill instructs users to modify ~/.claude/settings.json to install broad PreToolUse/PostToolUse hooks that execute a command on every tool invocation. Access to agent-wide configuration is security-sensitive because it can create persistent, global behavior changes and a privileged execution path that affects all future sessions, not just this project.

Content

Scanner excerpt · SKILL.md (reported line 210)May include surrounding context.

Enabling Observation Hooks (v2)

Add to your ~/.claude/settings.json:

json
{

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Directing users to modify ~/.claude/settings.json affects the agent's persistent user-level configuration directory, which can change behavior across all future sessions. In this skill's context, that is especially sensitive because it establishes automatic command-hook execution globally, increasing persistence and trust surface beyond a single project.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 177)May include surrounding context.

sessions_send

Send message to another session:

text
sessions_send(sessionKey="session-id", message="Learning: API requires X-Custom-Header")

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 49)May include surrounding context.

/projects # List known projects

text

### 3. Create Instincts Manually

```yaml
---

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This README is a markdown file, so vague-trigger checks apply. The example trigger value when to apply is an unspecific placeholder that gives no concrete activation phrase, scope, or exclusion criteria, which could lead users to define overly broad triggers that collide with normal conversation or ambiguous situations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The description says to use the skill when a command fails, when a user corrects you, when discovering patterns, or when learned behaviors need review. Several of these conditions, especially 'User corrects you' and 'Discovering patterns,' are very broad and lack limiting context or negative examples, which increases the chance of unintended invocation.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 91)May include surrounding context.

md
|-------|---------|----------|
| 0.3 | Tentative | Suggested but not enforced |
| 0.5 | Moderate | Applied when relevant |
| 0.7 | Strong | Auto-approved for application |
| 0.9 | Near-certain | Core behavior |

**Confidence increases when:**

Session Persistence

Medium
Category
Rogue Agent
Confidence
76% confidence
Finding

The evolve/promote workflow normalizes turning session-derived observations into persistent skills, commands, and global instincts. In this skill's context, that persistence can entrench unsafe or sensitive content into reusable artifacts, amplifying mistakes or leaked information across later sessions and projects.

Content

Scanner excerpt · SKILL.md (reported line 162)May include surrounding context.

bash
/evolve
# Analyzes instincts and suggests:
# - "Create skill: react-testing-workflow.md"
# - "Create command: /test-component"
# - "Promote prefer-functional-style to global (seen in 3 projects)"

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill explicitly encourages inter-session sharing via tools like session history, sending learnings to other sessions, and spawning sub-agents. In context, that broadens the disclosure surface for user/session content and can propagate sensitive information beyond the original interaction boundary.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
86% confidence
Finding

Creating a .learnings/ directory in the project establishes durable, potentially repo-visible storage for session-derived information. If users or agents log broad context there, sensitive operational details can be retained locally, committed accidentally, or exposed to collaborators and future agents.

Content

Scanner excerpt · SKILL.md (reported line 319)May include surrounding context.

Generic Setup (Other Agents)

For Claude Code, Codex, Copilot, or other agents, create .learnings/ in your project:

bash
mkdir -p .learnings

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This workflow tells the agent to persist user corrections, knowledge gaps, and contextual information into durable project files and then promote them into standing instruction files. That can encode sensitive user-provided information or proprietary project context into long-lived memory artifacts that may later be surfaced to other sessions or collaborators.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The logging templates request full context, inputs, error output, and user context, which are precisely the fields most likely to contain secrets, personal data, internal paths, stack traces, or proprietary material. Because the skill promotes durable logging and possible later sharing, this creates a concrete data-retention and inadvertent disclosure risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases overlap with ordinary conversation and instruct automatic logging based on common language like corrections, feature requests, and errors. In a skill centered on persistent learning and cross-session memory, this creates a real risk of capturing user content without sufficiently specific consent or filtering, especially when combined with hooks and durable storage.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The command explicitly supports --auto-create, which can modify the workspace by generating files, but the documentation does not clearly warn users that this option performs write actions. In an agent context, unclear mutation semantics are risky because users or higher-level automation may invoke the command expecting analysis-only behavior and unintentionally create or overwrite files.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrase "when handling user input or external data" is extremely broad and lacks clear activation boundaries. Because many skills and workflows involve some form of user input or external data, this could cause unintended or overly frequent invocation without clear exclusions or negative examples.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrase "when about to modify existing code" is a broad natural-language condition that could match many ordinary coding situations without clearly defining activation boundaries. Although the file provides examples, it does not specify concrete trigger phrases, scope constraints, or exclusion conditions that would prevent unintended invocation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest trigger "when writing React components" is a broad activation condition that could apply to a large range of routine coding activity without clearly defining when the skill should or should not activate. Although the file includes examples of preferred patterns, it does not narrow the trigger with explicit constraints or exclusion conditions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The hook persistently logs tool inputs and outputs to disk, which can capture prompts, file contents, secrets, tokens, credentials, proprietary code, and other sensitive session data without any visible consent, minimization, or redaction controls. In a self-improving agent, this is especially risky because the design encourages broad continuous collection across sessions and projects, increasing both privacy exposure and the blast radius if the local store is accessed by another process, user, or later evolution step.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
82% confidence
Finding

The guide instructs users to create persistent hook configuration in .claude/settings.json, causing behavior to carry across the session lifecycle for the repository. Persistence is not inherently malicious, but in this self-improving skill it increases the chance of unnoticed ongoing automatic execution and continued context capture after the original task is complete.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Static analysis

No suspicious patterns detected.