Back to skill

Security audit

Feihong Self Improving Agent

Security checks for vulnerabilities and agentic risk

Overview

This skill is intended to record agent learnings, but it stores and promotes conversation/error details into long-lived agent memory with broad hooks and weak redaction guidance.

Install only if you want persistent agent memory. Keep .learnings local by default, redact secrets and personal or business-sensitive data before logging, avoid global hooks unless you fully trust and review the scripts, and require explicit human approval before promoting any learning into agent instruction files.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:262
Finding

Untrusted Content Can Be Promoted into Persistent Agent Instructions

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:165
Finding

Error Logging Workflow Can Persist Credentials and Other Sensitive Data

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract-skill.sh:93
Finding

Symlink Traversal Bypasses the Skill Extraction Workspace Boundary

Content
View full analysis
"$SKILL_PATH/SKILL.md" << TEMPLATE ``` ### Technical Analysis The script attempts to confine writes to the current workspace by rejecting absolute paths and lexical `..` components. This only validates the path string; it does not validate the resolved filesystem destination. A relative component can be a symbolic link pointing outside the workspace. Shell redirection and `mkdir -p` follow such links, so an accepted relative `--output-dir` can resolve to an arbitrary directory writable by the invoking user. Skill-name validation prevents injection through `SKILL_NAME`, but it does not address symlink traversal through `SKILLS_DIR`. ### Attack Path 1. An attacker with the ability to prepare workspace files creates a symbolic link: ```bash ln -s /tmp/external-target out ``` 2. The user or agent invokes: ```bash ./scripts/extract-skill.sh example-skill --output-dir out ``` 3. The value `out` is relative ...[truncated 923 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (14)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description describes a reflective/knowledge-capture capability: recording failures, corrections, missing capabilities, API/tool failures, outdated knowledge, and improved approaches for future use. The supplied code does not implement any of that behavior. Instead, it is a command-line scaffolding utility that generates a new skill folder and template markdown file. While the comments mention creating a skill from a learning entry, the script neither captures learnings nor reviews them; it only creates boilerplate files for a skill. This is a materially different primary purpose, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Directing users to modify ~/.claude/settings.json touches a sensitive agent configuration directory that controls behavior across sessions. If that location or the referenced skill path is later tampered with, the agent will continue executing attacker-influenced hooks automatically, making this a high-value persistence and privilege surface.

Content

Scanner excerpt · references/hooks-setup.md (reported line 48)May include surrounding context.

Option 2: User-Level Configuration

Add to ~/.claude/settings.json for global activation:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The description says to use the skill whenever a command fails unexpectedly, the user corrects the agent, a capability is missing, an external tool fails, knowledge is outdated, or a better approach is discovered. These are very broad, common situations in ordinary agent interactions, and the file does not provide exclusion conditions or clear boundaries for when not to invoke the skill, increasing the chance of unintended invocation.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly encourages persistent recording of user corrections, requests, context, and promotion of those learnings into durable memory files. In an agent setting, that can capture secrets, proprietary code details, personal data, or sensitive workflow information in long-lived plaintext artifacts, increasing the blast radius of any later access or sharing.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
79% confidence
Finding

Persisting learning files under a user-level workspace creates durable memory across sessions, which can retain sensitive information far beyond the originating task. Persistence is not inherently unsafe, but in this skill it combines with broad logging guidance and cross-session use, making retention risk materially higher.

Content

Scanner excerpt · SKILL.md (reported line 64)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill instructs agents to read other sessions' transcripts and send learnings between sessions, which creates a direct pathway for unintended cross-session disclosure. Even if meant for convenience, session histories often contain confidential prompts, outputs, debugging artifacts, or credentials, so semantic transfer between sessions materially increases leakage risk.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The templates ask for 'full context,' inputs, parameters, error output, and user context, which strongly biases users or agents toward storing raw sensitive content. Plaintext operational logs commonly accumulate API keys, file paths, stack traces, customer data, and confidential requests, making these templates a practical data-retention vulnerability.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The 'Detection Triggers' section includes cues like 'Can you also...', 'Is there a way to...', and 'Unexpected output or behavior,' which are common phrases that can occur in many ordinary conversations. Because the section says to 'Automatically log when you notice' these patterns without stating limits or exceptions, it creates ambiguous activation criteria for the skill.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/examples.md (reported line 301)May include surrounding context.

When the above learning is extracted as a skill, it becomes:

File: skills/docker-m1-fixes/SKILL.md

markdown
---

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

Project-level hook configuration creates persistent automatic execution for future sessions, not just the current troubleshooting task. Even though this is opt-in, persistence raises risk because the behavior can outlive the user's intent and continue injecting instructions or processing tool output in later contexts.

Content

Scanner excerpt · references/hooks-setup.md (reported line 15)May include surrounding context.

Option 1: Project-Level Configuration

Create .claude/settings.json in your project root:

json
{

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The empty matcher causes the hook to fire on every user prompt, creating broad and persistent prompt-time influence over the agent. In the context of a self-improvement skill, this increases the chance of unneeded context injection, prompt bloat, and behavior steering across unrelated tasks, which can be abused if the hooked script or its output becomes unsafe.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Installing the hook in ~/.claude/settings.json enables the behavior across all projects and sessions, including contexts where the skill is irrelevant or where repositories are untrusted. This broadens exposure and makes any mistake, malicious modification, or over-collection in the hook affect every interaction rather than a single project.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Allowing --output-dir and then writing new directories/files gives this helper a general project-generation capability within the workspace. For a skill whose purpose is to capture corrections and learnings, repository scaffolding in arbitrary subdirectories is not an obvious or necessary capability.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest describes a skill focused on recording learnings, corrections, failures, and reviewing them before tasks. This script instead creates a new skill directory and writes a new SKILL.md template, which is a repository scaffolding capability rather than a direct implementation detail of storing or reviewing learnings.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.