Back to skill

Security audit

self-improving-science

Security checks for vulnerabilities and agentic risk

Overview

This skill mainly logs science and ML lessons with optional reminders, and I found no hidden exfiltration, destructive behavior, or deceptive authority expansion.

Install only if you want persistent science/ML learning logs. Keep hooks project-scoped, use narrow matchers, review hook code before enabling it, avoid logging secrets or raw sensitive data, and treat any generated skill scaffold as untrusted until manually reviewed.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (9)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a substantive research-methodology skill used to document and consult learnings about experiment flaws and ML/statistical issues. The supplied code does not analyze learnings, detect any of the listed conditions, review prior learnings, or capture methodology corrections. Instead, it is an extraction/helper utility that generates a new skill scaffold on disk from a name. Its primary purpose is materially different, and it exercises undeclared file-creation capabilities. While comments mention creating a skill from a research learning entry, the implementation only produces a template and does not perform the described scientific continuous-improvement behavior.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 522)May include surrounding context.

md
Extracted skills are untrusted until a human reviews the generated `SKILL.md`. Do not keep or publish an extracted skill without explicit user approval.

Session Persistence

Medium
Category
Rogue Agent
Confidence
83% confidence
Finding

The skill instructs creation of persistent files under ~/.openclaw/workspace/.learnings, which can store research notes across sessions. Even though the content warns against logging secrets, persistent memory in a home-directory workspace increases the chance that sensitive experimental details, proprietary methods, or regulated data summaries are retained longer than intended and later exposed to unrelated workflows or users on a shared environment.

Content

Scanner excerpt · SKILL.md (reported line 83)May include surrounding context.

└── FEATURE_REQUESTS.md

text

### Create Learning Files

```bash
mkdir -p ~/.openclaw/workspace/.learnings

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

L109 says the skill 'does not read other sessions' and should log notes 'in this workspace only,' but other parts of the file describe OpenClaw injecting MEMORY.md, daily memory/ files, AGENTS.md, SOUL.md, and TOOLS.md into sessions (L67-L80) and instruct promotion/review against those artifacts (for example L96-L105, L537-L542). That documentation creates an intent-level inconsistency about whether the skill is strictly workspace-local logging or also reads broader persistent memory artifacts.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
85% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILL.md (reported line 592)May include surrounding context.

md
### Ownership Rules
- This skill writes only to `.learnings/science/` in stackable mode.
- Do not read other skill folders, their SKILL.md files, or their log entries.
- Standalone mode writes to this project's `.learnings/*.md` log files only.
- Stackable mode writes only to the namespaced folder above and must not rewrite other skills' log entries.
- Promotion into `AGENTS.md`, `SOUL.md`, `TOOLS.md`, `MEMORY.md`, rules, hooks, or generated skills is not a logging write. Show a reviewed diff and apply only after explicit user approval.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The example matcher train|model|experiment|data|analysis|test is broad enough to trigger on many routine prompts unrelated to scientific self-improvement. That can cause frequent unsolicited hook execution, increasing unnecessary context injection and the chance that sensitive prompt content or tool output is processed more often than intended.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
79% confidence
Finding

The guide instructs users to install a hook into OpenClaw's persistent hooks directory, which enables behavior that survives beyond the current session and can automatically trigger on lifecycle events. Even though the stated purpose is legitimate, persistent hooks expand the attack surface because a compromised or later-modified hook could silently alter workspace files or influence future sessions without repeated user consent.

Content

Scanner excerpt · references/openclaw-integration.md (reported line 48)May include surrounding context.

Copy the hook to OpenClaw's hooks directory:

bash
mkdir -p .openclaw/hooks
cp -r hooks/openclaw .openclaw/hooks/self-improving-science

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

Lines L493-L507 present an 'Advanced Setup (With Error Detection)' heading, but the JSON is identical to the basic activator-only setup. This is then reinforced by L509, which says to leave PostToolUse disabled, so the documentation actively contradicts the stated purpose of that section.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file directs the user or skill behavior to create and update .learnings/, TOOLS.md, and AGENTS.md, which are file-write operations affecting workspace contents. The document provides setup steps and write destinations but does not include any explicit warning that these actions will persist or modify local project files.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.