Back to skill

Security audit

self-track

Security checks for vulnerabilities and agentic risk

Overview

The skill is not overtly malicious, but it gives broad self-tracking triggers authority to write persistent memory and create or push new skills without clear user approval.

Install only if you want the agent to maintain persistent learning records and potentially create or publish new skills. Before use, narrow the triggers and require confirmation for memory writes, vector-memory additions, new skill creation, commits, and pushes.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:25
Finding

Untrusted Research Content Can Be Written to Persistent Agent Memory

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 25–35
Vulnerability Type: Persistent memory poisoning
Risk Level: Medium

Vulnerable Instructions

markdown
### When I encounter something I don't know:
1. Add to `memory/gaps.md` with status "TODO"
2. Research (RSS feeds, web search, docs)
3. Attempt to solve
4. On success: mark gap "DONE" + date + notes
5. On failure: keep as TODO, note blockers

### After learning something significant:
1. Add to `memory/YYYY-MM-DD.md` under "## Learned"
2. Store in vector memory: `python3 scripts/ollama_mem.py add "insight" --category learning --importance 0.8`
3. Update `memory/gaps.md` if gap was closed
4. Update `MEMORY.md` if major milestone

Technical Analysis

The skill instructs the agent to research information from RSS feeds, web search results, and documentation, and then persist learned information in daily logs, vector memory, gap records, and long-term memory. It does not require source validation, provenance metadata, sanitization, separation of quoted content from agent instructions, or user approval before persistent writes.

An attacker who controls or influences a researched source could embed deceptive claims or instruction-like text in that source. If the agent interprets this content as a significant insight and stores it, later memory retrieval may present the hostile content as trusted historical context. The risk is especially relevant to vector memory because semantic retrieval can surface stored content in unrelated future sessions.

The project contains only SKILL.md; the referenced scripts/ollama_mem.py implementation is absent and therefore was not available for verification. Consequently, no claim is made that the storage utility itself executes stored content.

Attack Path

  1. The agent encounters a knowledge gap and follows the instruction to research RSS feeds, websites, or documentation.
  2. An attacker- ...[truncated 1103 chars]
Remediation
View remediation

Remediation Suggestions

  1. Require explicit user approval before information from external sources is added to long-term or vector memory.
  2. Store provenance metadata with every entry, including the source URL, retrieval date, author or publisher, and a trust classification.
  3. Keep raw research notes in an isolated, untrusted store rather than directly placing them in operational memory.
  4. Sanitize stored content by rejecting or quarantining imperative instructions, role directives, tool commands, credential requests, and text that attempts to modify agent policy.
  5. Require corroboration from multiple trusted sources before promoting research into MEMORY.md or assigning it high importance.
  6. Mark retrieved memories as untrusted data and prohibit treating them as system or developer instructions.
  7. Add review, expiration, correction, and deletion workflows so poisoned entries can be identified and removed.
  8. Restrict memory-writing utilities to designated files and structured fields, with audit logging for all persistent changes.
Vulnerability Patterns
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (2)

Self-Modification

High
Category
Rogue Agent
Confidence
95% confidence
Finding

The skill explicitly instructs the agent to create new capabilities by initializing a new skill, writing SKILL.md and resources, validating, and committing/pushing the result. This is a self-modification pathway that can expand agent behavior and persistence without strong authorization boundaries, making accidental or adversarially induced capability creation dangerous.

Content

Scanner excerpt · SKILL.md (reported line 51)May include surrounding context.

md
When I need a new capability:
1. `python3 /usr/local/lib/node_modules/openclaw/skills/skill-creator/scripts/init_skill.py <name> --path skills/ --resources references`
2. Write SKILL.md + resources
3. Test thoroughly
4. Validate: `python3 .../quick_validate.py skills/<name>`
5. Commit and push

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger phrases are broad and generic (e.g. learning, growing, improving, tracking, progress), which can cause the skill to activate in many unrelated contexts. Unintended activation is risky here because the skill encourages memory writes, research actions, and operational commands, so a benign conversation could spur unintended state changes or tool use.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.