Back to skill

Security audit

Self Improving Agent

Security checks for vulnerabilities and agentic risk

Overview

This skill is a local conversation-quality logging and reporting helper with some privacy-relevant persistence, but no evidence of hidden execution, exfiltration, privilege escalation, or destructive behavior.

Before installing, treat this as a local persistence tool: do not log secrets, private customer data, credentials, or sensitive prompts as improvement insights, and verify where your OpenClaw workspace stores improvement_log.md. The skill appears appropriate for voluntary quality tracking, but its automatic-analysis wording should be clarified by the publisher.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (6)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill indicates file read/write capability through its examples and configuration, but it does not declare any explicit tool scope or permissions. This creates an authorization and transparency gap: hosts or users cannot easily determine that the skill may persist data locally, increasing the risk of unintended file access or modification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill describes learning logs and self-improvement features but does not warn users that conversation-derived insights may be written to local files such as improvement_log.md. This can lead to unintentional retention of sensitive prompts, personal data, or confidential operational details beyond the original conversation context.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The statement that the skill 'automatically analyzes conversations after each session' is overly broad and does not define trigger conditions, consent boundaries, or data handling limits. In practice, this ambiguity can cause the skill to activate on sensitive conversations unexpectedly and process or persist content without clear user intent.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a self-improving system that analyzes conversation quality, identifies opportunities, and continuously optimizes response strategies. In practice, the implementation only analyzes text heuristically, writes notes to local files, generates reports, and suggests possible SOUL.md updates; it does not modify strategies, apply improvements, or implement any optimization loop.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The class docstring states an active capability to improve performance over time. However, the methods only analyze conversations, append to an improvement log, produce a report, and suggest possible SOUL.md updates; no code applies changes to agent behavior or configuration.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill description is so broad that it does not define when the skill should activate, what inputs it is allowed to process, or what safety boundaries govern its self-improvement behavior. In an agentic system, vague scope increases the chance of unintended invocation, privilege overreach, or unsafe optimization behavior because downstream components may grant the skill wide latitude based only on its description.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.