Back to skill

Security audit

self-evalutaed-agent

Security checks for vulnerabilities and agentic risk

Overview

This skill is not clearly malicious, but it should be reviewed carefully because it can automatically create persistent agent-facing tasks from workspace logs without clear approval or input-safety boundaries.

Install only if you are comfortable with an automated self-improvement workflow that writes persistent backlog and memory entries. Keep cron disabled until you have reviewed the scripts, restrict who can write error logs, require human approval before an Agent PM executes generated backlog tasks, and avoid storing secrets in procedural commands.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T02 · Agent Memory Poisoning

Error
Location
scripts/topic_selector.py:76
Finding

Persistent Agent Memory Poisoning Through Unsanitized Error Log Content

Content
View full analysis
0: error_types = error_analysis.get('error_types', {}) if 'cron_timeout' in error_types: return { 'topic': "Fix cron timeout issues - optimize execution", 'reason': f"Cron timeout errors detected ({error_types.get('cron_timeout')}x)", 'priority': 9, 'source': 'error-driven' } top_errors = error_types if top_errors: top_error = list(top_errors.keys())[0] return { 'topic': f"Error handling for {top_error}", 'reason': f"High error rate: {top_errors[top_error]} occurrences", 'priority': 8, 'source': 'error-driven' } ``` From `scripts/self_improvement_cycle.py:92-143`: ```python content = f"""# Self-Improvement Research **Generated:** {datetime.now().isoformat()} **Source:** Log analysis **Priority:** {topic['priority']}/10 ## Topic {topic['topic']} ## Why {topic['reason']} ## Key Improvements Needed Based on log analysis: 1. Analyze and fix root cause 2. Add monitoring for this issue 3. Verify fix works ## Backlog Tasks - [ ] 1. {topic['topic']} - type: implementation - priority: {topic['priority']} - impact: {topic['priority']}/10 - source: log_analysis - status: backlog """ with open(research_file, 'w') as f: f.write(content) # Add to backlog today = datetime.now().strftime('%Y-%m-%d') backlog_file = o ...[truncated 3098 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/topic_selector.py:20
Finding

Predictable Temporary File Allows Symbolic-Link File Overwrite

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (14)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README describes a system that 'learns from its own errors' and 'implements improvements' but does not warn users that it may modify code, scripts, or operational state. For an autonomous agent skill, undisclosed self-modification behavior is dangerous because users may install or schedule it without understanding that it can make persistent changes to their environment.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
84% confidence
Finding

The feature description indicates autonomous decision-making by automatically running on detected errors and selecting high-priority topics. While not inherently malicious, autonomous execution in a self-improvement context is risky because it can chain decisions into actions affecting the workspace without human validation.

Content

Scanner excerpt · README.md (reported line 7)May include surrounding context.

md
## Features

- **Auto-Trigger**: Automatically runs when new errors are detected
- **Topic Selection**: Analyzes error patterns and selects high-priority topics
- **Impact Measurement**: Records before/after metrics to measure improvement effectiveness
- **Procedural Memory**: Remembers working commands/scripts between sessions

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README states that the skill 'automatically runs when new errors are detected' but does not define scope, approval requirements, or safety boundaries. In a self-improving agent, ambiguous auto-triggering can cause unexpected execution and follow-on actions without operator awareness, increasing the chance of unsafe or unauthorized changes.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
86% confidence
Finding

The skill advertises autonomous decision making via automatic execution and topic prioritization from observed errors. In context, this is more dangerous because the architecture chains autonomous detection into research, backlog generation, and execution-oriented workflows, creating a path from passive monitoring to active system changes.

Content

Scanner excerpt · SKILL.md (reported line 11)May include surrounding context.

md
## Features

- **Auto-Trigger**: Automatically runs when new errors are detected
- **Topic Selection**: Analyzes error patterns and selects high-priority topics
- **Impact Measurement**: Records before/after metrics to measure improvement effectiveness
- **Procedural Memory**: Remembers working commands/scripts between sessions

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill states it will 'automatically run when new errors are detected' without defining scope, authorization boundaries, or what sources count as trusted triggers. Broad autonomous activation increases the chance of unintended execution, especially because the documented workflow proceeds from detection to research, backlog creation, and implementation-oriented actions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The feature list says the skill 'remembers working commands/scripts between sessions' but does not warn that this is persistent storage of operational history and commands. Persistent memory can retain sensitive commands, tokens in shell snippets, or unsafe procedures that may later be replayed without adequate review.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The 'What It Does' section describes generating research files, creating backlog tasks, and enabling downstream execution, but it does not clearly warn users that the skill can autonomously write files or create persistent artifacts. This can mislead operators about the level of autonomy and cause unreviewed changes or data persistence in sensitive workspaces.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill triggers subprocess execution automatically when it detects new errors, with no user confirmation, authorization check, or safety interlock. In operational environments, this can create an unsafe feedback loop where attacker-influenced log entries or benign failures repeatedly cause automated execution of a self-modifying or high-privilege recovery path.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/auto_trigger.py (reported line 66)May include surrounding context.

python
script_path = os.path.join(SCRIPTS_DIR, 'self_improvement_cycle.py')
    
    try:
        result = subprocess.run(
            ['python3', script_path],
            capture_output=True,
            text=True,

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The script automatically launches another Python program, giving it code-execution capability beyond passive log monitoring. In this skill context, that behavior is more dangerous because any compromise of the referenced script, configuration, or workspace turns mere error detection into automatic execution of additional code without further review.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The top-level docstring is written entirely in Russian, including the skill description, with no indication that another language is supported or that the locale is region-specific. Under the policy for all file types, this is a natural-language locale constraint because the file presents its purpose in a single language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module docstring includes a Russian-only description ('Универсальная версия - использует config.py') with no indication that users can choose language or locale. This creates a natural-language locale policy concern because the skill presents itself in a specific language without opt-in or documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The module’s natural-language description is written entirely in Russian and presents the skill behavior in that locale without any indication of user choice or region-specific justification. Under the policy, forcing a specific language without opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This Python file contains natural-language documentation entirely in Russian, including the primary module description. Under the policy criteria, forcing a specific language without user opt-in can be a locale/language policy issue, and no alternative language or opt-in mechanism is indicated in the file.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.