Back to skill

Security audit

Skill Evolution Loop

Security checks for vulnerabilities and agentic risk

Overview

This skill is a self-evolution tool that fits its stated purpose, but it scans private session history, creates persistent skills from chat text, and can delete skill directories with weak safeguards.

Review before installing. Use only in an environment where the operator explicitly wants local session history scanned for automation opportunities. Treat generated skills as drafts, inspect them before activation, avoid running non-dry-run gc until deletion has confirmation or quarantine, and do not rely on the documented protected cron-job safeguard unless enforcement is added.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T02 · Agent Memory Poisoning

Error
Location
engine.py:172
Finding

Persistent Skill Instruction Injection Through Untrusted Session Content

Content
View full analysis
100: continue # Skip system-formatted lines if re.match(r'^\[|^-|^\\*|^#|^\\||^>|^```|^\\{', line): continue if re.search(r'[0-9a-f]{8}-[0-9a-f]{4}', line): continue # Check for a verb and object has_verb = any(v in line for v in task_verbs) has_obj = any(o in line for o in task_objects) if has_verb and has_obj: task = line[:50].strip() date_str = datetime.fromtimestamp(f.stat().st_mtime).strftime("%Y-%m-%d") if task not in task_patterns: task_patterns[task] = { "count": 0, "dates": [], "examples": [], "has_tool": True } task_patterns[task]["count"] += 1 task_patterns[task]["dates"].append(date_str) if len(task_patterns[task]["examples"]) < 3: task_patterns[task]["examples"].append(task) ``` ```python def generate_skill_md(task_name, task_info): count = task_info.get("count", 1) dates = ', '.join(set(task_info.get("dates", []))) return f'''--- name: {task_name[:40]} description: | Automated task: {task_name} Frequency: {count} | Dates: {dates} triggers: - "{task_name}" - "auto-{task_name[:20]}" --- # {task_name} ## When to Use - When the user requests "{task_name}" or expresses a similar intent - After observing the repeated request {count} times ## Execution Process 1. Understand the user's intent and input parameters 2. Prepare the execution environment and dependencies 3. Perform the core operation 4. Verify the output 5. Report the result to the user ''' ``` ### Technical Analysis User-controlled text is read from historical session logs and interpolated directly into YAML frontmatter an ...[truncated 1944 chars]
Remediation
View remediation

other

Warning
Location
engine.py:128
Finding

Excessive Collection and Persistence of Private Session Content

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
engine.py:366
Finding

Unsafe Recursive Deletion Across the Global Skill Directory

Content
View full analysis
= max_age_days: is_referenced = False for other_dir in SKILLS_DIR.iterdir(): if other_dir == skill_dir or not other_dir.is_dir(): continue other_file = other_dir / "SKILL.md" if other_file.exists() and skill_dir.name in other_file.read_text(): is_referenced = True break gc_candidates.append({ "name": skill_dir.name, "age_days": age_days, "path": str(skill_dir), "is_referenced": is_referenced, "last_modified": mtime.isoformat() }) ``` ```python if not dry_run: removed = [] for c in gc_candidates: if not c["is_referenced"]: import shutil shutil.rmtree(c["path"]) removed.append(c["name"]) log(f"Removed: {c['name']} after {c['age_days']} inactive days") ``` ### Technical Analysis Garbage collection scans the entire shared `SKILLS_DIR` and determines ownership using only a plain-text marker inside `SKILL.md`. This marker is forgeable and is not tied to a trusted registry, file identity, signature, or dedicated managed directory. The age calculation uses the `SKILL.md` modification time as a substitute ...[truncated 1561 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
engine.py:112
Finding

Protected Cron Job Safeguard Is Documented but Not Enforced

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (13)

Vague Triggers

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The trigger 'gc' is extremely short and ambiguous, making accidental activation highly likely in routine technical conversations where 'gc' may refer to garbage collection generally. Because this skill includes a gc command that can mark or delete unused skills, a false activation could lead to unintended destructive maintenance actions.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The GC path performs irreversible deletion with shutil.rmtree when invoked non-dry-run, without any explicit confirmation, quarantine period, or recovery mechanism. In this skill's context, the engine manages user skills under ~/.hermes/skills, so a mistaken or automated invocation can destroy working automation assets and create a denial-of-service against the agent's capabilities.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill documents file-reading and file-writing behavior through commands and operational workflows, but it does not declare any explicit tool scope such as permissions or allowed-tools. That makes the runtime boundary unclear and increases the risk that an agent may invoke filesystem-capable tools more broadly than intended, especially in an automation-oriented skill that scans logs, writes reports, and performs garbage-collection actions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The visible natural-language description and instructions are written exclusively in Chinese, and the file does not offer an opt-in language choice or explain a required locale scope. Under the policy, language constraints should be optional or clearly justified when the skill is region-specific.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrase '技能进化' is broad enough to appear in normal discussion, which can cause accidental activation of a skill that performs scanning, report generation, and potentially destructive follow-on actions. In this context, overbroad invocation is more dangerous because the skill is tied to autonomous workflow evolution and eventual deletion logic.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The English trigger 'skill evolution' is generic and could match benign discussion about improving skills rather than a deliberate request to invoke this automation engine. Given that the skill orchestrates distillation and lifecycle management, broad activation language increases the chance of unintended entry into a privileged workflow.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill presents 'manual confirmation' and 'protected cron job skipping' as safety controls, but the same document also exposes direct commands for run and gc that can materially change or delete skills. This creates a misleading safety model: a user or agent may over-trust the stated safeguards even though destructive actions remain one command away, weakening defense-in-depth for an engine that evolves and removes artifacts automatically.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The module docstring, generated user-facing content, logs, reports, and CLI help are all written exclusively in Chinese, which imposes a specific language/locale by default. There is no indication that users can choose another language or that the locale restriction is intentional and justified for a region-specific tool.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The module docstring states that the elimination mechanism targets automatically generated skills that have not been triggered for 30 days. In code, run_gc determines staleness from SKILL.md modification time and whether the skill name appears in other skill files, which is materially different from measuring actual trigger or execution activity.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The docstring for run_gc says it will mark/delete automatically distilled skills that were not triggered for a threshold number of days. However, the code computes age solely from skill_file.stat().st_mtime and never consults any trigger history, execution logs, or invocation counters, so the documented intent contradicts the actual behavior.

Content

No source excerpt is available for this finding.

Overly Broad Trigger

Low
Category
Trigger Abuse
Confidence
97% confidence
Finding

The trigger 'gc' is too short and context-ambiguous, which substantially raises the probability of unintended activation. In a skill that includes cleanup and deletion semantics, accidental triggering can cause harmful maintenance actions or initiate a workflow that the user did not intend to run.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

This markdown file uses Chinese exclusively for the title and status text, and there is no indication that the user opted into Chinese or that the skill is intentionally limited to a Chinese-language audience. The policy requires flagging language or locale constraints when they are imposed without user choice or clear justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file presents all user-facing content exclusively in Chinese, and there is no indication that the skill or document is intentionally limited to Chinese-speaking users or a specific region. Under the language/locale policy, forcing a specific language without user opt-in can be a policy violation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.