Back to skill

Security audit

Hermes Agent

Security checks for vulnerabilities and agentic risk

Overview

This skill creates a local learning loop for OpenClaw and its persistence is disclosed, scoped, and aligned with that purpose.

Install only if you want OpenClaw to keep a durable local memory of operational lessons. Review ~/hermes-agent/ periodically, avoid storing sensitive details there, and use the built-in preference options if you want Hermes behavior only after failures, only on request, or disabled for certain projects.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger conditions are broad enough that the loop could activate during many ordinary interactions, causing memory reads and later writes without a clearly bounded scope. In an agent skill, ambiguous activation criteria increase the chance of unintended persistence behavior and hidden state accumulation across unrelated tasks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs the agent to record lessons into reflections.md and compress them into memory.md without any requirement for user notice, consent, or review. That creates a persistent memory channel that can silently store sensitive project details, user preferences, or erroneous instructions and then reuse them in future tasks.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · loop.md (reported line 51)May include surrounding context.

md
- one-off project hacks
- temporary incidents
- subjective style guesses without confirmation
- rules that only matter in one file or one hour

Static analysis

No suspicious patterns detected.