Back to skill

Security audit

Openclaw Plugin

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent security scanner, but it needs review because it can route intercepted messages to LLM classifiers, persists classifier-derived metadata, has fail-open security behavior, and uses an unpinned install dependency.

Review this before installing in any environment that handles private prompts, credentials, customer data, or high-privilege agents. Prefer a pinned and audited hopeid version, explicitly decide whether LLM-based classification may receive message content, disable or constrain external classifier routes if needed, and do not rely on strict blocking or Telegram alerts until the fail-open and implementation mismatch issues are fixed.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
index.ts:429
Finding

Fail-Open Initialization Race Allows Threat-Scanning Bypass

Content
View full analysis
{ quarantine = mgr; }).catch(err => { api.logger.warn(`[hopeIDS] Quarantine init warning: ${err.message}`); }); ``` ```ts const record = await quarantine!.create({ ts: new Date().toISOString(), agent: agentId, source: event.source ?? 'unknown', senderId: event.senderId, intent: intent || 'unknown', risk, patterns, contentHash: hashContent(event.prompt), }); ``` ```ts } catch (err: any) { api.logger.warn(`[hopeIDS] Scan failed: ${err.message}`); } ``` ### Technical Analysis The quarantine manager is initialized asynchronously without being awaited before the `before_agent_start` scanning hook becomes operational. The TypeScript non-null assertion in `quarantine!.create(...)` suppresses compile-time null checks but provides no runtime protection. If a message reaches the block branch before initialization completes, `quarantine` remains `null`. Calling `create` then raises a runtime exception. The outer catch block merely logs the exception and returns no blocking response. The host can consequently continue normal agent processing. The same fail-open behavior applies to other exceptions raised during heuristic scanning, semantic classification, or quarantine storage. This contradicts the documented security invariant that a blocked message must be fully aborted. ### Attack Path 1. The OpenClaw process starts and registers the plugin. 2. Quarantine initialization begins asynchronously. 3. Before initialization completes, an attacker submits a malicious prompt that exceeds the blocking threshold. 4. The scan enters the blocking branch. 5. `quarantine!.create(...)` attempts to access the still-null manager and throws. 6. The catch block suppre ...[truncated 735 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
index.ts:321
Finding

Attacker-Influenced Classifier Output Is Persisted as Trusted Metadata

Content
View full analysis
`• ${p}`).join('\n') : '• (no pattern metadata)'; ``` ### Technical Analysis The message being classified is attacker-controlled. Both supported LLM classifiers can generate free-form `red_flags` strings based on that message. Those strings are appended directly to `patterns` and persisted in the quarantine record without an allowlist, redaction, normalization, control-character removal, or length limits. Although the project states that quarantine contains metadata only, model-generated red flags may quote or reproduce parts of the original message. A crafted prompt can therefore induce a classifier to return embedded secrets, malicious instructions, arbitrary multiline strings, or other raw content as a red flag. The same values are incorporated into alert text and written through the logger. Consequently, attacker-influenced content can cross the intended message-to-metadata trust boundary and enter persistent quarantine storag ...[truncated 1263 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:253
Finding

Unpinned Third-Party Package Is Downloaded and Executed During Installation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (10)

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

The installation instruction uses npx hopeid setup without pinning a specific package version. This allows consumers of the skill to fetch and execute whatever version is current on npm at install time, increasing supply-chain risk if a future release is compromised, malicious, or unexpectedly changed. In a security-focused skill, this is more dangerous because users are especially likely to trust and run the command as-is.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The header comment advertises 'Telegram alerts for blocked messages' and states that alerts are metadata-only Telegram notifications. In implementation, the blocked-message path merely registers a dummy gateway method and writes the alert content to the logger instead of sending any Telegram message, which directly contradicts the stated behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code forwards message content to the external 'llm-task' classification path, which may route data to third-party model providers, without any explicit user-facing disclosure or consent at the call site. In a security-scanning plugin, intercepted prompts may contain secrets, personal data, or sensitive operational content, so undisclosed transmission increases privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The plugin forwards user message content to a classifier agent for analysis without a clear disclosure boundary. Even if described as sandboxed, the classifier still receives raw user input, which may include confidential or regulated data, making this a real data-exposure concern in the context of an interception/security-monitoring component.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The manifest says the skill will "Automatically scan messages before agent processing" but does not clearly define which messages are included, excluded, or under what contexts scanning occurs. For a manifest file, this is a vague activation condition because it suggests broad triggering without explicit scope constraints or negative examples.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest enables Telegram alerts for blocked messages by default and references use of a Telegram bot/chat without a clear user-facing disclosure about what message content or metadata may be transmitted externally. In a security-scanning plugin, blocked content may itself contain sensitive prompts, credentials, or personal data, so implicit forwarding can create an avoidable data-leak path.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest supports LLM-based semantic analysis through a configurable external endpoint, but it does not clearly warn that user messages may be transmitted to third-party or remote services for classification. Because this plugin analyzes potentially sensitive inbound content, sending that content to an LLM endpoint can expose private data, secrets, or regulated information beyond the local trust boundary.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The '/trust' command is described as 'Trust a sender (whitelist for future messages)', but the implementation itself includes a TODO indicating the trusted-sender list is not properly added/persisted. Because the comment and user-facing description promise a durable whitelist effect that the code does not fully implement, this is an intent-code divergence.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The dependency uses a caret range (^0.1.0), which allows installation of newer releases within the compatible semver range. That creates supply-chain risk because builds are not reproducible and a later compromised or breaking package version could be pulled in without review.

Content

Scanner excerpt · package.json (reported line 10)May include surrounding context.

json
"extensions": ["."]
  },
  "dependencies": {
    "hopeid": "^0.1.0"
  },
  "peerDependencies": {
    "openclaw": ">=2025.0.0"

Unverifiable Dependency: openclaw has 16 known advisory(ies) (CVE-2026-53846 (OpenClaw: Workspace .env npm_execpath could influence bundled runtime dependency); CVE-2026-32064 (OpenClaw's andbox browser noVNC observer lacked VNC authentication); CVE-2026-32006 (OpenClaw has a BlueBubbles group allowlist mismatch via DM pairing-store fallbac) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

The peer dependency on openclaw is expressed as a broad minimum version (>=2025.0.0), so consumers may install vulnerable OpenClaw versions that satisfy the range. Because the manifest does not constrain to a known-safe release, the plugin cannot assure compatibility only with patched versions despite known advisories in that ecosystem.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.