Back to skill

Security audit

superguard

Security checks for vulnerabilities and agentic risk

Overview

MoltGuard appears purpose-aligned as a security guard, but it installs an unpinned executable plugin and handles API credentials and remote security scanning with insufficient disclosure.

Review this carefully before installing, especially in shared, regulated, or sensitive environments. Treat the MoltGuard API key as a secret, avoid exposing /og_status or /og_claim output in logs or screenshots, and prefer a pinned or verified plugin release with clear documentation for remote data handling and credential storage.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:18
Finding

Unpinned Third-Party Plugin Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 18-22
Vulnerability Type: Unpinned executable dependency
Risk Level: Medium

bash
# Install the plugin
openclaw plugins install @openguardrails/moltguard

Technical Analysis

The Skill directs OpenClaw to download and install a third-party plugin without specifying an exact version or integrity digest. Consequently, the installed code is determined by the package registry at installation time rather than by the content reviewed in this audit.

The package contains only documentation and metadata, so the executable plugin implementation is not available for inspection. The version information is also inconsistent: the SKILL.md header declares version 1.0.0, while _meta.json declares version 6.8.16. This inconsistency further weakens traceability between the reviewed Skill and the executable dependency.

This is a supply-chain risk rather than evidence that the named dependency is currently malicious. A compromised publisher account, registry, package release, or dependency resolution process could replace the effective implementation after this Skill has been reviewed.

Attack Path

  1. An attacker compromises the package publisher, registry account, or release pipeline for @openguardrails/moltguard.
  2. The attacker publishes a modified release under the same package name or changes the version resolved by the unpinned installation command.
  3. A user or agent follows the Skill instructions and executes the installation command.
  4. OpenClaw downloads and installs the attacker-controlled release.
  5. The plugin executes with the filesystem, network, credential, and process privileges available to the OpenClaw plugin runtime.

Impact Assessment

Successful exploitation could permit arbitrary code execution within the OpenClaw process context. The resulting scope depends on the plugin sandbox and operating-system account, but could in ...[truncated 248 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin installation to a specific, reviewed plugin version rather than resolving the latest release.
  • Verify the downloaded artifact using a cryptographic integrity hash or signed provenance.
  • Reconcile the version declared in SKILL.md with the version in _meta.json.
  • Require explicit user confirmation before downloading, installing, or updating executable plugins.
  • Review the plugin source and transitive dependencies for the exact pinned release.
  • Run the plugin in a restricted sandbox with only the filesystem and network permissions required for detection.
  • Configure updates to fail closed if signature, provenance, or integrity verification fails.

other

Warning
Location
SKILL.md:82
Finding

Potential Disclosure of Sensitive Content to an External Detection Service

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 82-89
Vulnerability Type: Underspecified remote processing of security-sensitive content
Risk Level: Medium

text
All security detection is performed by Core:

**Core Risk Surfaces:**
1. **Prompt / Instruction Risk** — Prompt injection, malicious email/web instructions, unauthorized tasks
2. **Behavioral Risk** — Dangerous commands, file deletion, risky API calls
3. **Data Risk** — Secret leakage, PII exposure, sending sensitive data to LLMs

**Core Technology:**
- **Intent-Action Mismatch Detection** — Catches agents that say one thing but do another

Technical Analysis

The Skill explicitly states that all security detection is performed by a remote service named Core. The inspected files do not document the service endpoint, exact payload fields, local redaction behavior, retention period, access controls, jurisdiction, or whether prompts, files, commands, secrets, and personally identifiable information are transmitted in full.

Remote analysis may be legitimate for the declared guardrail functionality, but sending security context to an external service creates a separate confidentiality boundary. The documentation does not establish that payloads are minimized to the least information necessary or that users provide informed consent before processing begins.

The onboarding instructions also state that an API key is obtained automatically and stored under ~/.openclaw/credentials/moltguard/. Local storage of a plugin-specific API key is consistent with authentication needs, but the documentation-only package does not allow verification of file permissions, secret redaction, transport security, or whether commands that display the key prevent accidental disclosure.

Attack Path

  1. A user installs and activates MoltGuard as instructed.
  2. OpenClaw processes a prompt, file, webpage, command, or action containing confidential material, P ...[truncated 1074 chars]
Remediation
View remediation

Remediation Suggestions

  • Clearly document every remote endpoint, transmitted field, processing purpose, retention period, and data-sharing policy.
  • Obtain explicit user consent before enabling remote inspection.
  • Perform local filtering and redaction before transmission, including removal of credentials, authentication tokens, private keys, and unnecessary PII.
  • Send only derived features or the smallest contextual excerpt required for detection.
  • Provide a local-only detection mode for sensitive or regulated environments.
  • Enforce authenticated TLS with strict certificate validation and restrict outbound connections to documented Core endpoints.
  • Store API credentials with restrictive owner-only permissions and avoid displaying complete keys in status or claim commands.
  • Redact secrets from application logs, telemetry, dashboards, error reports, and diagnostic output.
  • Publish implementation details or source code sufficient to verify payload minimization and credential handling.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (3)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The installation section says the plugin 'works immediately' but does not disclose that onboarding later states credentials are automatically obtained and stored under the user's home directory. That omission can mislead users into installing without understanding that authentication material will be created locally, increasing the risk of accidental exposure on shared or poorly secured systems.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The /og_status command is described as showing the API key along with account details, but there is no warning that invoking it may expose secrets in the terminal, logs, screenshots, recordings, or copied chat output. In an agent setting, encouraging a command that reveals secrets without a handling warning materially increases credential disclosure risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The claim-agent flow instructs the user to retrieve and paste an Agent ID and API Key into a web workflow without any warning about secure handling. This normalizes manual exposure of credentials and creates opportunities for shoulder surfing, clipboard leakage, browser compromise, or accidental sharing in transcripts and support channels.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.