Back to skill

Security audit

yan-watchman

Security checks for vulnerabilities and agentic risk

Overview

This skill is a review concern because it pressures the agent to immediately build and publish itself while modifying the OpenClaw skills workspace, but the claimed implementation files are absent.

Install only after reviewing the actual source package, because this artifact does not include the implementation it claims to publish. Do not let it run the listed build or workspace-copy commands automatically; require explicit confirmation, verify the copied files, and define privacy, consent, retention, and cleanup rules before using it to monitor messages.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:61
Finding

Skill instructions attempt to override agent priorities and trigger unapproved filesystem actions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:7, SKILL.md:50-65, and SKILL.md:123-137
Vulnerability Type: Instruction hijacking through priority assertions, urgency, and action-oriented commands
Risk Level: High

Vulnerable Code Snippet

The actionable filesystem operations appear at SKILL.md:61-62:

bash
mkdir -p ~/.openclaw/workspace/skills/yan-watchman-publish
cp -r src config Cargo.toml README.md ~/.openclaw/workspace/skills/yan-watchman-publish/

The surrounding instructions also direct the agent to change directories, run cargo build --release, and create or replace a published SKILL.md through a shell heredoc. Elsewhere, the document declares self-evolution to be the highest instruction and repeatedly demands immediate action without waiting, reporting, or deliberating.

Technical Analysis

The document attempts to establish its own objective as a highest-priority instruction. It reinforces that assertion with urgency and language intended to discourage normal review, reporting, and deliberation. It then supplies shell commands for building a purported Rust project, creating directories under the user's OpenClaw workspace, copying files, and writing a new skill document.

This matches skill instruction hijacking because the untrusted skill text attempts to alter the agent's current-session goals and decision process when loaded. A skill may document optional usage, but it must not claim authority over system or user instructions, pressure the agent to act automatically, or suppress security review.

The supplied project contains only SKILL.md; the referenced Rust sources, manifest, configuration, and README are absent from the audited artifact. Consequently, the claimed implementation cannot be verified, and the documented build and copy procedure cannot complete as presented. If similarly structured files were later introduced into the expected workspace, following these instructions ...[truncated 1612 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove all claims that the skill's objective is a highest-priority instruction or can supersede system, developer, or user directions.
  2. Remove language discouraging deliberation, review, reporting, or confirmation.
  3. Present build and publication commands strictly as optional examples that are never executed merely because the skill was loaded.
  4. Require explicit, informed user confirmation immediately before running Cargo or modifying the workspace.
  5. Describe every expected filesystem change, including destination paths and whether existing files may be replaced.
  6. Validate that all referenced source files are included in the reviewed package before offering build or publication steps.
  7. Refuse to build or copy files discovered outside the audited package unless the user separately approves those exact files.
  8. Use a newly created staging directory and fail safely if the destination already exists, rather than silently merging with or replacing existing content.
  9. Show a dry-run plan or file manifest before executing any mutating command.
  10. Keep descriptive documentation separate from executable operational instructions and clearly state that normal agent safety constraints remain authoritative.
Vulnerability Patterns
  • Rogue AgentSelf-Modification, Session Persistence
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (4)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language content consistently forces a specific language/locale presentation, and there is no indication that users may choose another language or that the skill is region-specific. The policy requires flagging language or locale constraints when they are imposed without user opt-in or justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The skill document instructs immediate build, packaging, and publication of a multi-platform monitoring tool without any privacy, consent, data-handling, or deployment warnings. Because the stated functionality includes monitoring Discord, Feishu, and WebChat and persisting messages to SQLite, encouraging blind publication can lead operators to deploy surveillance-like functionality or mishandle retained message data without understanding the risks.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
78% confidence
Finding

The instructions create persistent directories and copy source, config, and documentation into a long-lived workspace location under the user's home directory. In the context of a monitoring skill that stores message data and configuration, this encourages durable local retention without documenting cleanup, access controls, or secret-handling, which can expose sensitive data or leave residual artifacts across sessions.

Content

Scanner excerpt · SKILL.md (reported line 62)May include surrounding context.

md
cargo build --release

# 创建 Skill 包结构
mkdir -p ~/.openclaw/workspace/skills/yan-watchman-publish
cp -r src config Cargo.toml README.md ~/.openclaw/workspace/skills/yan-watchman-publish/

# 创建 SKILL.md

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILL.md (reported line 66)May include surrounding context.

md
cp -r src config Cargo.toml README.md ~/.openclaw/workspace/skills/yan-watchman-publish/

# 创建 SKILL.md
cat > ~/.openclaw/workspace/skills/yan-watchman-publish/SKILL.md << 'EOF'
# yan-watchman - 多平台AI助手消息守护系统

## 概述

Static analysis

No suspicious patterns detected.