Back to skill

Security audit

Skill Perf

Security checks for vulnerabilities and agentic risk

Overview

This looks like a real OpenClaw performance-measurement skill, but it reads broad local OpenClaw state and automatically exposes generated reports through an unauthenticated local web server.

Install only if you are comfortable with this skill spawning OpenClaw subagents, reading local OpenClaw session/workspace metadata, persisting benchmark reports, and creating an unauthenticated report server. Prefer reviewing or patching the report server to bind only to 127.0.0.1, serve only the current report, shut down predictably, and make transcript/bootstrap analysis opt-in with redaction.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/snapshot.py:1525
Finding

Generated reports are served on all network interfaces without access control

Content
View full analysis
/...` URL and the function description, both of which imply that the report is available only locally. The HTTP server also serves the entire parent report directory, enabling unauthenticated directory listing and access to other reports stored there. The comment states that the process exits after 30 seconds, but no timeout, process termination, or lifecycle-management code is implemented. The child process can consequently remain available until explicitly terminated or reclaimed by the operating system. The use of `getsockname()` only obtains an ephemeral port and do ...[truncated 1590 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/snapshot.py:822
Finding

Report generation reads sensitive Agent state and complete session transcripts beyond token-measurement needs

Content
View full analysis
Optional[Path]: """从 openclaw.json 读取指定 agent 的 workspace 路径""" cfg_path = OPENCLAW_DIR / "openclaw.json" if not cfg_path.exists(): return None try: cfg = json.loads(cfg_path.read_text(encoding="utf-8")) agents = cfg.get("agents", {}).get("list", []) for a in agents: if a.get("id") == agent_id: ws = a.get("workspace", "") if ws: p = Path(ws).expanduser() return p if p.is_absolute() else OPENCLAW_DIR / p # 默认 workspace default_ws = cfg.get("agents", {}).get("defaults", {}).get("workspace", "") if default_ws: p = Path(default_ws).expanduser() return p if p.is_absolute() else OPENCLAW_DIR / p except Exception: pass return OPENCLAW_DIR / "workspace" def _scan_bootstrap_files(agent_id: str) -> list: """ 扫描 agent workspace 中的 bootstrap 文件,返回: [{"name": "AGENTS.md", "path": "...", "chars": 1234, "tokens": 456, "exists": True}, ...] """ ws = _get_agent_workspace(agent_id) result = [] for fname in _BOOTSTRAP_FILES: fpath = ws / fname if ws else None if fpath and fpath.exists(): try: text = fpath.read_text(encoding="utf-8", errors="replace") chars = len(text) tokens = estimate_file_tokens_text(text) result.append({ ...[truncated 4652 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (31)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

This mismatch is security-relevant because the documented purpose is narrow performance measurement, while the static findings indicate undeclared local data traversal and report-serving behavior. A skill presented as a benchmark tool but able to enumerate session artifacts, deleted logs, prompts, and configuration files can expose sensitive local context and operational metadata beyond user expectations.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

This mismatch is security-relevant because the documented purpose is narrow performance measurement, while the static findings indicate undeclared local data traversal and report-serving behavior. A skill presented as a benchmark tool but able to enumerate session artifacts, deleted logs, prompts, and configuration files can expose sensitive local context and operational metadata beyond user expectations.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

This mismatch is security-relevant because the documented purpose is narrow performance measurement, while the static findings indicate undeclared local data traversal and report-serving behavior. A skill presented as a benchmark tool but able to enumerate session artifacts, deleted logs, prompts, and configuration files can expose sensitive local context and operational metadata beyond user expectations.

Content

No source excerpt is available for this finding.

Self-Modification

High
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · references/TOKEN_GUIDE.md (reported line 170)May include surrounding context.

md
|------|------|----------|
| Tool list + descriptions | 所有可用工具的短描述 | ~4–6 KB |
| Skills list (metadata) | 技能清单(仅 metadata,指令按需加载) | ~3–5 KB |
| Self-update instructions | 自更新指引 | ~1 KB |
| Workspace bootstrap files | AGENTS.md, SOUL.md, TOOLS.md, IDENTITY.md, USER.md, HEARTBEAT.md, MEMORY.md 等 | ~8–12 KB |
| Time + reply tags + heartbeat | UTC/本地时间、回复标签、心跳行为 | ~1 KB |
| Runtime metadata | host/OS/model/thinking | ~0.5 KB |

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The README shows a very broad natural-language invocation pattern ('帮我测量 ... 的 token 消耗') and the metadata further states the skill should trigger whenever users mention measurement, performance, cost, efficiency, comparison, or optimization. In an agent environment, this can cause over-triggering on loosely related requests, leading the skill to spawn subagents and consume resources unexpectedly, increasing cost and creating opportunities for unintended execution paths.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill instructs the agent to read local skill files, spawn subagents, run shell scripts, generate reports, and expose a localhost HTTP report, but it declares no tool restrictions. In a skill system, missing scope declarations can cause the runtime to grant broader-than-expected access, increasing the blast radius if the skill is invoked unexpectedly or its inputs are attacker-controlled.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger terms are extremely broad and overlap with ordinary user requests about performance, cost, or efficiency, making accidental invocation more likely. Because this skill can spawn subagents, run shell commands, inspect local state, and generate network-accessible reports, overbroad triggering increases the chance that powerful actions occur in contexts where the user did not intend to run such a skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The title and the entire operational template are written as requirements in Chinese, and the document includes fixed output phrases and instructions without offering any language or locale choice. Under the policy rule, forcing a specific language is a natural-language policy concern unless the locale restriction is explicitly justified as region-specific.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/TOKEN_GUIDE.md (reported line 220)May include surrounding context.

cacheRetention: "short" # none | short | long

text

- `cacheRetention: "long"` + heartbeat `every: "55m"` → 保持 cache 热,降低 cache write 成本
- `cacheRetention: "none"` → 适合突发性通知 agent,避免无效 cache write
- 每个 agent 可独立配置:`agents.list[].params.cacheRetention`

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/TOKEN_GUIDE.md (reported line 220)May include surrounding context.

cacheRetention: "short" # none | short | long

text

- `cacheRetention: "long"` + heartbeat `every: "55m"` → 保持 cache 热,降低 cache write 成本
- `cacheRetention: "none"` → 适合突发性通知 agent,避免无效 cache write
- 每个 agent 可独立配置:`agents.list[].params.cacheRetention`

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/TOKEN_GUIDE.md (reported line 433)May include surrounding context.

md
| 百炼托管 qwen 系列 | `bailian` + model 含 `qwen` | **0.20×** | 百炼隐式缓存,折扣 80% |

> 📌 当前 `report_html.py` 的 `_detect_billing_model()` 将 `bailian` 提供商统一识别为 `bailian`(0.20×),
> 与 Kimi 原生(0.10×)区分开。如 `cacheWrite > 0` 说明走了显式缓存,应改用 0.10× 计算。

---

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/TOKEN_GUIDE.md (reported line 433)May include surrounding context.

md
| 百炼托管 qwen 系列 | `bailian` + model 含 `qwen` | **0.20x** | 百炼隐式缓存,折扣 80% |

> 📌 当前 `report_html.py` 的 `_detect_billing_model()` 将 `bailian` 提供商统一识别为 `bailian`(0.20x),
> 与 Kimi 原生(0.10x)区分开。如 `cacheWrite > 0` 说明走了显式缓存,应改用 0.10x 计算。

---

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

技能清单声明该 skill 会通过 OpenClaw 的 sessions_spawn 启动双 subagent 并发测量,并自动扣除系统底噪、生成 HTML 报告和置信度评级。但这里的主流程仅循环调用 run_once 顺序执行 runs 次测试,汇总后保存为 JSON;代码中也没有并发测量、噪声校正、HTML 生成或置信度计算逻辑。该实现与技能对外宣称的核心行为存在明显语义落差。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script persists benchmark results to disk and includes the full user-supplied task text in the saved JSON. If tasks contain secrets, internal URLs, credentials, customer data, or other sensitive prompts, this creates a local data exposure risk through later filesystem access, backups, indexing, or accidental sharing of result files.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · scripts/calibrate.py (reported line 16)May include surrounding context.

python
python3 calibrate.py show

    # 计算某个 skill 的底噪估算值
    python3 calibrate.py noise --skill-md ~/.openclaw/skills/html-extractor/SKILL.md
"""

import argparse

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script reads and analyzes detailed OpenClaw session internals, including skills lists, bootstrap files, system prompt composition, workspace-injected files, transcripts, and deleted session artifacts. In a skill whose stated purpose is token/performance measurement, this is excessive data access that can reveal sensitive prompt content, internal agent configuration, file paths, and historical conversation material unrelated to the requested benchmark.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/snapshot.py (reported line 1509)May include surrounding context.

python
# 报告成功生成,清理 pending 中的对应条目
    if PENDING_FILE.exists():
        try:
            plist = json.loads(PENDING_FILE.read_text(encoding="utf-8"))
            plist = [p for p in plist if p.get("session_key") != session_key]
            if plist:
                PENDING_FILE.write_text(json.dumps(plist, ensure_ascii=False, indent=2), encoding="utf-8")

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/snapshot.py (reported line 1510)May include surrounding context.

python
# 报告成功生成,清理 pending 中的对应条目
    if PENDING_FILE.exists():
        try:
            plist = json.loads(PENDING_FILE.read_text(encoding="utf-8"))
            plist = [p for p in plist if p.get("session_key") != session_key]
            if plist:
                PENDING_FILE.write_text(json.dumps(plist, ensure_ascii=False, indent=2), encoding="utf-8")

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/snapshot.py (reported line 1511)May include surrounding context.

python
# 报告成功生成,清理 pending 中的对应条目
    if PENDING_FILE.exists():
        try:
            plist = json.loads(PENDING_FILE.read_text(encoding="utf-8"))
            plist = [p for p in plist if p.get("session_key") != session_key]
            if plist:
                PENDING_FILE.write_text(json.dumps(plist, ensure_ascii=False, indent=2), encoding="utf-8")

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/snapshot.py (reported line 1512)May include surrounding context.

python
# 报告成功生成,清理 pending 中的对应条目
    if PENDING_FILE.exists():
        try:
            plist = json.loads(PENDING_FILE.read_text(encoding="utf-8"))
            plist = [p for p in plist if p.get("session_key") != session_key]
            if plist:
                PENDING_FILE.write_text(json.dumps(plist, ensure_ascii=False, indent=2), encoding="utf-8")

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill adds an unrelated capability to start a local HTTP server for report exposure, which is not necessary for token accounting. This increases attack surface and may disclose generated reports or filesystem-derived content through an ad hoc service that is not authenticated, monitored, or clearly scoped.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
89% confidence
Finding

The code launches a background local HTTP server via subprocess to expose generated HTML reports. Even though it uses a fixed argument list rather than shell interpolation, this expands the skill’s capabilities beyond passive measurement and can unintentionally expose report contents to other local users or network peers depending on host binding behavior and environment configuration.

Content

Scanner excerpt · scripts/snapshot.py (reported line 1544)May include surrounding context.

python
url = f"http://localhost:{port}/{filename}"

    # 后台启动 python -m http.server,30 秒后自动退出
    subprocess.Popen(
        ["python3", "-m", "http.server", str(port), "--directory", serve_dir],
        stdout=subprocess.DEVNULL,
        stderr=subprocess.DEVNULL,

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script's descriptive comments, parameter documentation, and user-facing status messages are presented in Chinese only. This forces a specific language for interaction and documentation without offering the user a language choice or documenting a justified locale constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file contains operational skill documentation exclusively in Chinese, and there is no indication that the skill is region-specific or that users may choose another language. Under the policy, forcing a specific language without opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file title and all instructional content are written exclusively in Chinese, with no indication that users may choose another language or that the document is intended only for a Chinese-specific audience. Under the stated policy, forcing a specific language without opt-in is a natural-language locale policy violation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.