Back to skill

Security audit

微博热搜采集 | Weibo Hot Search

Security checks for vulnerabilities and agentic risk

Overview

This skill does what it claims at a high level, but it contains exploitable command and HTML injection risks that users should review before installing.

Install only if you are comfortable running a Weibo scraper that uses your OpenClaw browser profile and writes local data. Before using --with-content or opening generated reports from fresh scraped data, the publisher should remove shell=True, strictly validate Weibo URLs, avoid innerHTML for remote values, and document exact network and file-write scope.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/fetch-hot-search.py:126
Finding

Shell Command Injection Through a Remotely Supplied Topic URL

Content
View full analysis

Vulnerability Details

File Location: scripts/fetch-hot-search.py, lines 16-19 and 126-129
Vulnerability Type: OS command injection
Risk Level: High

Vulnerable Code

python
def run_command(cmd, timeout=30):
    try:
        result = subprocess.run(
            cmd,
            shell=True,
            capture_output=True,
            text=True,
            timeout=timeout
        )
        return result.returncode == 0, result.stdout, result.stderr
    except Exception as e:
        return False, "", str(e)
python
success, _, _ = run_command(
    f"openclaw browser open --profile openclaw '{url}'",
    timeout=15
)

Technical Analysis

The run_command() function passes a string to subprocess.run() with shell=True. In fetch_topic_content(), the topic URL is interpolated directly into that shell command.

The URL originates in a browser snapshot of remote Weibo content. The parser only verifies that the extracted value contains the substring s.weibo.com; it does not establish that the entire value is a valid URL with an exact approved hostname, nor does it prevent shell metacharacters such as apostrophes, semicolons, command substitutions, or comment characters.

Although the URL is enclosed in single quotes, an attacker-controlled apostrophe can terminate that quoted argument. Subsequent shell syntax is then interpreted by the operating system shell. A malicious snapshot value conceptually shaped like the following would pass the substring check while breaking out of the command argument:

text
https://s.weibo.com/topic';malicious_command;# 

The vulnerability is reached when detailed topic collection is enabled through --with-content.

Attack Path

  1. An attacker causes maliciously structured link data to appear in content represented by the OpenClaw browser snapshot.
  2. parse_hot_search() extracts the attacker-controlled /url v ...[truncated 959 chars]
Remediation
View remediation

Remediation Suggestions

Remove shell interpretation entirely. Change the command helper to accept an argument list and invoke subprocess.run() with its default shell=False behavior:

python
def run_command(args, timeout=30):
    try:
        result = subprocess.run(
            args,
            shell=False,
            capture_output=True,
            text=True,
            timeout=timeout,
            check=False,
        )
        return result.returncode == 0, result.stdout, result.stderr
    except Exception as e:
        return False, "", str(e)

success, _, _ = run_command(
    [
        "openclaw",
        "browser",
        "open",
        "--profile",
        "openclaw",
        url,
    ],
    timeout=15,
)

Validate the URL before invoking OpenClaw:

  • Parse it with urllib.parse.urlparse().
  • Permit only https.
  • Require an exact approved hostname, such as s.weibo.com; do not use substring matching.
  • Reject embedded credentials, control characters, malformed ports, and unexpected hostname suffixes.
  • Normalize protocol-relative URLs before validation.
  • Apply the same non-shell execution pattern to all other OpenClaw commands, even where their current arguments are constants.
  • Add regression tests containing apostrophes, semicolons, command substitutions, newlines, and malformed hostnames.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/generate_html.py:430
Finding

Stored HTML and JavaScript Injection in the Generated Report

Content
View full analysis

Vulnerability Details

File Location: scripts/generate_html.py, lines 371 and 430-449
Vulnerability Type: Stored cross-site scripting and generated HTML injection
Risk Level: High

Vulnerable Code

Remotely sourced database records are embedded directly into an executable script context:

python
<script>
    const items = {json.dumps(items, ensure_ascii=False)};
    const channelColors = {json.dumps(CHANNEL_COLORS, ensure_ascii=False)};
    let selectedChannel = '';

The same values are subsequently concatenated into HTML and assigned through innerHTML:

javascript
let html = '';
Object.values(grouped)
    .sort((a, b) =>
        b.date.localeCompare(a.date) ||
        a.channel_id.localeCompare(b.channel_id)
    )
    .forEach(group => {
        const bgColor = channelColors[group.channel_id] || '#666';
        html += group.items.map(item => `
            <a href="${fixUrl(item.url)}" class="hot-item" target="_blank">
                <div class="rank ${getRankClass(item.rank)}">${item.rank}</div>
                <div class="item-content">
                    <div class="item-title">${item.title}</div>
                    <div class="item-meta">
                        <span class="channel-badge"
                              style="background: ${bgColor}">${group.channel}</span>
                        ${item.heat
                            ? `<span class="heat-badge">🔥 ${formatHeat(item.heat)}</span>`
                            : ''}
                        ${item.tag
                            ? `<span class="tag-badge">${item.tag}</span>`
                            : ''}
                        <span style="color: #999;">${item.date}</span>
                    </div>
                </div>
            </a>
        `).join('');
    });

container.innerHTML = html;

Technical Analysis

Titles, URLs, tags, channel names, and da ...[truncated 2698 chars]

Remediation
View remediation

Remediation Suggestions

Avoid constructing markup from untrusted strings. Create elements through DOM APIs and place text into textContent:

javascript
const title = document.createElement('div');
title.className = 'item-title';
title.textContent = item.title;

const link = document.createElement('a');
link.className = 'hot-item';
link.target = '_blank';
link.rel = 'noopener noreferrer';
link.href = validatedUrl;
link.appendChild(title);

container.replaceChildren(link);

Apply the following hardening measures:

  • Validate URLs before storing them and again before rendering them.
  • Permit only https URLs with an exact approved hostname such as s.weibo.com.
  • Reject javascript:, data:, file:, and other unexpected schemes.
  • Do not rely on innerHTML for rendering remote values.
  • If HTML insertion is unavoidable, use a well-maintained sanitizer with a minimal allowlist.
  • When embedding JSON in inline script, escape characters significant to HTML parsing, including less-than signs, greater-than signs, ampersands, U+2028, and U+2029. Prefer loading JSON from a separate non-executable resource.
  • Add a restrictive Content Security Policy that disallows inline script and event handlers.
  • Add rel="noopener noreferrer" to links opened with target="_blank".
  • Regenerate data/index.html after fixing scripts/generate_html.py.
  • Add tests using script terminators, image error handlers, quote characters, malicious URL schemes, and malformed links.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (17)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill makes strong claims about multi-channel automated collection, persistence, and HTML reporting, but the behavior described by analysis suggests those features are not actually implemented and may rely on manual extraction or placeholders. This mismatch is dangerous because it can mislead operators into granting trust, credentials, or execution privileges to a skill whose real behavior is different from its declared behavior, reducing auditability and increasing the chance of unsafe use.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill makes strong claims about multi-channel automated collection, persistence, and HTML reporting, but the behavior described by analysis suggests those features are not actually implemented and may rely on manual extraction or placeholders. This mismatch is dangerous because it can mislead operators into granting trust, credentials, or execution privileges to a skill whose real behavior is different from its declared behavior, reducing auditability and increasing the chance of unsafe use.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
97% confidence
Finding

This is a concrete tool-parameter abuse issue because command strings are composed and executed through the shell. Since fetch_topic_content opens URLs extracted from remote page snapshots, an attacker controlling or influencing those values could inject shell syntax and execute arbitrary commands on the host running the skill.

Content

Scanner excerpt · scripts/fetch-hot-search.py (reported line 18)May include surrounding context.

python
def run_command(cmd, timeout=30):
    try:
        result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout)
        return result.returncode == 0, result.stdout, result.stderr
    except Exception as e:
        return False, "", str(e)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill documents capabilities that require shell execution, network access, and file writes, but it does not declare any explicit tool scope or permissions boundary. In an agent environment, that creates an authorization gap where a reviewer or runtime may not understand or constrain what the skill can do, increasing the chance of unintended network access, local file modification, or command execution.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes collecting Weibo hot-search data and producing persisted/HTML outputs, but this file implements a general shell-command wrapper using subprocess.run(..., shell=True). Spawning shell commands is a broader execution capability than necessary for the stated data-collection purpose, even if currently used to drive an external browser tool.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
95% confidence
Finding

The helper executes arbitrary shell strings with shell=True, which is dangerous because any value interpolated into the command becomes shell-interpreted. In this file, URLs derived from page content are later embedded into commands, so a crafted or unexpected URL could trigger command injection and arbitrary local command execution.

Content

Scanner excerpt · scripts/fetch-hot-search.py (reported line 18)May include surrounding context.

python
def run_command(cmd, timeout=30):
    try:
        result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=timeout)
        return result.returncode == 0, result.stdout, result.stderr
    except Exception as e:
        return False, "", str(e)

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest emphasizes multi-channel hot-search data collection, persistence, and HTML reporting. This function opens each topic URL and extracts underlying post content, which is a materially broader behavior than collecting hot-search榜单 entries themselves.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes multi-channel hot search data collection with persistence and visualization support, but this script does not actually collect any hot search data. It launches a browser to the generic hot search page, instructs the user to manually extract data, and saves a placeholder file containing no items.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest says the skill supports hot search, social, entertainment, and life channels simultaneously, but this script only opens https://weibo.com/hot/search and later writes "channel": "hot". There is no code path for fetching or even selecting the other declared channels.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The generated page is explicitly marked as lang="zh-CN" and uses a Chinese-only title and UI labels, which imposes a specific language/locale choice on users. The file does not provide any opt-in, alternate locale, or documentation that this is intentionally limited to a region-specific audience.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/save_to_db.py (reported line 25)May include surrounding context.

python
try:
        # 运行Python脚本获取数据
        result = subprocess.run(
            [sys.executable, str(script_path), 
             '--limit', str(limit),
             '--output', '-'],  # 输出到stdout

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The document declares lang="zh-CN", and the visible UI text throughout the page is exclusively Chinese. Under the policy, forcing a specific language or locale without user opt-in can be a natural-language policy violation unless the locale constraint is explicitly justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The code opens remote Weibo URLs via browser automation and snapshots page content, and with --with-content it also fetches detailed topic posts. While this behavior is central to the tool's purpose, the code provides no explicit warning that it will connect to external services and retrieve content from them.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

This code creates parent directories and writes fetched data to a user-specified path, but there is no confirmation prompt or explicit warning near the write operation. Although file output is a normal feature of the script, the code itself does not disclose that it will create directories and write a file when --output is used.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The top comments label the file as a Weibo Hot Search Fetcher and say it uses OpenClaw to scrape Weibo hot search, which implies actual data retrieval. In practice, the script merely opens a browser page and emits instructions plus placeholder JSON, so the documentation overstates what the code does.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The file-level description and all user-facing text are written in Chinese, indicating the tool is designed around a fixed language/locale. Under the policy, forcing a specific language without opt-in can be a natural-language policy violation when no alternative or choice is offered.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module title and description are written only in Chinese/Chinese-mixed text, and all user-facing console output in the file is likewise Chinese. This creates a language/locale constraint without any visible opt-in or alternative, which matches the policy category for forced language use.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.