Back to skill

Security audit

WeChat Article Extract

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does what it claims, but its optional image download can fetch arbitrary URLs from article HTML and save unbounded responses.

Install only if you are comfortable with a local script fetching WeChat pages and writing extracted output. Avoid using --download-images on untrusted saved HTML until the skill restricts image hosts, validates redirects, and enforces download size limits.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/extract_wechat_article.py:151
Finding

Server-Side Request Forgery Through Unrestricted Image Downloads

Content
View full analysis

Vulnerability Details

File Location: scripts/extract_wechat_article.py, lines 151–154 and 316–332
Vulnerability Type: Server-Side Request Forgery (SSRF) and unbounded response handling
Risk Level: High

Vulnerable Code

python
def normalize_image_url(value: Any) -> str:
    url = html.unescape(str(value or "").strip())
    if url.startswith("//"):
        url = f"https:{url}"
    return url if url.startswith("http") else ""
python
def download_images(record: dict[str, Any], output_dir: Path, timeout: int = 30) -> list[dict[str, str]]:
    output_dir.mkdir(parents=True, exist_ok=True)
    downloaded: list[dict[str, str]] = []
    for idx, entry in enumerate(record.get("imageEntries") or [], start=1):
        url = str(entry.get("sourceUrl") or "").strip()
        if not url:
            continue
        request = Request(url, headers={**DEFAULT_HEADERS, "Referer": "https://mp.weixin.qq.com/"})
        with urlopen(request, timeout=timeout) as response:
            body = response.read()
            suffix = guess_image_suffix(response.headers.get("Content-Type", ""), url)
        file_path = output_dir / f"{idx:02d}{suffix}"
        file_path.write_bytes(body)
        downloaded.append({"marker": entry.get("marker", ""), "path": str(file_path), "sourceUrl": url})
    return downloaded

Technical Analysis

Image locations are taken from article HTML and considered valid whenever their string begins with http. The code does not parse and strictly validate the scheme, restrict destination hostnames, resolve and classify destination IP addresses, constrain destination ports, or revalidate redirects.

The --html-file workflow allows a locally supplied, potentially attacker-controlled HTML document to populate imageEntries with arbitrary HTTP destinations. When --download-images is enabled, urlopen() requests each destination. Because Python's default URL opener follows redirects, even a superficially trusted destinati ...[truncated 1971 chars]

Remediation
View remediation

Remediation Suggestions

  1. Parse URLs with urlparse() and permit only exact http and https schemes.
  2. Prefer an explicit allowlist of required WeChat image CDN hostnames rather than accepting arbitrary hosts.
  3. Resolve each hostname and reject all loopback, private, link-local, multicast, reserved, and unspecified IPv4 and IPv6 addresses.
  4. Protect against DNS rebinding by connecting only to a validated resolved address while preserving the expected HTTP host and TLS hostname semantics.
  5. Disable automatic redirects or implement a redirect handler that repeats scheme, hostname, port, and resolved-IP validation for every hop.
  6. Reject embedded credentials, unexpected ports, malformed hostnames, and excessive redirect chains.
  7. Enforce maximum image count, per-image byte size, and aggregate download size.
  8. Stream responses in bounded chunks instead of calling unrestricted response.read().
  9. Require an approved image MIME type and consider verifying file signatures before writing.
  10. Add regression tests for localhost, RFC1918 addresses, IPv6 loopback, link-local metadata endpoints, encoded or unusual host representations, redirects to internal destinations, oversized responses, and non-image content.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract_wechat_article.py:41
Finding

Redirect-Based SSRF in WeChat Article Fetching

Content
View full analysis

Vulnerability Details

File Location: scripts/extract_wechat_article.py, lines 41–42 and 55–68
Vulnerability Type: Redirect-based Server-Side Request Forgery and unbounded response handling
Risk Level: Medium

Vulnerable Code

python
def is_public_wechat_article(url: str) -> bool:
    parsed = urlparse(str(url or "").strip())
    return parsed.scheme in {"http", "https"} and parsed.hostname == "mp.weixin.qq.com" and parsed.path.startswith("/s")
python
def fetch_html(url: str, timeout: int = 20, max_retries: int = 3, retry_delay: float = 1.0) -> dict[str, Any]:
    if not is_public_wechat_article(url):
        raise ValueError("URL must be a public mp.weixin.qq.com/s article link")

    clean_url = strip_tracking_params(url)
    last_error = ""
    status_code = 0
    for attempt in range(1, max(1, max_retries) + 1):
        try:
            request = Request(clean_url, headers=DEFAULT_HEADERS)
            with urlopen(request, timeout=timeout) as response:
                status_code = getattr(response, "status", 0) or 200
                charset = response.headers.get_content_charset() or "utf-8"
                body = response.read().decode(charset, errors="replace")
            return {"html": body, "source_url": clean_url, "status_code": status_code, "attempt": attempt}

Technical Analysis

The initial input validation restricts the hostname to mp.weixin.qq.com, which substantially reduces direct SSRF exposure. However, both HTTP and HTTPS are accepted even though the documented workflow specifies HTTPS, and urlopen() follows redirects by default.

Only the initial URL is validated. Redirect destinations are not checked for an approved scheme, hostname, port, or public IP address. Consequently, a valid initial WeChat URL that returns an attacker-influenced redirect could make the process request localhost, private-network services, or link-local endpoints.

The response body is also loaded entirely into memory with ...[truncated 1296 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require https for article URLs, matching the declared workflow.
  2. Disable automatic redirects or use a custom redirect handler.
  3. Revalidate every redirect target, including its scheme, hostname, port, and resolved IP addresses.
  4. If redirects outside mp.weixin.qq.com are unnecessary, reject them entirely.
  5. Reject loopback, private, link-local, multicast, reserved, and unspecified IPv4 and IPv6 destinations.
  6. Set a conservative redirect-hop limit.
  7. Stream the response in bounded chunks and stop once a configured maximum HTML size is reached.
  8. Validate the response content type before parsing.
  9. Add tests for cross-origin redirects, redirects to localhost and metadata addresses, HTTP-to-HTTPS policy, redirect loops, and oversized responses.
Vulnerability Patterns
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (4)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill invokes a local Python script that can read and write files, access the network, and execute via shell, but the manifest does not declare any tool scope or permission boundaries. This creates an overbroad and opaque execution surface: an agent or reviewer cannot easily constrain what the skill is allowed to do, and future changes to the bundled script could expand behavior without corresponding policy visibility.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_extract_wechat_article.py (reported line 39)May include surrounding context.

python
with tempfile.TemporaryDirectory() as tmpdir:
            html_path = Path(tmpdir) / "article.html"
            html_path.write_text(SAMPLE_HTML, encoding="utf-8")
            output = subprocess.check_output(
                [
                    sys.executable,
                    str(SCRIPT),

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_extract_wechat_article.py (reported line 65)May include surrounding context.

python
with tempfile.TemporaryDirectory() as tmpdir:
            html_path = Path(tmpdir) / "article.html"
            html_path.write_text(SAMPLE_HTML, encoding="utf-8")
            output = subprocess.check_output(
                [
                    sys.executable,
                    str(SCRIPT),

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The default request headers hard-code Accept-Language: zh-CN,zh;q=0.9, which imposes a specific locale preference in outbound requests. The file does not offer the user a language/locale option or explain why this locale is required, so it matches the language/locale policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.