T09 · Insecure Skill Coding Practices
- Location
scripts/extract_wechat_article.py:151- Finding
Server-Side Request Forgery Through Unrestricted Image Downloads
- Content
View full analysis
Vulnerability Details
File Location:
scripts/extract_wechat_article.py, lines 151–154 and 316–332
Vulnerability Type: Server-Side Request Forgery (SSRF) and unbounded response handling
Risk Level: HighVulnerable Code
python def normalize_image_url(value: Any) -> str: url = html.unescape(str(value or "").strip()) if url.startswith("//"): url = f"https:{url}" return url if url.startswith("http") else ""python def download_images(record: dict[str, Any], output_dir: Path, timeout: int = 30) -> list[dict[str, str]]: output_dir.mkdir(parents=True, exist_ok=True) downloaded: list[dict[str, str]] = [] for idx, entry in enumerate(record.get("imageEntries") or [], start=1): url = str(entry.get("sourceUrl") or "").strip() if not url: continue request = Request(url, headers={**DEFAULT_HEADERS, "Referer": "https://mp.weixin.qq.com/"}) with urlopen(request, timeout=timeout) as response: body = response.read() suffix = guess_image_suffix(response.headers.get("Content-Type", ""), url) file_path = output_dir / f"{idx:02d}{suffix}" file_path.write_bytes(body) downloaded.append({"marker": entry.get("marker", ""), "path": str(file_path), "sourceUrl": url}) return downloadedTechnical Analysis
Image locations are taken from article HTML and considered valid whenever their string begins with
http. The code does not parse and strictly validate the scheme, restrict destination hostnames, resolve and classify destination IP addresses, constrain destination ports, or revalidate redirects.The
--html-fileworkflow allows a locally supplied, potentially attacker-controlled HTML document to populateimageEntrieswith arbitrary HTTP destinations. When--download-imagesis enabled,urlopen()requests each destination. Because Python's default URL opener follows redirects, even a superficially trusted destinati ...[truncated 1971 chars]- Remediation
View remediation
Remediation Suggestions
- Parse URLs with
urlparse()and permit only exacthttpandhttpsschemes. - Prefer an explicit allowlist of required WeChat image CDN hostnames rather than accepting arbitrary hosts.
- Resolve each hostname and reject all loopback, private, link-local, multicast, reserved, and unspecified IPv4 and IPv6 addresses.
- Protect against DNS rebinding by connecting only to a validated resolved address while preserving the expected HTTP host and TLS hostname semantics.
- Disable automatic redirects or implement a redirect handler that repeats scheme, hostname, port, and resolved-IP validation for every hop.
- Reject embedded credentials, unexpected ports, malformed hostnames, and excessive redirect chains.
- Enforce maximum image count, per-image byte size, and aggregate download size.
- Stream responses in bounded chunks instead of calling unrestricted
response.read(). - Require an approved image MIME type and consider verifying file signatures before writing.
- Add regression tests for localhost, RFC1918 addresses, IPv6 loopback, link-local metadata endpoints, encoded or unusual host representations, redirects to internal destinations, oversized responses, and non-image content.
- Parse URLs with
