T09 · Insecure Skill Coding Practices
- Location
scripts/douyin_download.py:29- Finding
Insufficient URL Validation Enables Arbitrary Outbound Requests
- Content
View full analysis
str: """Extract the first URL from shared text.""" m = re.search(r'https?://[^\s]+', text) if not m: raise ValueError("No valid URL found") return m.group(0) def resolve_short_url(url: str) -> str: """Follow redirects and obtain the final URL.""" req = urllib.request.Request(url, headers={"User-Agent": MOBILE_UA}) try: resp = urllib.request.urlopen(req, timeout=10) return resp.url except urllib.error.HTTPError as e: return e.url or url ``` ```python # Resolve short URLs first. if 'v.douyin.com' in url or len(url) < 50: print(f"[2/4] Resolving short URL...") url = resolve_short_url(url) print(f" Expanded: {url}") else: print(f"[2/4] Skipping short URL resolution") ``` ### Technical Analysis The script extracts any HTTP or HTTPS URL from user-controlled input without validating its hostname or resolved IP address. It then sends a request when either of these weak conditions is satisfied: 1. The URL contains the substring `v.douyin.com`; or 2. The complete URL is shorter than 50 characters. The length condition allows arbitrary short URLs such as `http://127.0.0.1:8000/` to be requested. The substring condition can also be bypassed with attacker-controlled hostnames or URL components containing `v.douyin.com`. `urllib.request.urlopen()` follows HTTP redirects by default, but neither the initial URL nor redirect destinations are checked. Consequently, an attacker-controlled public endpoint can redirect the request to a loopback, private-network, link-local, or cloud metadata address. The response body from this particular request is not exposed to the caller, and later video-ID extraction will generally fail for a non-Douyin URL ...[truncated 1517 chars]- Remediation
View remediation
