T05 · Unauthorized Access and Privilege Escalation
- Location
scrapling.py:98- Finding
Unrestricted URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
scrapling.py:98-125
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: HighVulnerable Code
python clean_url = url.strip() self.results = [] self.output_file = output_file if output_file and not self.validate_output_path(output_file): return {"error": "invalid_output_path", "message": "路径不合法", "results": []} try: if mode == "get": # HTTP 请求模式 page = Fetcher.get(clean_url, timeout=DEFAULT_TIMEOUT) elif mode == "stealthy": # 隐身模式 page = StealthyFetcher.fetch( clean_url, headless=headless, timeout=DEFAULT_TIMEOUT, solve_cloudflare=solve_cloudflare ) elif mode == "dynamic": # 浏览器自动化 page = DynamicFetcher.fetch( clean_url, headless=headless, timeout=DEFAULT_TIMEOUT ) elif mode == "spider": # 爬虫模式 spider = Spider( name="demo", start_urls=[clean_url], concurrent_requests=DEFAULT_CONCURRENT, delay=DEFAULT_DELAY )Technical Analysis
The user-controlled URL is stripped of whitespace and passed directly to HTTP, stealth, browser, or spider fetchers. The code does not validate the URL scheme, hostname, resolved IP address, destination port, or redirect targets.
Consequently, the documented restriction to public websites is not enforced in code. An attacker who can control the URL can direct the fetcher toward loopback interfaces, private networks, link-local services, or cloud instance metadata endpoints. Non-HTTP schemes may also be reachable if an underlying fetcher supports them.
A timeout limits request duration but does not prevent access to unauthorized network locations. The same destination policy must be applied after DNS resolution and to every redire ...[truncated 1184 chars]
- Remediation
View remediation
Remediation Suggestions
- Permit only explicitly supported schemes, normally
httpandhttps. - Parse URLs with a strict URL parser and reject malformed URLs, embedded credentials, and ambiguous host representations.
- Resolve the hostname before connecting and reject loopback, private, link-local, multicast, unspecified, and reserved IPv4 and IPv6 ranges.
- Explicitly block cloud metadata destinations and other environment-specific sensitive endpoints.
- Revalidate every redirect destination and every spider-discovered URL.
- Protect against DNS rebinding by connecting only to the validated resolved address while preserving the intended HTTP host and TLS verification.
- Consider an explicit domain allowlist or a controlled outbound proxy when the permitted targets are known.
- Apply network-level egress controls as defense in depth.
- Permit only explicitly supported schemes, normally
