T05 · Unauthorized Access and Privilege Escalation
- Location
scripts/scrapling_fetch.py:162- Finding
Arbitrary URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
scripts/scrapling_fetch.py:42,scripts/scrapling_fetch.py:80-86, andscripts/scrapling_fetch.py:162-171
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: HighThe command-line URL is passed directly to Scrapling's HTTP and browser-based fetchers without validating the URL scheme, destination hostname, resolved IP address, port, or redirect targets.
Vulnerable Code
User-controlled URL declaration and dispatch:
python parser.add_argument("url", help="目标网页 URL") parser.add_argument("max_chars", nargs="?", type=int, default=30000, help="最大输出字符数(默认: 30000)") parser.add_argument("--mode", choices=["basic", "stealth"], default="basic", help="抓取模式: basic(快速)/ stealth(隐身,绕过反爬)") parser.add_argument("--json", action="store_true", help="JSON 格式输出") parser.add_argument("--debug", action="store_true", help="显示调试信息") args = parser.parse_args() try: # 抓取页面 if args.mode == "stealth": html, selector = fetch_stealth(args.url, args.debug) else: html, selector = fetch_basic(args.url, args.debug)Direct request in basic mode:
python page = Fetcher.get(url)Direct request in stealth mode:
python try: page = StealthyFetcher.fetch( url, headless=True, network_idle=True, wait=3, )Technical Analysis
The
urlpositional argument is entirely caller-controlled. Both fetching modes use that value as a network destination without applying an allowlist or rejecting loopback, private, link-local, reserved, or otherwise sensitive addresses.Consequently, when this skill operates in a network environment more privileged than the caller, it can act as a request proxy into that environment. Potential targets include:
- Services bound to the runtime's loopback interface.
- Private RFC1918 ne ...[truncated 2313 chars]
- Remediation
View remediation
Remediation Suggestions
Introduce a centralized URL validation and safe-fetch layer used by both fetching modes:
- Parse URLs with a standards-compliant parser and permit only explicitly required schemes, normally
httpandhttps. - Reject malformed URLs, embedded credentials, missing hostnames, ambiguous numeric IP representations, and unsupported ports.
- Resolve the hostname before connecting and reject every address classified as loopback, private, link-local, multicast, reserved, unspecified, or otherwise non-global.
- Validate every returned DNS address rather than checking only the first address.
- Disable automatic redirects where possible. If redirects are required, apply the complete validation process to every redirect target and enforce a low redirect limit.
- Protect against DNS rebinding by ensuring the validated address is the address actually used for the connection. Prefer network-layer egress controls over relying exclusively on application checks.
- Use an explicit hostname allowlist when the expected target set is known.
- Block access to cloud metadata, localhost, and private networks through firewall or container egress policy as defense in depth.
- Apply request timeouts, response-size limits, and concurrency limits to reduce denial-of-service exposure.
- Add tests covering loopback addresses, IPv6 loopback, private ranges, link-local ranges, alternate IP encodings, attacker-controlled redirects, and DNS rebinding scenarios.
- Return a generic validation error without fetching or disclosing prohibited destination details.
Example validation should occur before either of these calls:
python safe_url = validate_public_http_url(args.url) if args.mode == "stealth": html, selector = fetch_stealth(safe_url, args.debug) else: html, selector = fetch_basic(safe_url, args.debug)The same checks must also be enforced inside the network layer for redirect destinations rather ...[truncated 48 chars]
- Parse URLs with a standards-compliant parser and permit only explicitly required schemes, normally
