T09 · Insecure Skill Coding Practices
- Location
scripts/scrape_web.py:15- Finding
Unrestricted URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
scripts/scrape_web.py, lines 15-16, 40, and 60
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: HighVulnerable Code
python def fetch_via_http(url: str) -> str: try: import httpx except Exception as e: raise RuntimeError("Missing dependency. install with: pip install httpx") from e r = httpx.get( url, follow_redirects=True, timeout=30, headers={ # 尽量告诉服务器我们要文本 "Accept": "text/plain,text/markdown,text/html,*/*", "User-Agent": "Mozilla/5.0", }, ) r.raise_for_status() # 自动按响应头编码解码;若为空,httpx 会做合理推断 return r.textpython page = StealthyFetcher.fetch(url, headless=True, network_idle=True)python parser.add_argument("--url", required=True, help="Target URL")Technical Analysis
The user-controlled
--urlvalue is passed directly to bothhttpx.get()andStealthyFetcher.fetch()without validating the URL scheme, destination hostname, resolved IP address, or destination port.The direct HTTP implementation also enables
follow_redirects=True. Consequently, validating only an initial public URL would not be sufficient: an attacker-controlled public endpoint could redirect the request to a loopback, private, link-local, or cloud metadata address.The affected process can therefore be used as a network proxy to access resources available from the host's network context but unavailable to the attacker directly.
Attack Path
- An attacker supplies a URL targeting an internal resource, such as a loopback service, private-network host, or cloud metadata endpoint.
- Alternatively, the attacker supplies a public URL that redirects to an internal destination.
- The URL is forwarded without destination validation to Scrapling or
httpx. - The request executes with the ...[truncated 752 chars]
- Remediation
View remediation
Remediation Suggestions
- Accept only explicitly permitted schemes, normally
httpsand, if required,http. - Reject URLs containing embedded credentials or malformed hostnames.
- Resolve the hostname before connecting and reject loopback, private, link-local, multicast, unspecified, and reserved addresses for both IPv4 and IPv6.
- Protect against DNS rebinding by ensuring the address actually used for the connection remains within the validated address set.
- Disable automatic redirects or validate every redirect destination using the same scheme, hostname, port, and resolved-address controls.
- Restrict destination ports to those required for web scraping.
- Prefer an explicit domain allowlist where the expected targets are known.
- Run the scraper in an outbound network sandbox that cannot access cloud metadata endpoints, internal management networks, or local administrative services.
- Apply the same validation policy to both
StealthyFetcher.fetch()and thehttpxfallback so that error handling cannot bypass the restrictions.
- Accept only explicitly permitted schemes, normally
