T09 · Insecure Skill Coding Practices
- Location
scripts/scrapling_fetch.py:156- Finding
Unrestricted URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
scripts/scrapling_fetch.py, lines 156–163
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: Highpython def fetch_one(url, max_chars, as_json, overrides, Fetcher, HTML2Text): parts = urlparse(url) if parts.scheme not in ('http', 'https'): return {'ok': False, 'url': url, 'error': 'url must start with http:// or https://'} Fetcher.configure(auto_match=True) fetcher = Fetcher() page = fetcher.get(url)Technical Analysis
The URL validation only verifies that the supplied scheme is HTTP or HTTPS. It does not prevent requests to:
- Loopback addresses such as
127.0.0.1or::1 - RFC 1918 private networks
- Link-local addresses
- Cloud metadata services such as
169.254.169.254 - Internal hostnames
- Public hostnames that resolve to restricted addresses
- Redirect destinations that lead to restricted addresses
The requested response is subsequently converted to Markdown and returned through standard output or JSON. Therefore, an attacker who controls or influences the input URL may use the Skill as a network proxy to access resources visible from the execution environment.
Accepting arbitrary public webpage URLs is necessary for the declared content-extraction functionality. However, access to local, private, link-local, and metadata networks exceeds the minimum privileges required for ordinary public webpage extraction.
Attack Path
- An attacker supplies a URL referencing an internal service, loopback endpoint, cloud metadata endpoint, or attacker-controlled hostname that resolves to a restricted address.
urlparse()accepts the input because its scheme ishttporhttps.fetcher.get(url)sends the request from the environment running the Skill.- Scrapling receives content from the otherwise inaccessible endpoint.
- The script extracts and converts t ...[truncated 1174 chars]
- Loopback addresses such as
- Remediation
View remediation
Remediation Suggestions
- Resolve the destination hostname before making a request.
- Reject loopback, private, link-local, multicast, reserved, and unspecified IPv4 and IPv6 addresses.
- Explicitly deny known cloud metadata destinations, including
169.254.169.254. - Re-resolve and revalidate every redirect destination before following it.
- Mitigate DNS rebinding by ensuring the validated address is the address actually used for the connection.
- Prefer an explicit domain allowlist when the expected target sites are known.
- Apply connection and response timeouts.
- Limit response size before loading or converting the complete document.
- Restrict the number of URLs accepted in batch mode.
- Where possible, run the Skill in a network sandbox that cannot reach internal services or metadata endpoints.
- Add automated tests covering IPv4, IPv6, alternate IP representations, redirects, internal DNS names, and DNS rebinding scenarios.
