T09 · Insecure Skill Coding Practices
- Location
scripts/fetch_page.py:7- Finding
Unrestricted URL Fetching Enables SSRF and Local Resource Access
- Content
View full analysis
Vulnerability Details
File Location:
scripts/fetch_page.py:7-18,scripts/extract_text.py:21-32, andscripts/save_output.py:35-47
Vulnerability Type: Server-Side Request Forgery (SSRF) and unsafe URL scheme handling
Risk Level: HighComplete Vulnerable Code Snippets
scripts/fetch_page.py:7-18python parser.add_argument("--url", required=True, help="Public URL to fetch") args = parser.parse_args() req = urllib.request.Request( args.url, headers={ "User-Agent": "Mozilla/5.0 (compatible; ScraperSkill/1.0)" } ) with urllib.request.urlopen(req, timeout=20) as resp: html = resp.read().decode("utf-8", errors="replace")scripts/extract_text.py:21-32python parser.add_argument("--url", required=True, help="Public URL to fetch and clean") args = parser.parse_args() req = urllib.request.Request( args.url, headers={ "User-Agent": "Mozilla/5.0 (compatible; ScraperSkill/1.0)" } ) with urllib.request.urlopen(req, timeout=20) as resp: html = resp.read().decode("utf-8", errors="replace")scripts/save_output.py:35-47python parser.add_argument("--url", required=True, help="Public URL") parser.add_argument("--title", required=True, help="Local title") args = parser.parse_args() req = urllib.request.Request( args.url, headers={ "User-Agent": "Mozilla/5.0 (compatible; ScraperSkill/1.0)" } ) with urllib.request.urlopen(req, timeout=20) as resp: html = resp.read().decode("utf-8", errors="replace")Technical Analysis
All three scripts pass a user-controlled URL directly to
urllib.request.Requestandurllib.request.urlopenwithout validating its scheme, hostname, resolved IP address, destination port, or redirect targets.Consequently, an attacker may direct requests toward loopback, private, link-local, or otherwise reserved network addresses. Depending on the handlers available to
urllib.request, unsupported non-web scheme ...[truncated 2514 chars]- Remediation
View remediation
Remediation Suggestions
- Parse every URL with
urllib.parse.urlsplitand permit only explicitly supported schemes, preferablyhttpsand, only if required,http. - Reject URLs containing embedded credentials, malformed hostnames, fragments where inappropriate, or destination ports outside an explicit allowlist.
- Resolve the hostname before connecting and inspect every returned IPv4 and IPv6 address with the
ipaddressmodule. - Reject loopback, private, link-local, multicast, reserved, unspecified, and other non-global addresses.
- Explicitly block known metadata destinations, including link-local cloud metadata addresses.
- Disable automatic redirects or implement a custom redirect handler that reparses, resolves, and validates every redirect destination.
- Account for DNS rebinding by ensuring that the validated address is the address actually used for the connection, preferably through network-level egress controls.
- Apply an outbound firewall or proxy policy that prevents the scraper process from reaching local, private, management, and metadata networks.
- Stream response bodies in bounded chunks and enforce a strict maximum response size instead of using unbounded
resp.read(). - Validate response content types before processing and provide controlled error handling without printing sensitive response details.
- Enforce robots, authorization, and public-resource restrictions in code where those guarantees are claimed, rather than relying solely on documentation.
- Apply the same centralized URL-validation function to all three affected scripts to prevent inconsistent protections.
- Parse every URL with
