T09 · Insecure Skill Coding Practices
- Location
tools/sitemap_extractor.py:1- Finding
Unrestricted Sitemap URL Fetch Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
tools/sitemap_extractor.py, lines 1-9
Vulnerability Type: Server-Side Request Forgery caused by unrestricted outbound requests
Risk Level: HighVulnerable Code
python #!/usr/bin/env python3 import sys, re, requests url = sys.argv[1] if len(sys.argv) > 1 else "" if not url: print("Usage: sitemap_extractor.py https://example.com/sitemap.xml") sys.exit(1) txt = requests.get(url, timeout=20).text for loc in re.findall(r"<loc>(.*?)</loc>", txt): print(loc)Technical Analysis
The script passes a command-line URL directly to
requests.get()without validating its scheme, hostname, resolved IP address, port, or redirect destination. It therefore permits requests to arbitrary network locations reachable from the agent runtime.No controls reject loopback, private, reserved, or link-local destinations. The code also follows HTTP redirects by default, allowing an initially public URL to redirect to an internal service. A timeout limits request duration but does not limit response size, so the entire response is loaded into memory before parsing.
Although the script only prints content found between
locelements, an internal or attacker-controlled response can deliberately wrap sensitive information in those elements. The request itself can also be used for blind internal-service discovery or interaction with HTTP endpoints.Attack Path
- An attacker supplies a sitemap URL pointing directly to an internal address, localhost service, cloud metadata endpoint, or attacker-controlled redirector.
- The user or agent invokes
tools/sitemap_extractor.pywith that URL. - The script issues the request using the agent host's network identity and network access.
- If redirects are used,
requestsfollows them automatically to the final internal destination. - The complete response is read into memory.
- Any response ...[truncated 944 chars]
- Remediation
View remediation
Remediation Suggestions
- Accept only explicitly permitted URL schemes, preferably
https, withhttpenabled only when necessary. - Parse the URL before use and reject embedded credentials, malformed hosts, unexpected ports, and non-HTTP schemes.
- Resolve the hostname and reject every loopback, private, link-local, multicast, reserved, unspecified, and otherwise non-public IP address.
- Prevent DNS-rebinding bypasses by connecting only to validated resolved addresses and checking all returned address records.
- Disable automatic redirects or validate the scheme, host, port, and resolved addresses at every redirect hop.
- Prefer restricting sitemap retrieval to the origin of the website explicitly approved by the user.
- Stream the response and enforce a strict maximum body size.
- Call
raise_for_status()and accept only expected XML content types. - Parse XML with a hardened parser rather than a regular expression, with external entity processing disabled.
- Run the utility in an environment with outbound network restrictions that deny access to internal and metadata ranges.
- Accept only explicitly permitted URL schemes, preferably
