T09 · Insecure Skill Coding Practices
- Location
scripts/extract_web_text.py:121- Finding
Unrestricted URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
scripts/extract_web_text.py, lines 121-126
Vulnerability Type: Server-Side Request Forgery and unbounded response processing
Risk Level: Mediumpython def fetch_html(url: str, timeout: int, insecure: bool) -> str: request = Request(url, headers={"User-Agent": USER_AGENT}) context = ssl._create_unverified_context() if insecure else ssl.create_default_context() with urlopen(request, timeout=timeout, context=context) as response: charset = response.headers.get_content_charset() or "utf-8" return response.read().decode(charset, errors="replace")Technical Analysis
The user-controlled URL is passed directly to
urllib.request.urlopenwithout validating its scheme, destination address, resolved IP address, or redirect targets. Consequently, the extractor can make requests to loopback interfaces, private networks, link-local services, and cloud instance metadata endpoints.Scheme validation is also absent. The implementation should explicitly limit input to HTTP and HTTPS rather than relying on the behavior of the URL library. Redirects require independent validation because an initially public URL can redirect to a prohibited internal destination.
The response is consumed using an unrestricted
response.read(). The--max-charsoption limits extracted text only after the entire HTTP response has been downloaded and decoded, so it does not prevent excessive memory use from a large response.Attack Path
- An attacker supplies a URL that targets an internal resource, such as a loopback service, private-network administrative interface, or link-local metadata endpoint.
- The Agent follows the Skill workflow and invokes
extract_web_text.pywith that URL. urlopensends the request from the Agent's execution environment, where the target may be reachable even though it is inaccessible to the attacker.- The internal resp ...[truncated 1109 chars]
- Remediation
View remediation
Remediation Suggestions
- Parse the URL before creating the request and permit only
httpandhttps. - Reject URLs containing embedded credentials.
- Resolve the hostname and reject every address that is loopback, private, link-local, multicast, reserved, or unspecified.
- Revalidate the destination after every redirect. Prefer a custom redirect handler that refuses redirects to prohibited hosts or addresses.
- Consider an explicit host allowlist when the deployment has a limited set of permitted sources.
- Stream the response in bounded chunks and stop after a configured maximum number of bytes. Do not call unrestricted
response.read(). - Validate the response content type and reject unexpected binary content.
- Apply separate connection and total-transfer time limits where supported.
- Require explicit user confirmation before accessing non-public or unusual destinations.
- Add tests covering loopback addresses, IPv6 loopback, private ranges, link-local addresses, DNS rebinding, redirects to internal services, and oversized responses.
- Parse the URL before creating the request and permit only
