T09 · Insecure Skill Coding Practices
- Location
scripts/content-security-filter.py:160- Finding
Unrestricted URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
scripts/content-security-filter.py, lines 160-168; user-controlled invocation at lines 183-187
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: HighVulnerable Code
python def fetch_url(url: str) -> str: headers = { "User-Agent": "Mozilla/5.0 (compatible; SecurityFilter/1.0)", "Accept": "text/html,text/plain" } req = urllib.request.Request(url, headers=headers) try: with urllib.request.urlopen(req, timeout=15) as resp: raw = resp.read(500_000) # 500KB max return raw.decode('utf-8', errors='ignore') except Exception as e: return f"[FETCH_ERROR: {e}]"The attacker-controlled value reaches this function through:
python elif args.url: if not args.quiet: print(f"Fetching: {args.url}", file=sys.stderr) content = fetch_url(args.url)Technical Analysis
The
--urlargument is passed directly tourllib.request.Requestandurllib.request.urlopenwithout validating:- The permitted URL scheme
- The resolved destination IP address
- Whether the destination is loopback, private, link-local, or otherwise reserved
- Redirect destinations
- Cloud instance metadata addresses
- Whether DNS resolution changes between validation and connection
Although fetching web content is part of the declared functionality, the documented purpose is to scan external content. Access to arbitrary destinations, including services reachable only from the machine running the Skill, exceeds the minimum network privileges needed for that purpose.
The fetched response is passed to
scan_content()and included in its JSON output under thesanitizedproperty. Safe content is returned without truncation, while content classified as unsafe is truncated to 5,000 characters. The fetch itself reads up to 500,000 bytes.The implementat ...[truncated 2017 chars]
- Remediation
View remediation
Remediation Suggestions
-
Restrict URL schemes
- Parse URLs with
urllib.parse.urlsplit. - Permit only
httpandhttps. - Reject URLs containing embedded credentials, malformed hosts, or unsupported ports.
- Parse URLs with
-
Block non-public destinations
- Resolve the hostname before connecting.
- Use Python's
ipaddressmodule to reject every resolved address classified as loopback, private, link-local, reserved, multicast, unspecified, or otherwise non-global. - Explicitly deny cloud metadata destinations, including well-known link-local metadata addresses.
-
Validate redirects
- Disable automatic redirects or implement a custom redirect handler.
- Reapply scheme, hostname, port, and resolved-address validation to every redirect target.
- Enforce a small redirect limit.
-
Mitigate DNS rebinding
- Ensure the address validated is the address used for the connection.
- Reject a hostname if any of its resolved addresses are prohibited.
- Consider routing requests through a hardened outbound proxy that enforces destination policy.
-
Apply least-privilege network controls
- Run the Skill in a sandbox with outbound access limited to required public HTTP/HTTPS destinations.
- Block access to localhost, private subnets, internal DNS zones, and metadata services at the firewall or container-network layer.
-
Reduce disclosure
- Do not return fetched response bodies unless explicitly required.
- If content must be returned, apply a strict output limit and redact sensitive patterns.
- Consider returning only findings by default and requiring an explicit trusted option to include sanitized content.
-
Add security tests
- Test direct requests and redirects to loopback, RFC 1918, IPv6 local, link-local, and metadata addresses.
- Include alternate IP representations, hostnames resolving to mixed public/private addresses, and DNS-rebinding sc ...[truncated 8 chars]
-
