T09 · Insecure Skill Coding Practices
- Location
web_crawl.py:45- Finding
Server-Side Request Forgery Through Unrestricted URL Crawling
- Content
View full analysis
Dict[str, Any]: """ Crawl and extract content from a URL Args: url: Target URL mode: Extraction mode (text, markdown, links, structured, full) max_length: Maximum content length selector: Optional CSS selector for targeted extraction """ try: resp = requests.get( url, headers=self.headers, timeout=self.timeout, allow_redirects=True ) ``` ### Technical Analysis The caller-controlled `url` is passed directly to `requests.get` without validation of its scheme, hostname, destination port, or resolved IP address. The crawler also enables automatic redirects through `allow_redirects=True` without validating each redirect destination. Consequently, a caller can make the process issue requests from the crawler's network context to destinations that may not be reachable externally, including: - Loopback services such as `127.0.0.1` or `::1` - Private network ranges - Link-local services - Cloud instance metadata endpoints - Internal hostnames resolved through local DNS - Public endpoints that redirect to an otherwise prohibited internal address A hostname-only check would not be sufficient because DNS rebinding, alternate IP representations, IPv6 addresses, and redirects can bypass superficial validation. Validation must be performed against the resolved destination immediately before every connection. The same vulnerable primitive is reachable through `crawl_url`, `parallel_crawl`, the command-line interface, and website-analysis functionality. ### Attack Path 1. An attacker supplies a URL to a feature that invokes the crawler. 2. The URL identi ...[truncated 1405 chars]- Remediation
View remediation
