T09 · Insecure Skill Coding Practices
- Location
universal_crawler_v2.py:111- Finding
Server-Side Request Forgery Through Weak Domain Validation and Unrestricted Redirects
- Content
View full analysis
Vulnerability Details
File Location:
universal_crawler_v2.py:111-115,universal_crawler_v2.py:291-304,universal_crawler_v2.py:399-428
Vulnerability Type: Server-Side Request Forgery (SSRF) and insufficient URL authorization
Risk Level: HighVulnerable Code
python def is_allowed(url: str, allowed: List[str], excluded: List[str]) -> bool: domain = get_domain(url) if allowed and not any(d in domain for d in allowed): return False if any(re.search(p, url, re.I) for p in excluded): return False return Truepython def extract_by_requests(self, url: str) -> Optional[PageContent]: """使用requests提取""" import requests try: headers = { "User-Agent": random.choice(USER_AGENTS), "Accept": "text/html,application/xhtml+xml", } response = requests.get(url, headers=headers, timeout=self.config.timeout) response.raise_for_status() response.encoding = response.apparent_encoding or "utf-8" return self._parse_html(response.text, url) except Exception as e: self.failed.append({"url": url, "error": str(e)}) return Nonepython def _download_image(self, img_url: str, save_dir: Path) -> Optional[str]: """下载图片并返回文件名""" try: import requests import mimetypes img_hash = hashlib.md5(img_url.encode('utf-8')).hexdigest() headers = { "User-Agent": random.choice(USER_AGENTS), "Referer": self.config.start_url } response = requests.get(img_url, headers=headers, timeout=10) if response.status_code == 200: content_type = response.headers.get('content-type', '').split(';')[0].strip() ext = mimetypes.guess_extension(content_type) if not ext: path = urlparse(img_url) ...[truncated 2659 chars]- Remediation
View remediation
Remediation Suggestions
- Replace substring matching with normalized hostname-boundary validation:
- Permit an exact hostname match.
- If subdomains are intended, require the hostname to end with
"." + allowed_domain. - Normalize case and trailing dots before comparison.
- Resolve each destination and reject loopback, private, link-local, multicast, unspecified, and reserved IP ranges for both IPv4 and IPv6.
- Disable automatic redirects and validate every
Locationdestination before following it. - Apply the same URL policy to page URLs, redirects, image URLs, and every browser-based crawler backend.
- Restrict schemes to
httpandhttps; reject URLs containing unexpected credentials or malformed host components. - Consider an explicit outbound proxy or network-level egress policy that prevents access to internal networks and metadata endpoints.
- Add tests for deceptive hostnames, DNS rebinding scenarios, redirects to private addresses, IPv6 literals, and encoded IP representations.
- Replace substring matching with normalized hostname-boundary validation:
