T09 · Insecure Skill Coding Practices
- Location
src/utils.py:85- Finding
Unrestricted URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
src/utils.py:85-90; request sink atsrc/service.py:131-144
Vulnerability Type: Server-Side Request Forgery caused by insufficient destination validation
Risk Level: HighVulnerable Code
src/utils.py:85-90:python def validate_url(value: str) -> str: """Validate that a full HTTP or HTTPS URL was provided.""" parsed = urlparse(value) if parsed.scheme not in {"http", "https"} or not parsed.netloc: raise ValueError(f"Invalid URL '{value}'. Use a full http:// or https:// URL.") return valuesrc/service.py:131-144:python async def crawl_page(url: str, use_js: bool, query: str | None) -> dict: """Crawl a single page and return structured output.""" if CRAWL4AI_IMPORT_ERROR is not None: return { "error": build_setup_error("Crawl4AI is not available in this runtime.") } browser_config = BrowserConfig(headless=True, verbose=False) run_config = build_run_config(use_js, query) try: with redirect_stdout(io.StringIO()), redirect_stderr(io.StringIO()): async with AsyncWebCrawler(config=browser_config, verbose=False) as crawler: result = await crawler.arun(url=url, config=run_config)Technical Analysis
validate_url()only verifies that the scheme is HTTP or HTTPS and that a network location is present. It does not reject loopback, private, link-local, multicast, reserved, or cloud metadata addresses. It also does not ensure that DNS resolution returns only public addresses.The validated value is passed directly to
AsyncWebCrawler.arun(), which makes the request from the network environment of the Skill process. Consequently, a caller can use the scraper as a server-side request primitive. Examples of accepted destinations include localhost services, RFC 1918 private addresses, and link-local metadata addresses such ...[truncated 2239 chars]- Remediation
View remediation
Remediation Suggestions
- Resolve the hostname before crawling and reject every address that is not globally routable. Block loopback, private, link-local, reserved, multicast, unspecified, and IPv4-mapped IPv6 variants.
- Explicitly deny localhost names and cloud metadata endpoints, including
169.254.169.254and relevant IPv6 link-local addresses. - Reject URLs containing ambiguous host syntax, embedded credentials, malformed ports, or unsupported hostname representations.
- Disable automatic redirects or validate every redirect destination using the same policy before following it.
- Mitigate DNS rebinding by binding the validated hostname to checked IP addresses and ensuring the actual connection does not use a newly resolved prohibited address.
- Apply outbound firewall or sandbox rules that prevent the browser process from connecting to loopback, private networks, metadata services, and other sensitive ranges.
- If operationally feasible, use an explicit domain allowlist rather than accepting arbitrary public destinations.
- Restrict browser subresource requests because JavaScript pages can request destinations other than the top-level URL.
- Add automated tests covering IPv4, IPv6, alternative address representations, redirects, and DNS rebinding scenarios.
