T09 · Insecure Skill Coding Practices
- Location
scripts/search.py:89- Finding
Unvalidated Search Result URLs Permit Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
scripts/search.py:89-96andscripts/search.py:190-193
Vulnerability Type: Server-Side Request Forgery through unvalidated remote URL fetching
Risk Level: HighVulnerable Code:
python def _extract_content(url: str) -> Optional[str]: """Fetch a URL and extract clean readable text using trafilatura.""" try: downloaded = trafilatura.fetch_url(url, config=_traf_config) if not downloaded: return None return trafilatura.extract(downloaded, include_links=False, include_images=False, include_tables=True, deduplicate=True) or None except Exception: return Nonepython futures = {} pool = ThreadPoolExecutor(max_workers=max(1, len(results))) for i, r in enumerate(results): if r.get("url"): futures[i] = pool.submit(_extract_content, r["url"])Technical Analysis
URLs supplied by the Serper search response are passed directly to
trafilatura.fetch_url()without validation. The implementation does not:- Restrict URLs to HTTP and HTTPS.
- Resolve hostnames and reject loopback, private, link-local, reserved, or multicast addresses.
- Block cloud instance metadata endpoints.
- Revalidate destinations after HTTP redirects.
- Enforce an explicit response-body size limit.
Search results constitute remotely influenced input. An attacker may operate an indexed page or compromise a page that appears in search results. That page can redirect the content-fetching client to an internal service. If the underlying fetching library follows redirects, the Skill may issue requests that the user could not issue directly from outside the execution environment.
The API credential is not attached to these page-fetching requests; it is only sent to the hard-coded Serper API endpoint. Consequently, this issue does not directly disclose the Serper cred ...[truncated 1591 chars]
- Remediation
View remediation
Remediation Suggestions
- Accept only
httpandhttpsURLs and reject URLs containing credentials or malformed hostnames. - Resolve the destination hostname before connecting and reject every address in loopback, private, link-local, reserved, multicast, and unspecified ranges for both IPv4 and IPv6.
- Disable automatic redirects or validate the scheme, hostname, DNS resolution, and IP address after every redirect.
- Explicitly block common metadata destinations, including
169.254.169.254, even where link-local filtering is already present. - Protect against DNS rebinding by ensuring that validation and connection use the same resolved public address.
- Set strict connection, read, total-operation, and response-size limits.
- Run page extraction in a sandbox with outbound network policy restricted to public Internet destinations.
- Log rejected destinations without logging credentials or sensitive response data.
- Add tests covering direct private addresses, encoded IP representations, IPv6 loopback, DNS names resolving to private addresses, and redirect chains to prohibited destinations.
- Accept only
