T09 · Insecure Skill Coding Practices
- Location
scripts/sync_debox_docs.py:426- Finding
Server-Side Request Forgery Through Untrusted Documentation Resources
- Content
View full analysis
str | None: if urlsplit(url).scheme not in {"http", "https"}: return None try: url = urlunsplit((*urlsplit(url)[:4], "")) request = Request(request_url(url), method="HEAD", headers={"User-Agent": USER_AGENT}) with urlopen(request, timeout=TIMEOUT) as response: if response.status >= 400: return f"{url}: HTTP {response.status}" return None except HTTPError as exc: if exc.code in {403, 405}: try: request = Request(request_url(url), headers={"User-Agent": USER_AGENT, "Range": "bytes=0-0"}) with urlopen(request, timeout=TIMEOUT): return None ``` Image URLs extracted from documentation are also fetched without destination validation: ```python for parser in parsed_pages.values(): for source, alt in parser.images: image_usage.setdefault(source, set()).add(parser.base_url) if alt: image_alts.setdefault(source, set()).add(alt) if source in image_sources: continue try: data, content_type = fetch(source, MAX_IMAGE_BYTES, "image/*") if not content_type.startswith("image/"): raise ValueError(f"expected image but received {content_type}") image_sources[source] = (data, content_type) except Exception as exc: # noqa: BLE001 broken_images.append(f"{source}: {exc}") ``` All HTTP and HTTPS links extracted from document articles are checked concurrently: ```python check_urls = sorted({ urlunsplit((*urlsplit(url)[:4], "")) for url in all_links if urlsplit(url).scheme in {"http", ...[truncated 2941 chars]- Remediation
View remediation
