T09 · Insecure Skill Coding Practices
- Location
scripts/create_awesome_repo.py:43- Finding
Server-Side Request Forgery in Generated URL Verifier
- Content
View full analysis
list[str]: text = path.read_text(encoding="utf-8") urls = set() for match in re.finditer(r"\[([^\]]*)\]\(([^)\s]+)\)", text): url = match.group(2).strip() if url.startswith(("http://", "https://")): urls.add(url) for match in re.finditer(r"https?://[^\s>)]+", text): urls.add(match.group(0).rstrip(".,")) return sorted(urls) def check_url(url: str, timeout: int) -> dict[str, object]: request = urllib.request.Request( url, headers={"User-Agent": "Mozilla/5.0 (compatible; awesome-url-checker/1.0)"}, ) started = time.time() try: with urllib.request.urlopen(request, timeout=timeout) as response: status = int(response.status) final_url = response.geturl() ``` The extracted URLs are subsequently requested without destination validation: ```python urls = extract_urls(readme) if args.limit: urls = urls[: args.limit] print(f"Found {len(urls)} URLs") results = [] for index, url in enumerate(urls, 1): result = check_url(url, args.timeout) results.append(result) ``` ### Technical Analysis The generator embeds this code into the generated repository as `verify_urls.py`. The verifier treats every string beginning with `http://` or `https://` as safe and passes it directly to `urllib.request.urlopen()`. No controls prevent requests to: - Loopback addresses such as `127.0.0.1` or `[::1]` - Private network ranges - Link-local addresses - Cloud instance metadata endpoints - Reserved or multicast addresses - Internal DNS names - Nonstandard HTTP ports - Public URLs that redirect to an internal destination `urllib.request.urlopen()` follows HTTP redirects by defau ...[truncated 2230 chars]- Remediation
View remediation
