T09 · Insecure Skill Coding Practices
- Location
cli.py:77- Finding
Unrestricted URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
Vulnerability Details
File Location:
cli.py, lines 77-86; invocation at lines 269-271
Vulnerability Type: Server-Side Request Forgery
Risk Level: HighVulnerable Code
python def extract_web_text(url): """提取网页正文""" import requests from bs4 import BeautifulSoup headers = {"User-Agent": "Mozilla/5.0"} r = requests.get(url, timeout=15, headers=headers) r.encoding = r.apparent_encoding soup = BeautifulSoup(r.text, "html.parser") for tag in soup(["script", "style", "nav", "footer", "header", "aside"]): tag.decompose() text = soup.get_text(separator="\n", strip=True) lines = [l.strip() for l in text.split("\n") if len(l.strip()) > 15] return "\n".join(lines)The user-controlled URL reaches this function through:
python if args.url: print(f"\U0001f4d6 读取网页:{args.url}") text = extract_web_text(args.url)Technical Analysis
The application passes a user-controlled URL directly to
requests.get()without validating the URL scheme, destination hostname, resolved IP address, port, or redirect chain. Python Requests also follows redirects by default.Consequently, a URL may target loopback addresses, private network ranges, link-local services, cloud instance metadata endpoints, or internal hostnames accessible from the Agent's execution environment. An initially public URL could also redirect to one of these destinations.
This behavior exceeds the minimum network privileges required for reading public webpages because no boundary distinguishes public web content from internal network resources. Extracted response text is subsequently supplied to Edge TTS, creating a potential path by which readable internal content is transmitted to an external TTS provider.
Attack Path
- An attacker asks the Agent to read a URL controlled by the attacker or directly supplies an internal-service URL.
- The Ski ...[truncated 1277 chars]
- Remediation
View remediation
Remediation Suggestions
- Accept only explicitly supported
httpandhttpsURLs. - Reject URLs containing credentials, ambiguous host syntax, unsupported ports, or malformed hostnames.
- Resolve the destination hostname before connecting and reject every address in loopback, private, link-local, multicast, unspecified, and reserved ranges for both IPv4 and IPv6.
- Disable automatic redirects or validate the scheme, hostname, and resolved address again at every redirect hop.
- Apply the same destination policy immediately before connection to reduce DNS-rebinding risk.
- Consider an allowlist of approved public domains when the operating environment permits it.
- Route webpage retrieval through a restricted proxy with egress controls rather than granting the Skill unrestricted network access.
- Impose response-size and content-type limits before parsing or forwarding content.
- Require explicit user confirmation before fetching a URL rather than automatically treating every standalone URL as a TTS request.
- Clearly disclose that extracted text is sent to Edge TTS, and do not submit content obtained from internal or sensitive sources to the external provider.
- Accept only explicitly supported
