T09 · Insecure Skill Coding Practices
- Location
scripts/fetch_thread.py:790- Finding
Arbitrary URL Fetching Enables SSRF and External Disclosure of Retrieved Internal Content
- Content
View full analysis
dict: """Fetch a generic web page and extract references with anchor context.""" result = { "url": url, "type": "web_page", "title": "", "body": "", "state": None, "labels": [], "comments": [], "refs": [], "links": [], "metadata": {}, } try: req = Request(url, method="GET", headers={ "User-Agent": "Mozilla/5.0 (compatible; fetch-thread/1.0)" }) with urlopen(req, timeout=20) as resp: html = resp.read().decode("utf-8", errors="replace") ``` The recursive tracker passes seed URLs and discovered URLs directly to this fetcher: ```python # Fetch content sys.stderr.write(f"[chain_tracker] depth={depth} fetching: {url}\n") try: data = fetch_thread.fetch_thread_url(url) except Exception as e: sys.stderr.write(f"[chain_tracker] fetch failed: {e}\n") continue # Record node node = { "url": url, "depth": depth, "type": data.get("type", "unknown"), "title": data.get("title", ""), "body": (data.get("body", "") or "")[:2000], "comments": data.get("comments", [])[:10], "score": score, "reason": reason, } nodes.append(node) # Update knowledge state knowledge_state = _update_knowledge(knowledge_state, node, creds) ``` Fetched content is then included in a prompt sent to the configured external LLM endpoint: ```python def _update_knowledge(knowledge_state: str, node: dict, creds: dict) -> str: """Ask LLM to update knowledge_state after reading a new node.""" title = node.get("title", "") body = (node.get("bo ...[truncated 4152 chars]- Remediation
View remediation
