T09 · Insecure Skill Coding Practices
- Location
scripts/save_paper.py:98- Finding
Weak arXiv URL Validation Enables Server-Side Request Forgery and Unintended Data Upload
- Content
View full analysis
Vulnerability Details
File Location:
scripts/save_paper.py, lines 98-127
Vulnerability Type: Server-Side Request Forgery caused by insufficient URL and redirect validation
Risk Level: HighVulnerable Code
python # 下载并附加 PDF if 'arxiv.org' in args.url: try: import urllib.request import tempfile # 将摘要链接转换为 PDF 链接 pdf_url = args.url.replace('/abs/', '/pdf/') if not pdf_url.endswith('.pdf'): pdf_url += '.pdf' print(f"正在下载 PDF...") # 设置 User-Agent 以支持下载 opener = urllib.request.build_opener() opener.addheaders = [('User-agent', 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) Chrome/120.0.0.0')] urllib.request.install_opener(opener) # 创建安全的文件名 safe_title = "".join(c for c in args.title if c.isalnum() or c in (" ", "-", "_")).strip() safe_title = safe_title[:50] # 限制长度 safe_filename = f"{safe_title}.pdf" # 使用临时目录,但指定文件名 with tempfile.TemporaryDirectory() as td: pdf_path = os.path.join(td, safe_filename) urllib.request.urlretrieve(pdf_url, pdf_path) print(f"正在上传 PDF 附件({safe_filename})...") zot.attachment_simple([pdf_path], item_key) print("PDF 已附加。") except Exception as e: print(f"附加 PDF 失败:{e}", file=sys.stderr)Technical Analysis
The script decides whether a URL is an arXiv resource by checking whether the untrusted string supplied through
--urlcontains the substringarxiv.org. A substring match does not establish that th ...[truncated 2289 chars]- Remediation
View remediation
Remediation Suggestions
- Parse the input with
urllib.parse.urlsplitinstead of using substring matching. - Require
httpsand an exact, case-normalized hostname allowlist, such asarxiv.organd explicitly approved arXiv subdomains. - Reject URLs containing credentials, nonstandard ports, fragments, malformed hostnames, or unexpected path formats.
- Extract and strictly validate an arXiv identifier, then construct the PDF URL from a fixed trusted base URL rather than modifying the supplied URL.
- Resolve the destination hostname and reject loopback, private, link-local, multicast, reserved, and unspecified IP addresses.
- Disable automatic redirects or validate the scheme, hostname, port, and resolved IP address after every redirect.
- Apply connection and read timeouts, a maximum response size, and a download quota.
- Verify the response status,
Content-Type, and PDF file signature before uploading the file. - Where possible, enforce an outbound network policy that only permits access to approved Zotero and arXiv endpoints.
- Parse the input with
