T09 · Insecure Skill Coding Practices
- Location
scripts/extract_figure_urls.py:58- Finding
Unrestricted Figure Retrieval Enables Server-Side Request Forgery
- Content
View full analysis
]+src=["\']([^"\']+)["\']', block, re.IGNORECASE) img_urls = [] for src in srcs: if src.startswith('data:'): continue full_url = urljoin(base_url, src) img_urls.append(full_url) ``` The documented workflow subsequently retrieves the extracted URL: ```bash # reader/html.md:81-85 mkdir -p {workspace}/figures curl -sL "{FIGURE_URL}" -o {workspace}/figures/{FIGURE_NAME} ``` ### Technical Analysis Figure URLs originate in arbitrary HTML supplied through a user-selected webpage. `urljoin()` resolves absolute, scheme-relative, and relative URLs, but the result is not validated before it is presented for download. There are no restrictions on: - URL scheme - Destination hostname or port - Loopback, private, link-local, or reserved IP addresses - Cloud instance metadata addresses - Cross-origin image retrieval - Redirect destinations The documented `curl -L` operation follows redirects, so validating only the initial URL would also be insufficient. Although an agent is expected to select relevant figures manually, the workflow explicitly directs it to download URLs emitted by this parser and does not require a security review of those destinations. ### Attack Path 1. An attacker hosts an apparent academic paper or modifies a supported HTML paper page. 2. The page includes a `` containing an image such as: ```htmlExperimental results ``` 3. `extract_figure_urls.py` accepts and prints the URL without checking its destination. 4. Following `reader/html.md`, the agent invokes `curl -sL` on the extracted UR ...[truncated 992 chars]
- Remediation
View remediation
