T09 · Insecure Skill Coding Practices
- Location
pdfagent/tools/html_to_pdf.py:90- Finding
Unrestricted Remote HTML Fetching Enables Server-Side Request Forgery
- Content
View full analysis
str: source_path = Path(source) if source_path.exists(): return source_path.read_text(encoding="utf-8") if source.startswith("http://") or source.startswith("https://"): with urlopen(source) as resp: return resp.read().decode("utf-8", errors="ignore") return source ``` The vulnerable functionality is directly exposed through the CLI: ```python @app.command("html-to-pdf") def html_to_pdf_cmd( source: str = typer.Argument(..., help="URL or HTML file path"), out: Path = typer.Option(..., "--out"), json_output: bool = typer.Option(False, "--json"), usage_file: Optional[Path] = typer.Option(None, "--usage-file"), ): meter = UsageMeter("html_to_pdf", [Path(source)] if Path(source).exists() else []) try: html_to_pdf(source, out) ``` ### Technical Analysis The `source` argument is controlled by the caller and may contain an arbitrary HTTP or HTTPS URL. `_load_html()` passes that URL directly to `urllib.request.urlopen()` without validating the destination hostname or resolved IP address. There are no controls to prevent requests to: - Loopback addresses such as `127.0.0.1` or `::1`. - RFC 1918 private networks. - Link-local addresses. - Cloud instance metadata services. - Internal DNS names. - Reserved or otherwise non-public address ranges. Redirect destinations are not independently validated. Consequently, an initially public URL could redirect the request to a prohibited internal destination. The request also has no explicit connection or read timeout and no maximum response-size limit. The primary HTML conversion paths pass the same uncontrolled source to `wkhtmltopdf ...[truncated 1561 chars]- Remediation
View remediation
