T09 · Insecure Skill Coding Practices
- Location
scripts/generate_report.py:74- Finding
Stored HTML and Script Injection Through Untrusted Patent Links
- Content
View full analysis
str: """Render a patent number as a clickable link.""" if not pn: return pn or "" if url: href = url elif patent_id: href = f"https://eureka.zhihuiya.com/view/#/fullText?patentId={patent_id}" else: import urllib.parse href = f"https://eureka.zhihuiya.com/patent-search/search?q={urllib.parse.quote(pn)}" return ( f'{esc(pn)}' ) ``` ### Technical Analysis The `url` and `patent_id` fields originate from report input data and are incorporated into the quoted `href` attribute without HTML attribute escaping. Although the displayed patent number is escaped with `esc(pn)`, the actual link target is not. An attacker who can influence a retrieved patent record or the normalized input JSON can supply an attribute-breaking value such as: ```text https://example.invalid/" onmouseover="alert(document.domain) ``` This produces attacker-controlled HTML attributes in the generated report. The renderer also does not restrict URL schemes. Consequently, a value using a browser-executable scheme such as `javascript:` could execute when a reader follows the link. HTML escaping alone is insufficient for that second case; explicit scheme validation is required. Other report sections escape URL attribute values, but they also lack a protocol allowlist. The patent renderer is more directly exploitable because it does not perform even attribute escaping. The base64 operation elsewhere in this script is not part of this vulnerability. It only embeds an RDKit-generated molecular S ...[truncated 1582 chars]- Remediation
View remediation
' f'{esc(pn)}' ) ``` 2. **Allowlist safe URL schemes** Use `urllib.parse.urlsplit()` and accept only `https` and, if operationally necessary, `http`. Reject `javascript:`, `data:`, `file:`, and other schemes. Invalid links should be rendered as escaped plain text rather than as anchors. ```python from urllib.parse import urlsplit def safe_web_url(value: str) -> str: value = (value or "").strip() try: parsed = urlsplit(value) except ValueError: return "" if parsed.scheme.lower() not in {"https", "http"}: return "" if not parsed.netloc: return "" return value ``` 3. **Validate patent identifiers** Validate `patent_id` against the identifier format expected by the destination service before including it in a generated URL. Encode it with `urllib.parse.quote()` even after format validation. 4. **Apply the same URL policy globally** Centralize URL rendering and apply the safe-scheme allowlist to patent, paper, clinical-trial, deal, news, and source links. Existing HTML escaping in those sections does not prevent executable URL schemes. 5. **Add security regression tests** Test at least the following inputs: - Quotes that attempt to break out of `href` - Injected event-handler attributes - `javascript:` URLs - `data:` URLs - Malformed URLs - Unexpected patent identifiers - Valid HTTP and HTTPS links 6. **Add defense-in-depth controls** When reports are served over HTTP, apply a restrictive Content Security Policy that blocks inline script and limits navigation and resource loading to approve ...[truncated 138 chars]
