T09 · Insecure Skill Coding Practices
- Location
scripts/extract_router.py:52- Finding
Unvalidated URLs are delegated to network-capable tools and third-party extraction proxies
- Content
View full analysis
RoutePlan: url = url.strip() parsed = urlparse(url) path = parsed.path or "" query = parsed.query or "" target = f"{path}?{query}" ``` ```python return RoutePlan( input_url=url, source_type="网页", handler="proxy_cascade", fallback_chain=["r.jina.ai", "defuddle.md", "web_fetch"], notes="通用网页优先代理级联,失败后再走本地回退。", save_name=normalize_save_name(title, "网页"), output_format="markdown", extraction_steps=[ "先用 r.jina.ai 去噪抽取", "失败则用 defuddle.md 结构化净化", "再失败则 web_fetch", "最后 browser fallback", ], failure_modes=[ "JS 重渲染导致空白", "页面噪音过重", "抽取层只返回杂乱 HTML", ], ) ``` The same behavior is declared in `README.md`, lines 67-73: ```markdown ### 4) 通用网页 按顺序尝试: 1. `r.jina.ai` 2. `defuddle.md` 3. `web_fetch` 4. browser fallback ``` ### Technical Analysis The router accepts an arbitrary string, parses it, and places the original value into a plan for network-capable proxy, fetch, and browser tools. It does not enforce an `http` or `https` scheme, reject embedded credentials, validate the destination host, resolve and reject private addresses, or impose redirect restrictions. Although the Python scripts only generate execution specifications and do not themselves initiate network requests, the documented OpenClaw workflow directs the subsequent tool layer to execute the resulting plan. Consequently, the effective security boundary includes the delegated `r.jina.ai`, `defuddle.md`, `web_fetch`, and browser operations. This creates two related risks: 1. **Internal-resource access:** An input such as a loopback, private-network, link- ...[truncated 2948 chars]- Remediation
View remediation
