Back to skill

Security audit

content-extraction

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent content-extraction skill, but it needs Review because it can route arbitrary URLs through external services and local saves without clear user consent or network boundaries.

Install only where it is acceptable for user-supplied URLs and extracted content to be processed by OpenClaw web/browser/Feishu tools and possibly third-party cleanup services. Avoid private, signed, internal, or sensitive document URLs unless you add confirmation gates, disable third-party fallbacks, and control local Markdown retention.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract_router.py:52
Finding

Unvalidated URLs are delegated to network-capable tools and third-party extraction proxies

Content
View full analysis
RoutePlan: url = url.strip() parsed = urlparse(url) path = parsed.path or "" query = parsed.query or "" target = f"{path}?{query}" ``` ```python return RoutePlan( input_url=url, source_type="网页", handler="proxy_cascade", fallback_chain=["r.jina.ai", "defuddle.md", "web_fetch"], notes="通用网页优先代理级联,失败后再走本地回退。", save_name=normalize_save_name(title, "网页"), output_format="markdown", extraction_steps=[ "先用 r.jina.ai 去噪抽取", "失败则用 defuddle.md 结构化净化", "再失败则 web_fetch", "最后 browser fallback", ], failure_modes=[ "JS 重渲染导致空白", "页面噪音过重", "抽取层只返回杂乱 HTML", ], ) ``` The same behavior is declared in `README.md`, lines 67-73: ```markdown ### 4) 通用网页 按顺序尝试: 1. `r.jina.ai` 2. `defuddle.md` 3. `web_fetch` 4. browser fallback ``` ### Technical Analysis The router accepts an arbitrary string, parses it, and places the original value into a plan for network-capable proxy, fetch, and browser tools. It does not enforce an `http` or `https` scheme, reject embedded credentials, validate the destination host, resolve and reject private addresses, or impose redirect restrictions. Although the Python scripts only generate execution specifications and do not themselves initiate network requests, the documented OpenClaw workflow directs the subsequent tool layer to execute the resulting plan. Consequently, the effective security boundary includes the delegated `r.jina.ai`, `defuddle.md`, `web_fetch`, and browser operations. This creates two related risks: 1. **Internal-resource access:** An input such as a loopback, private-network, link- ...[truncated 2948 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
scripts/extract_router.py:38
Finding

Source routing uses substring matching against the complete URL instead of validated hostnames

Content
View full analysis
RoutePlan: url = url.strip() parsed = urlparse(url) path = parsed.path or "" query = parsed.query or "" target = f"{path}?{query}" if WECHAT_RE.search(url): return RoutePlan( input_url=url, source_type="公众号", handler="browser", fallback_chain=["r.jina.ai", "defuddle.md", "web_fetch"], notes="公众号文章优先走浏览器抓取,处理反爬和动态内容。", save_name=normalize_save_name(title, "公众号"), output_format="markdown", extraction_steps=[ "打开文章页面", "等待正文区域加载完成", "提取标题 / 作者 / 发布时间 / 正文 / 图片", "生成 Markdown + frontmatter", ], failure_modes=[ "登录墙", "反爬白页", "正文容器缺失", ], ) if FEISHU_RE.search(url): ``` ### Technical Analysis Platform detection searches the entire raw URL rather than `parsed.hostname`. Therefore, a platform name appearing in an attacker-controlled username, path, query string, or unrelated hostname is sufficient to select a specialized handler. Examples that can be incorrectly classified include: ```text https://attacker.example/?next=https://www.feishu.cn/docx/example https://feishu.cn.attacker.example/docx/example https://attacker.example/path/youtube.com/watch ``` This violates least-privilege ro ...[truncated 1738 chars]
Remediation
View remediation
bool: if not host: return False normalized = host.rstrip(".").lower() return normalized == domain or normalized.endswith("." + domain) ``` 3. Apply the helper to approved platform domains: ```python host = parsed.hostname if is_domain(host, "mp.weixin.qq.com"): ... elif is_domain(host, "feishu.cn") or is_domain(host, "larksuite.com"): ... elif is_domain(host, "youtube.com") or is_domain(host, "youtu.be"): ... ``` 4. Normalize internationalized domain names consistently and reject invalid or ambiguous hostnames. 5. Reject URLs containing unexpected user-information fields. 6. Require every specialized downstream tool to independently validate the destination origin rather than trusting the router's classification. 7. Add negative tests for: - `feishu.cn.attacker.example` - `attacker-feishu.cn` - trusted domain names appearing only in paths or queries - mixed-case and trailing-dot hostnames - user-information tricks such as `https://feishu.cn@attacker.example/` ]]>
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (17)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this is a content extraction skill for multiple source types, implying it actually retrieves and extracts content. The supplied code explicitly states and demonstrates that it does not perform extraction itself; it only classifies the input URL and generates a plan with handler, fallback chain, steps, and failure modes. There are no network calls, browser actions, API calls, or transcript/document retrievals. Thus the primary purpose is materially different: routing/planning rather than extraction.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README explicitly recommends saving long extracted content to a local Markdown file, but it does not warn users that extracted material may contain sensitive or copyrighted data and will persist on disk. In a content-extraction skill, silent local persistence increases the chance of unintended data retention, later disclosure, or mishandling of confidential document contents.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill routes user-provided URLs and document content through external tools and services such as browser, Feishu, web_fetch, YouTube transcript tooling, and third-party cleanup endpoints like r.jina.ai/defuddle.md without warning that data may be transmitted off-host. This is dangerous because users may supply private document links or sensitive web resources, causing unintended disclosure to external services or logs during extraction.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill advertises executable extraction behavior involving networked sources, but it declares no explicit tool scope such as permissions or allowed-tools. In an agent environment, this creates ambiguous authority boundaries and can enable unintended network access or overly broad execution if the runtime defaults are permissive.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The markdown content is entirely written in Chinese, including the goal and mapping rules, with no indication that the skill supports other languages or that the Chinese-only constraint is intentional and context-specific. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The document content is written in Chinese and does not offer any language or locale choice, which can constitute a natural-language policy violation under the language/locale rule. There is no indication that the skill is region-specific or that Chinese is required for a justified compliance or product constraint.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The input description "article/blog/news page" is ambiguous because it refers to broad everyday web content categories rather than concrete URL patterns or boundaries. This makes it unclear which pages should activate the skill and which should not.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The example instructs the extractor to try third-party retrieval services like r.jina.ai and defuddle.md without noting trust, privacy, or data-handling implications. In a content-extraction skill, automatically sending user-supplied URLs to external services can leak sensitive URLs or fetched content and may violate user expectations or organizational policy.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The phrase "any supported URL" is a very broad invocation condition and does not define the exact supported patterns or exclusions in this extractor section. Without a precise list or negative examples, the activation scope is unclear and could lead to unintended matching.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file contains core usage instructions entirely in Chinese, including the goal, usage steps, and next-phase plans. Under the policy, forcing a specific language without user opt-in or an explicitly justified regional scope is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module docstring says the executable 'turns a URL into a concrete extraction workflow' and explicitly 'does not talk to external OpenClaw tools directly'; the code then only builds and prints an ExtractionSpec. That behavior is materially narrower than the manifest's description of an extraction skill for URLs, Feishu, YouTube, and web pages, which implies actual content extraction rather than planning output.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code emits source types, notes, extraction steps, and failure modes exclusively in Chinese string literals. Because there is no option to choose language or indication that the skill is intentionally region-specific, it effectively forces a locale/language on users.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The fallback strategy explicitly routes failed WeChat article extraction through third-party services such as r.jina.ai and defuddle.md without any indication that the target URL and potentially article content will be transmitted off-platform. This creates a real privacy and data-handling risk, especially if users supply sensitive, private, or access-controlled links and do not expect external disclosure during fallback behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill states that it may save long extracted content locally but provides no user warning, consent model, or storage constraints. This can lead to unexpected persistence of sensitive or copyrighted data on disk, especially when extracting content from private collaboration platforms or authenticated sessions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The document presents source labels and sections in Chinese for multiple extractors without indicating that language is configurable or user-selected. This may conflict with language or locale policies when users have not opted into Chinese output.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The expected behavior includes choosing a save path based on title, which implies writing files, but the document does not warn users that extracted content may be saved locally. For markdown files, behaviors affecting user data or system state should be disclosed.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The expected-behavior section is written as mandatory instructions in Chinese, including output expectations, without indicating that language choice is optional or user-selected. This can violate a language/locale policy if the organization requires user choice rather than forcing a specific language.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.