Back to skill

Security audit

Web Crawl

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed web crawler, but it can automatically fetch arbitrary URLs, including internal network targets, without clear controls or confirmation.

Install only in an environment where unrestricted outbound crawling is acceptable. Before use, add or require controls that block localhost, private IP ranges, cloud metadata addresses, unusual ports, and redirects to prohibited destinations, and require explicit user confirmation before parallel or research-driven crawling of untrusted URLs.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
web_crawl.py:45
Finding

Server-Side Request Forgery Through Unrestricted URL Crawling

Content
View full analysis
Dict[str, Any]: """ Crawl and extract content from a URL Args: url: Target URL mode: Extraction mode (text, markdown, links, structured, full) max_length: Maximum content length selector: Optional CSS selector for targeted extraction """ try: resp = requests.get( url, headers=self.headers, timeout=self.timeout, allow_redirects=True ) ``` ### Technical Analysis The caller-controlled `url` is passed directly to `requests.get` without validation of its scheme, hostname, destination port, or resolved IP address. The crawler also enables automatic redirects through `allow_redirects=True` without validating each redirect destination. Consequently, a caller can make the process issue requests from the crawler's network context to destinations that may not be reachable externally, including: - Loopback services such as `127.0.0.1` or `::1` - Private network ranges - Link-local services - Cloud instance metadata endpoints - Internal hostnames resolved through local DNS - Public endpoints that redirect to an otherwise prohibited internal address A hostname-only check would not be sufficient because DNS rebinding, alternate IP representations, IPv6 addresses, and redirects can bypass superficial validation. Validation must be performed against the resolved destination immediately before every connection. The same vulnerable primitive is reachable through `crawl_url`, `parallel_crawl`, the command-line interface, and website-analysis functionality. ### Attack Path 1. An attacker supplies a URL to a feature that invokes the crawler. 2. The URL identi ...[truncated 1405 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (14)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose says this is a web crawling and extraction skill, but the behavior reportedly extends into research planning, report synthesis, and broken or undefined crawl paths. This mismatch is dangerous because reviewers and policy systems may authorize the skill under a narrower trust model than the code actually exercises, enabling unintended autonomous actions or unsafe execution paths.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · research.py (reported line 51)May include surrounding context.

python
Comprehensive research report
    """
    # Step 1: Search (will be done via OpenClaw's web_search tool)
    # For now, we return instructions on how to use this
    
    instructions = f"""# Deep Research Instructions

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · research.py (reported line 85)May include surrounding context.

python
Comprehensive research report
    """
    # Step 1: Search (will be done via OpenClaw's web_search tool)
    # For now, we return instructions on how to use this
    
    instructions = f"""# Deep Research Instructions

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This markdown example shows a user request and assistant response entirely in Chinese, which implies the skill may respond in a fixed language during research workflows. There is no accompanying note that language should follow user preference or that this example is region-specific, so it can violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README advertises direct and parallel crawling of arbitrary URLs but does not warn that this causes outbound network requests using user-supplied destinations. In an agent environment, that omission can enable SSRF-like behavior, privacy leakage, or unexpected access to internal or sensitive endpoints if users or prompt content influence the URLs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README says the tool will be used automatically whenever the user needs 'deep research,' which is a broad activation cue that can match ordinary requests. In an agent setting, this can cause unnecessary network access and automated crawling of external URLs without clear user intent, increasing the chance of overreach and unintended data transmission.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill advertises and enables network-capable crawling behavior but does not declare any explicit tool scope, permissions, or allowed-tools boundary. That creates unclear execution authority and increases the chance the agent invokes broader network functionality than intended, especially in environments where tool access is policy-driven.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Activation keywords such as 'research' and 'analyze website' are broad and likely to match many ordinary user requests unrelated to this specific skill. Overbroad triggering can cause the agent to invoke network crawling unexpectedly, expanding data access and tool usage beyond user intent or the minimal necessary capability.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The 'When to Use' section includes ambiguous conditions like topic research and systematic analysis, which can be interpreted far more broadly than simple content extraction. In context, that ambiguity makes the skill more dangerous because it is network-capable and may be selected for general research tasks without clear user consent to crawl external sites.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The command help text says crawl <url> [mode] - Crawl single URL, implying the script implements that operation, but the crawl branch calls crawl_url, which is not defined anywhere in this file. This is more than incomplete documentation: the documented command path contradicts the actual executable behavior and would fail at runtime rather than performing the advertised crawl.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The entire skill description, usage guidance, and user-facing invocation example are presented only in Chinese, with no indication that other languages are supported or that Chinese is required for a documented regional purpose. This can violate a language/locale policy when users are not given an explicit opt-in or choice.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This code constructs user-facing instructions telling the user to invoke web_search and parallel_crawl, both of which involve external network activity. The instructions describe how to run those steps but omit any warning that search queries may be transmitted to external services and URLs will be fetched from third-party sites.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This code file performs a network-style crawl via crawler.crawl(url, ...), which fetches content from a user-supplied URL. Although the docstring says the function analyzes a website, there is no explicit user-facing warning, confirmation, or disclosure near the operation that external requests will be made and remote content retrieved.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The crawler hard-codes an Accept-Language header of "en-US,en;q=0.9,zh-CN;q=0.8" for every request. This imposes a specific locale preference on all outbound requests without giving the user a choice or documenting a justified region-specific need.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.