Back to skill

Security audit

test_skill

Security checks for vulnerabilities and agentic risk

Overview

This is a real web crawler, but its network and installation controls are loose enough that users should review it before installing.

Install and run this only in a constrained environment such as a virtualenv/container with limited filesystem and network access. Use a dedicated output directory, avoid privileged installs, review dependency versions first, and do not crawl targets unless you are authorized and comfortable with the crawler potentially fetching redirected or embedded URLs outside the apparent site boundary.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
universal_crawler_v2.py:111
Finding

Server-Side Request Forgery Through Weak Domain Validation and Unrestricted Redirects

Content
View full analysis

Vulnerability Details

File Location: universal_crawler_v2.py:111-115, universal_crawler_v2.py:291-304, universal_crawler_v2.py:399-428
Vulnerability Type: Server-Side Request Forgery (SSRF) and insufficient URL authorization
Risk Level: High

Vulnerable Code

python
def is_allowed(url: str, allowed: List[str], excluded: List[str]) -> bool:
    domain = get_domain(url)
    if allowed and not any(d in domain for d in allowed):
        return False
    if any(re.search(p, url, re.I) for p in excluded):
        return False
    return True
python
def extract_by_requests(self, url: str) -> Optional[PageContent]:
    """使用requests提取"""
    import requests
    
    try:
        headers = {
            "User-Agent": random.choice(USER_AGENTS),
            "Accept": "text/html,application/xhtml+xml",
        }
        response = requests.get(url, headers=headers, timeout=self.config.timeout)
        response.raise_for_status()
        response.encoding = response.apparent_encoding or "utf-8"
        
        return self._parse_html(response.text, url)
    except Exception as e:
        self.failed.append({"url": url, "error": str(e)})
        return None
python
def _download_image(self, img_url: str, save_dir: Path) -> Optional[str]:
    """下载图片并返回文件名"""
    try:
        import requests
        import mimetypes
        
        img_hash = hashlib.md5(img_url.encode('utf-8')).hexdigest()
        headers = {
            "User-Agent": random.choice(USER_AGENTS),
            "Referer": self.config.start_url
        }
        response = requests.get(img_url, headers=headers, timeout=10)
        if response.status_code == 200:
            content_type = response.headers.get('content-type', '').split(';')[0].strip()
            ext = mimetypes.guess_extension(content_type)
            if not ext:
                path = urlparse(img_url)
...[truncated 2659 chars]
Remediation
View remediation

Remediation Suggestions

  1. Replace substring matching with normalized hostname-boundary validation:
    • Permit an exact hostname match.
    • If subdomains are intended, require the hostname to end with "." + allowed_domain.
    • Normalize case and trailing dots before comparison.
  2. Resolve each destination and reject loopback, private, link-local, multicast, unspecified, and reserved IP ranges for both IPv4 and IPv6.
  3. Disable automatic redirects and validate every Location destination before following it.
  4. Apply the same URL policy to page URLs, redirects, image URLs, and every browser-based crawler backend.
  5. Restrict schemes to http and https; reject URLs containing unexpected credentials or malformed host components.
  6. Consider an explicit outbound proxy or network-level egress policy that prevents access to internal networks and metadata endpoints.
  7. Add tests for deceptive hostnames, DNS rebinding scenarios, redirects to private addresses, IPv6 literals, and encoded IP representations.

T09 · Insecure Skill Coding Practices

Warning
Location
universal_crawler_v2.py:399
Finding

Unbounded and Unconditional Download of Remote Content

Content
View full analysis

Vulnerability Details

File Location: universal_crawler_v2.py:68, universal_crawler_v2.py:399-428, universal_crawler_v2.py:469-478
Vulnerability Type: Uncontrolled resource consumption and insufficient remote-content validation
Risk Level: Medium

Vulnerable Code

python
# 内容
min_content_length: int = 100
download_images: bool = False
python
def _download_image(self, img_url: str, save_dir: Path) -> Optional[str]:
    """下载图片并返回文件名"""
    try:
        import requests
        import mimetypes
        
        img_hash = hashlib.md5(img_url.encode('utf-8')).hexdigest()
        headers = {
            "User-Agent": random.choice(USER_AGENTS),
            "Referer": self.config.start_url
        }
        response = requests.get(img_url, headers=headers, timeout=10)
        if response.status_code == 200:
            content_type = response.headers.get('content-type', '').split(';')[0].strip()
            ext = mimetypes.guess_extension(content_type)
            if not ext:
                path = urlparse(img_url).path
                ext = os.path.splitext(path)[1]
            if not ext or len(ext) > 5:
                ext = ".jpg"
            if ext in ['.jpe', '.jpeg']: ext = '.jpg'
            
            filename = f"{img_hash}{ext}"
            filepath = save_dir / filename
            filepath.write_bytes(response.content)
            return filename
    except:
        pass
    return None
python
images_dir = save_dir / "images"
images_dir.mkdir(exist_ok=True)

original_images = content.images
local_images = []
for img_url in content.images:
    local_name = self._download_image(img_url, images_dir)
    if local_name:
        local_images.append(f"./images/{local_name}")
    else:
        local_images.append(img_url)
content.images = local_images

Technical Analysis

The downloader accesses `respo ...[truncated 1947 chars]

Remediation
View remediation

Remediation Suggestions

  1. Check self.config.download_images before creating the image directory or initiating downloads.
  2. Use stream=True and write fixed-size chunks instead of accessing response.content.
  3. Enforce:
    • A maximum size for each image.
    • A maximum cumulative image size per page.
    • A maximum cumulative download size for the entire crawl.
    • A maximum number of images per page.
  4. Reject responses whose declared Content-Length exceeds the configured limit, while still enforcing limits during streaming because that header can be absent or false.
  5. Permit only an explicit set of image media types and verify file signatures before storage.
  6. Stop reading immediately when a limit is exceeded, delete any partial file, and record a clear failure reason.
  7. Store downloads on a filesystem with an appropriate quota and run the crawler with constrained memory and disk resources.

T08 · Insecure Dependencies

Warning
Location
requirements.txt:4
Finding

Mutable and Unverified Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: requirements.txt:4-10, install.py:34-53
Vulnerability Type: Unpinned dependency and software supply-chain exposure
Risk Level: Medium

Vulnerable Code

text
# Core dependencies
requests>=2.28.0
beautifulsoup4>=4.11.0

# Advanced Crawler (Required)
crawl4ai
playwright>=1.30.0
nest-asyncio>=1.5.0
python
# 1. Install requirements.txt with passed arguments (e.g. --break-system-packages)
print(f"Installing requirements from requirements.txt with args: {unknown_args}...")
try:
    if os.path.exists("requirements.txt"):
        # Add --ignore-installed psutil to avoid "Cannot uninstall psutil" error on Debian/Ubuntu
        # when system psutil is present but pip tries to upgrade it.
        install_args = ["-r", "requirements.txt", "--ignore-installed", "psutil"]
        install(install_args, extra_args=unknown_args)
    else:
        print("Warning: requirements.txt not found!")
except Exception as e:
    print(f"Error installing requirements: {e}")
    sys.exit(1)

# 2. Install Playwright Browsers
print("Installing Playwright browsers...")
try:
    # Check if playwright is installed
    subprocess.check_call([sys.executable, "-m", "playwright", "install", "chromium"])
except Exception as e:
    print(f"Error installing playwright browsers: {e}")

Technical Analysis

The dependency manifest uses lower-bound constraints and leaves crawl4ai completely unpinned. It does not provide package hashes. As a result, installation may resolve to package versions that were not present or reviewed when the Skill was audited.

Python package installation can execute package build hooks and installation logic. A compromised future release, compromised package index, or unsafe alternate index can therefore execute code with the privileges of the user running install.py.

The installer also forwards all unrecognized command-l ...[truncated 1622 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin every direct and transitive dependency to an exact reviewed version using a generated lock file.
  2. Require cryptographic hashes during installation, such as with pip's --require-hashes.
  3. Regenerate and review the lock file through a controlled dependency-update process.
  4. Restrict installation to explicitly trusted package indexes and avoid accepting arbitrary index or trusted-host options from generic forwarded arguments.
  5. Replace unrestricted parse_known_args() forwarding with an allowlist of required pip options.
  6. Install into a dedicated virtual environment or container using a non-privileged account.
  7. Pin the Playwright package and associated browser revision, and verify browser artifact integrity through a controlled cache or trusted build pipeline.
  8. Run dependency vulnerability and provenance checks in continuous integration before publishing the Skill.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (19)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 18)May include surrounding context.

md
- See `requirements.txt` for Python packages.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill is presented as a universal crawler for BBC and other sites, but it lacks a clear privacy and integrity warning about making network requests to arbitrary domains, interacting with dynamic pages, and potentially collecting untrusted remote content. In an agent context, this omission is risky because users may invoke the skill without appreciating that it performs broad external network access and processes attacker-controlled web data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill explicitly states that it downloads images locally and stores crawled content in a hierarchical on-disk structure, but it does not warn users that running it will persist third-party content and media to local storage. This can create privacy, copyright, disk-consumption, and data-handling risks, especially when users crawl arbitrary sites without realizing files will be retained.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill usage content is presented entirely in Chinese, including headings, instructions, and warnings, without offering an alternate language or indicating that the user can choose their preferred locale. Per SQP-3, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly documented and justified.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
91% confidence
Finding

This subprocess invocation executes pip with additional arguments taken directly from unvalidated command-line input via parse_known_args(). Although shell injection is avoided because a list is passed to check_call, an attacker or unsuspecting user can supply dangerous pip flags such as alternate package indexes, trusted hosts, editable/VCS installs, or --break-system-packages, causing installation of untrusted code or unsafe modification of the environment. In an installer context, this is more dangerous because pip install routinely executes arbitrary code from packages during installation.

Content

Scanner excerpt · install.py (reported line 19)May include surrounding context.

python
if extra_args:
        cmd.extend(extra_args)
        
    subprocess.check_call(cmd)

def main():
    parser = argparse.ArgumentParser(description="Install dependencies for BBC Crawler MaxClaw")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · install.py (reported line 53)May include surrounding context.

python
print("Installing Playwright browsers...")
    try:
        # Check if playwright is installed
        subprocess.check_call([sys.executable, "-m", "playwright", "install", "chromium"])
    except Exception as e:
        print(f"Error installing playwright browsers: {e}")
        # Don't exit here, browser install failure might be non-critical if using existing browsers

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code exposes a respect_robots configuration flag but never fetches or enforces robots.txt before crawling. This is a real security/compliance issue because operators may assume restricted paths will be avoided, but the crawler will still access them, potentially hitting sensitive or disallowed endpoints and creating legal, policy, or abuse-risk exposure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language instructions in this file are presented only in Chinese, including installation and usage guidance, with no indication that the user can select another language. Under SQP-3, forcing a specific language without user opt-in is a policy concern unless the locale constraint is clearly justified, which is not documented here.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This markdown file states that images are automatically downloaded locally and links are rewritten, and later explains that crawl results are saved into an output directory. For markdown files, SQP-2 applies when behaviours affecting user data or system integrity are described without a warning; here the README explains the behavior but does not warn users that running the skill will create and modify local files.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This markdown file documents output formats and shows files being created, including article markdown files and failed_urls.txt, but it does not explicitly warn users that running the skill will save potentially large amounts of scraped data locally. For markdown files, SQP-2 applies when descriptions omit warnings about behaviors that could affect user data or system integrity; here the disk-writing behavior is described operationally but not disclosed as a caution.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The echoed description presents a specific intent tied to a BBC Crawler, but the operative command is a generic pip install -r requirements.txt that has no built-in verification that the requirements correspond to that stated purpose. This is a documentation-to-code mismatch because the script's user-facing text asserts narrower intent than the code guarantees.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

requests>=2.28.0 allows any newer version to be installed, including versions not yet reviewed by the skill author. This creates non-reproducible environments and can expose consumers to known or future dependency vulnerabilities if an unsafe release is selected.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
# BBC Crawler Dependencies

# Core dependencies
requests>=2.28.0
beautifulsoup4>=4.11.0

# Advanced Crawler (Required)

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
87% confidence
Finding

requests has known advisories, but because the manifest does not pin an exact version, it is impossible to verify whether the installed release is affected. In practice, this means deployments could silently resolve to a vulnerable version or remain on one, especially in unconstrained environments.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
90% confidence
Finding

beautifulsoup4>=4.11.0 is not strictly pinned, so installations may vary over time and across systems. While this is often operationally convenient, it weakens reproducibility and increases exposure to unreviewed upstream changes or newly introduced vulnerable releases.

Content

Scanner excerpt · requirements.txt (reported line 5)May include surrounding context.

text
# Core dependencies
requests>=2.28.0
beautifulsoup4>=4.11.0

# Advanced Crawler (Required)
crawl4ai

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The dependency crawl4ai is completely unpinned, so installation may resolve to any currently available release. Because this package reportedly has multiple known advisories and appears to be a network-capable crawler dependency, leaving it unpinned increases supply-chain risk and makes builds non-reproducible and potentially vulnerable.

Content

Scanner excerpt · requirements.txt (reported line 8)May include surrounding context.

text
beautifulsoup4>=4.11.0

# Advanced Crawler (Required)
crawl4ai
playwright>=1.30.0
nest-asyncio>=1.5.0

Unverifiable Dependency: crawl4ai has 16 known advisory(ies) (CVE-2026-57571 (Crawl4AI: Arbitrary file write (path traversal) in crawler downloads can lead to); CVE-2026-56260 (Crawl4AI: Multiple Docker API Vulnerabilities - File Write, SSRF, Auth Bypass, X); CVE-2025-28197 (Crawl4AI SSRF vulnerability) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
95% confidence
Finding

crawl4ai is flagged as having multiple known advisories, including severe issues, and the manifest does not constrain it to any reviewed version. Given that this is a crawler dependency with network-facing functionality, an unpinned vulnerable release could materially increase exposure to SSRF, file-write, or related compromise scenarios depending on how the skill uses it.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
91% confidence
Finding

playwright>=1.30.0 permits installation of arbitrary later versions, which can introduce unreviewed code changes into a browser automation component. In a crawler context, browser-driving dependencies have broad system and network interaction, so unpinned versions raise both supply-chain and operational security risk.

Content

Scanner excerpt · requirements.txt (reported line 9)May include surrounding context.

text
# Advanced Crawler (Required)
crawl4ai
playwright>=1.30.0
nest-asyncio>=1.5.0

# Removed Scrapling due to deprecation

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
89% confidence
Finding

nest-asyncio>=1.5.0 is also only minimally bounded, allowing future releases to be pulled without review. Even for lower-risk utility libraries, this can create reproducibility problems and expand the attack surface through dependency drift.

Content

Scanner excerpt · requirements.txt (reported line 10)May include surrounding context.

text
# Advanced Crawler (Required)
crawl4ai
playwright>=1.30.0
nest-asyncio>=1.5.0

# Removed Scrapling due to deprecation
# scrapling>=0.4.2

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This code file includes user-facing descriptions and runtime messages in Chinese, such as the module docstring and many printed status/error messages, without indicating that other languages are supported. Under the policy rule, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.