Back to skill

Security audit

Website Scraper Pro

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent single-page web scraping skill, but it can fetch arbitrary URLs from the user's environment without meaningful destination limits or clear network-disclosure safeguards.

Review before installing. Use this only in a sandbox or environment with restricted outbound network access, avoid internal/private/localhost/cloud-metadata URLs, and consider pinning dependencies or using a prebuilt reviewed environment before running it on sensitive systems.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
src/utils.py:85
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: src/utils.py:85-90; request sink at src/service.py:131-144
Vulnerability Type: Server-Side Request Forgery caused by insufficient destination validation
Risk Level: High

Vulnerable Code

src/utils.py:85-90:

python
def validate_url(value: str) -> str:
    """Validate that a full HTTP or HTTPS URL was provided."""
    parsed = urlparse(value)
    if parsed.scheme not in {"http", "https"} or not parsed.netloc:
        raise ValueError(f"Invalid URL '{value}'. Use a full http:// or https:// URL.")
    return value

src/service.py:131-144:

python
async def crawl_page(url: str, use_js: bool, query: str | None) -> dict:
    """Crawl a single page and return structured output."""
    if CRAWL4AI_IMPORT_ERROR is not None:
        return {
            "error": build_setup_error("Crawl4AI is not available in this runtime.")
        }

    browser_config = BrowserConfig(headless=True, verbose=False)
    run_config = build_run_config(use_js, query)

    try:
        with redirect_stdout(io.StringIO()), redirect_stderr(io.StringIO()):
            async with AsyncWebCrawler(config=browser_config, verbose=False) as crawler:
                result = await crawler.arun(url=url, config=run_config)

Technical Analysis

validate_url() only verifies that the scheme is HTTP or HTTPS and that a network location is present. It does not reject loopback, private, link-local, multicast, reserved, or cloud metadata addresses. It also does not ensure that DNS resolution returns only public addresses.

The validated value is passed directly to AsyncWebCrawler.arun(), which makes the request from the network environment of the Skill process. Consequently, a caller can use the scraper as a server-side request primitive. Examples of accepted destinations include localhost services, RFC 1918 private addresses, and link-local metadata addresses such ...[truncated 2239 chars]

Remediation
View remediation

Remediation Suggestions

  1. Resolve the hostname before crawling and reject every address that is not globally routable. Block loopback, private, link-local, reserved, multicast, unspecified, and IPv4-mapped IPv6 variants.
  2. Explicitly deny localhost names and cloud metadata endpoints, including 169.254.169.254 and relevant IPv6 link-local addresses.
  3. Reject URLs containing ambiguous host syntax, embedded credentials, malformed ports, or unsupported hostname representations.
  4. Disable automatic redirects or validate every redirect destination using the same policy before following it.
  5. Mitigate DNS rebinding by binding the validated hostname to checked IP addresses and ensuring the actual connection does not use a newly resolved prohibited address.
  6. Apply outbound firewall or sandbox rules that prevent the browser process from connecting to loopback, private networks, metadata services, and other sensitive ranges.
  7. If operationally feasible, use an explicit domain allowlist rather than accepting arbitrary public destinations.
  8. Restrict browser subresource requests because JavaScript pages can request destinations other than the top-level URL.
  9. Add automated tests covering IPv4, IPv6, alternative address representations, redirects, and DNS rebinding scenarios.

T08 · Insecure Dependencies

Warning
Location
src/main.py:3
Finding

Unpinned Runtime Dependency Creates a Supply-Chain Risk

Content
View full analysis

Vulnerability Details

File Location: src/main.py:3-8; related setup instructions in SKILL.md
Vulnerability Type: Unpinned third-party dependency and runtime package resolution
Risk Level: Medium

Vulnerable Code

src/main.py:3-8:

python
# /// script
# requires-python = ">=3.12"
# dependencies = [
#   "crawl4ai",
# ]
# ///

Related setup commands documented in SKILL.md:

bash
uv run --with crawl4ai crawl4ai-setup
uv run --with crawl4ai python -m playwright install chromium

Technical Analysis

The inline dependency declaration specifies crawl4ai without an exact version or integrity hash. Running the documented uv run command can therefore resolve a package version available from the configured package index at execution time rather than a dependency set fixed during review.

Crawl4AI also has transitive dependencies, and the documented browser setup downloads additional browser artifacts. Without a reviewed lockfile, exact version constraints, hashes, and trusted source configuration, subsequent executions may install code or artifacts that differ from those audited.

Python dependencies execute within the local process when imported and can also perform package-defined installation or initialization behavior. A compromised upstream release, compromised package index, dependency-resolution mistake, or malicious transitive dependency could consequently execute with the permissions granted to the Skill process.

No evidence was found that the project intentionally references a malicious or typosquatted package. The finding concerns the unsafe mutability and lack of integrity controls in the dependency acquisition process.

Attack Path

  1. An upstream package account, distribution channel, configured package index, or transitive dependency is compromised, or an incompatible release is published.
  2. A user invokes the documented uv run or setup command after t ...[truncated 1252 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin Crawl4AI to a reviewed exact version rather than using an unconstrained package name.
  2. Generate and commit a lockfile containing exact transitive dependency versions.
  3. Require cryptographic hashes for downloaded Python distributions where supported.
  4. Configure uv to use only an approved, authenticated package index or internal dependency mirror.
  5. Pin and verify Playwright or Chromium versions and artifact checksums rather than downloading mutable browser artifacts implicitly.
  6. Perform dependency updates through a controlled process that includes source review, vulnerability scanning, and regression testing.
  7. Build an immutable environment or container image in CI and run the Skill from that prebuilt artifact instead of resolving packages at invocation time.
  8. Execute installation and scraping with a nonprivileged account inside a sandbox with minimal filesystem access, no unnecessary credentials, and restricted outbound networking.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill invokes a local Python script that performs outbound web requests and may access environment-dependent tooling, but the skill manifest does not declare any explicit tool scope such as allowed tools or permissions. This weakens reviewability and policy enforcement because an agent or orchestrator cannot easily constrain or pre-approve the skill's network-capable behavior from the manifest alone.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code performs a network operation by invoking the web crawler against a user-provided URL, which can transmit user-selected targets and retrieve remote data. Within this file, the operation is silent: there is no confirmation prompt, no visible logging/print statement, and stdout/stderr are explicitly suppressed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill directs the agent to fetch arbitrary user-provided URLs but does not warn that using the skill triggers outbound connections to third-party servers. That omission matters because requests may disclose IP address, user agent, timing, and other network metadata to the target site, which is especially relevant when scraping sensitive, internal, or attacker-controlled URLs.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.