Back to skill

Security audit

Crawl4AI Web Crawler

Security checks for vulnerabilities and agentic risk

Overview

This web-crawling skill is mostly coherent, but it asks users to install and run powerful crawler/browser tooling with broad network, profile, and remote-LLM capabilities without enough safety scoping.

Install only in an isolated environment, pin the package and Docker image before use, avoid crawling private/internal or authenticated sites unless explicitly intended, do not reuse your normal browser profile, keep downloads and cache in disposable directories, and assume remote LLM extraction may send page content to the selected provider.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:20
Finding

Unpinned Third-Party Package Installation and Execution

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 20-22
Vulnerability Type: Uncontrolled third-party dependency installation
Risk Level: Medium

Vulnerable Code:

bash
pip install -U crawl4ai
crawl4ai-setup          # Automatically installs the Playwright browser
crawl4ai-doctor         # Verifies the installation

Technical Analysis

The installation instructions use pip install -U crawl4ai without a pinned version, cryptographic hashes, a lockfile, or a restricted package index. The -U option explicitly installs or upgrades to the newest package version available at execution time. Consequently, the installed code can differ from the version that existed when this Skill was audited.

The instructions then execute crawl4ai-setup and crawl4ai-doctor, which are entry points supplied by the installed package. Python packages may also execute build or installation-related code depending on their distribution format and build backend. If the package, one of its transitive dependencies, its publisher account, or the package repository is compromised, following these instructions could execute attacker-controlled code.

Attack Path

  1. An attacker compromises the upstream package, a transitive dependency, a publisher account, or the dependency distribution channel.
  2. The attacker publishes a malicious release that is newer than the previously trusted version.
  3. A user follows the Skill instructions and runs pip install -U crawl4ai.
  4. Pip resolves and installs the mutable malicious release or dependency.
  5. Malicious behavior executes during installation, import, or execution of crawl4ai-setup or crawl4ai-doctor.
  6. The payload runs under the privileges of the user or environment performing the installation.

Impact Assessment

Successful exploitation could provide arbitrary code execution with the installing user's privileges. Depending on those privileges and the env ...[truncated 528 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin Crawl4AI and all transitive dependencies to reviewed versions instead of using -U.
  • Maintain a lockfile or requirements file containing cryptographic hashes, and install with pip install --require-hashes.
  • Use an explicitly configured, trusted package index or an internally controlled dependency mirror.
  • Verify package provenance, publisher identity, release signatures, and expected checksums before installation.
  • Review package-provided entry points before running crawl4ai-setup or crawl4ai-doctor.
  • Perform installation in an unprivileged virtual environment or isolated container with limited filesystem access, no unnecessary secrets, and restricted network access.
  • Use automated dependency scanning and define a controlled process for reviewing and approving upgrades.

T09 · Insecure Skill Coding Practices

Warning
Location
references/api-reference.md:18
Finding

HTTPS Certificate Validation Disabled by Default

Content
View full analysis

Vulnerability Details

File Location: references/api-reference.md, line 18
Vulnerability Type: Disabled TLS certificate verification
Risk Level: Medium

Vulnerable Code:

python
ignore_https_errors: bool = True,

Technical Analysis

The documented BrowserConfig default instructs the browser to ignore HTTPS certificate errors. This weakens endpoint authentication because expired, self-signed, hostname-mismatched, or otherwise invalid certificates can be accepted rather than causing navigation to fail.

A network-positioned attacker can exploit this configuration by presenting a fraudulent certificate and substituting the target page. Because the crawler supports JavaScript-enabled browsing, authenticated browser profiles, cookies, storage state, downloads, and content extraction, substituted pages can execute attacker-controlled client-side content in the browser context and poison returned crawl data.

This setting does not by itself bypass correctly enforced application authorization or directly execute native code on the host. Its primary effect is loss of transport authenticity and integrity, with additional consequences determined by the browser configuration and data available in the active session.

Attack Path

  1. A user crawls an HTTPS URL while using the documented configuration.
  2. An attacker obtains a network interception position, such as control over a hostile Wi-Fi access point, proxy, DNS path, or compromised gateway.
  3. The attacker redirects or intercepts the target connection and presents an invalid certificate.
  4. The browser accepts the certificate because HTTPS errors are ignored.
  5. The attacker returns modified HTML, JavaScript, links, or downloadable content.
  6. The crawler processes the substituted content and may expose session-associated requests or return attacker-controlled extraction results to downstream systems.

Impact Assessment

Exploitation can c ...[truncated 733 chars]

Remediation
View remediation

Remediation Suggestions

  • Change the secure default to ignore_https_errors=False.
  • Require explicit, narrowly scoped opt-in when invalid certificates are unavoidable in a trusted development environment.
  • Display a prominent warning that certificate-error bypass must not be combined with production credentials, cookies, storage state, or persistent browser profiles.
  • For private services, install and trust the appropriate private certificate authority rather than disabling certificate validation.
  • Restrict crawling through trusted networks and authenticated proxies, and validate proxy certificates.
  • Fail closed on certificate errors and record the failure without returning untrusted page content.
  • Add tests confirming that expired, hostname-mismatched, self-signed, and untrusted certificates are rejected by the default configuration.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (10)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill presents powerful scraping, browser automation, deep crawling, screenshotting, managed browser use, proxying, and persistent profile capabilities without user-facing warnings about privacy, authorization, legal constraints, or system/network impact. In context, this omission increases the likelihood that users or downstream agents will use the skill to collect sensitive data, hit internal resources, or perform high-impact crawling without informed consent or safeguards.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger words are broad and generic (scrape, crawl, extract webpage), which can cause the skill to activate for loosely related user requests. In an agentic environment, unintended invocation of a web crawler raises the chance of unreviewed outbound requests, over-collection of data, and privacy-impacting actions without clear user intent.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 280)May include surrounding context.

python
import requests

resp = requests.post("http://localhost:11235/crawl",
    json={"urls": ["https://example.com"], "priority": 10})

task_id = resp.json()["task_id"]

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 280)May include surrounding context.

python
import requests

resp = requests.post("http://localhost:11235/crawl",
    json={"urls": ["https://example.com"], "priority": 10})

task_id = resp.json()["task_id"]

Internal Network Request

Medium
Category
Server-Side Request Forgery
Confidence
83% confidence
Finding

The documented API accepts arbitrary urls and submits them to a local crawler service, which can then fetch attacker-influenced destinations. In this skill's context, that capability is core functionality, but without restrictions or warnings it can facilitate SSRF-style access to internal services, cloud metadata endpoints, or other sensitive network locations if users or upstream agents pass untrusted targets.

Content

Scanner excerpt · SKILL.md (reported line 280)May include surrounding context.

python
import requests

resp = requests.post("http://localhost:11235/crawl",
    json={"urls": ["https://example.com"], "priority": 10})

task_id = resp.json()["task_id"]

Internal Network Request

Medium
Category
Server-Side Request Forgery
Confidence
70% confidence
Finding

Code issues a request to a loopback, link-local, or private-range host. This can reach internal services not meant to be exposed and is a common SSRF pivot.

Content

Scanner excerpt · SKILL.md (reported line 284)May include surrounding context.

json={"urls": ["https://example.com"], "priority": 10})

task_id = resp.json()["task_id"] result = requests.get(f"http://localhost:11235/task/{task_id}") print(result.json())

text

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The reference exposes persistent browser state, cookies, storage_state, user_data_dir, downloads_path, and file download features without warning that these options can store authentication material, session data, and sensitive downloaded content on disk. In a scraping tool that may access logged-in sites, this increases the chance that operators unknowingly persist credentials or sensitive artifacts in insecure locations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The API reference documents sending crawled page content to external LLM providers and handling API tokens, but it does not warn that scraped content may be transmitted off-box to third-party services. In a web-crawling skill, users may process sensitive pages, authentication-protected content, or proprietary data, so omission of a clear disclosure creates a real risk of unintended data exfiltration and policy violations.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.