Back to skill

Security audit

crawl

Security checks for vulnerabilities and agentic risk

Overview

This remote crawling skill is mostly coherent, but it sends API keys and browsing/search data to a plaintext HTTP endpoint and trusts server-provided download URLs, so it needs review before use.

Install only after reviewing the trust boundary. Do not use the public HTTP endpoint for API keys or sensitive browsing targets; prefer a self-hosted HTTPS OpenCrawl server, avoid authenticated/internal/private pages, and understand that crawl content may be processed by remote workers and stored in R2.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
tools/crawl.py:16
Finding

API Credential and User Data Transmitted Over Plaintext HTTP

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
tools/crawl.py:60
Finding

Unrestricted Server-Controlled Download URL Enables Client-Side SSRF

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
requirements.txt:1
Finding

Open-Ended Dependency Version Prevents Reproducible Installation

Content
View full analysis
=2.28.0 ``` ### Technical Analysis The project accepts any current or future release of `requests` newer than or equal to version 2.28.0. As a result, installations performed at different times may retrieve different dependency versions without a corresponding review or source change in this project. No malicious dependency, typo-squatted package, or currently exploited version was identified during this audit. The risk arises from the lack of reproducibility and from automatically accepting future versions whose behavior or security properties have not been evaluated. Transitive dependencies are likewise not locked or hash-verified. ### Attack Path 1. A future compromised, vulnerable, or unexpectedly incompatible version of `requests` is published and satisfies `>=2.28.0`. 2. A user installs the Skill's dependencies without a lock file or hash verification. 3. The package resolver selects the affected release. 4. The Skill imports and executes that dependency while processing network requests. 5. Any resulting impact depends on the defect or malicious behavior in the selected release. ### Impact Assessment The potential privilege scope is the Python process running the Skill, including its network access and environment variables such as `OPENCRAWL_API_KEY`. The actual impact is conditional on a future dependency compromise or vulnerability; the audited requirement itself does not establish that the current `requests` package is malicious. ]]>
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (16)

Tainted flow: 'API_URL' from os.environ.get (line 16, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
92% confidence
Finding

The crawl request sends the Bearer API key to a server endpoint whose base URL is taken directly from the OPENCRAWL_API_URL environment variable. If that variable is modified by an attacker or misconfigured, the tool will transmit credentials and user-supplied crawl targets to an arbitrary host, enabling credential theft and covert data exfiltration.

Content

Scanner excerpt · tools/crawl.py (reported line 35)May include surrounding context.

python
if selector:
            body["selector"] = selector

        res = requests.post(
            f"{API_URL}/api/crawl",
            headers={
                "Authorization": f"Bearer {API_KEY}",

Tainted flow: 'API_URL' from os.environ.get (line 16, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
92% confidence
Finding

The search function posts user queries and the Bearer token to an API host controlled by the OPENCRAWL_API_URL environment variable. A malicious or poisoned environment can redirect these requests to an attacker-controlled server, exposing the API key and potentially sensitive search terms.

Content

Scanner excerpt · tools/crawl.py (reported line 81)May include surrounding context.

python
error_exit("OPENCRAWL_API_KEY environment variable not set")

    try:
        res = requests.post(
            f"{API_URL}/api/search",
            headers={
                "Authorization": f"Bearer {API_KEY}",

Tainted flow: 'API_URL' from os.environ.get (line 16, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
91% confidence
Finding

The balance check sends the Authorization bearer token to the host specified by OPENCRAWL_API_URL without trust validation. That makes the token vulnerable to exfiltration if the environment variable is attacker-influenced or if the default endpoint is replaced in deployment.

Content

Scanner excerpt · tools/crawl.py (reported line 120)May include surrounding context.

python
error_exit("OPENCRAWL_API_KEY environment variable not set")

    try:
        res = requests.get(
            f"{API_URL}/api/balance",
            headers={"Authorization": f"Bearer {API_KEY}"},
            timeout=10,

Tainted flow: 'API_URL' from os.environ.get (line 16, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · tools/crawl.py (reported line 138)May include surrounding context.

python
def check_status():
    try:
        res = requests.get(f"{API_URL}/api/status", timeout=10)
        data = res.json()
    except Exception as e:
        error_exit(str(e))

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared purpose is webpage crawling, but the documented behavior also includes multi-engine search and account/platform queries. This mismatch undermines informed consent and policy enforcement because users or orchestrators may approve the skill for one purpose while it performs broader network actions, increasing the chance of unintended data disclosure or abuse.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The README encourages users to send crawl jobs to a public server but does not clearly warn that requested URLs, rendered content, and extracted data are transmitted to and processed by an external third-party service, with results stored remotely. In a browsing/crawling skill, this omission can lead users to unknowingly expose proprietary, personal, or authenticated content to infrastructure outside their control.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The README states that community-contributed workers run real Chrome browsers and mentions cookie isolation, but it does not warn users not to crawl sensitive, internal, or authenticated pages. Because this skill routes browsing activity through external worker infrastructure, users may expose session-bearing content, confidential documents, or internal application data to untrusted systems if they assume the service is safe for private targets.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill declares required environment variables and clearly relies on outbound network access, but it does not define any explicit tool scope or permission boundaries. In an agent setting, this weakens least-privilege controls and can cause the skill to be invoked with broader capabilities than users expect, especially since it sends data to a remote third-party service.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description presents the skill as a crawler, while the body documents additional operational features such as search, balance checks, and status queries. In a security-sensitive agent ecosystem, incomplete disclosure of capabilities is itself risky because it can bypass user expectations, review workflows, or automated classification based on the manifest.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill states that rendered content is sent to an OpenCrawl server and then uploaded to Cloudflare R2, but it does not present this as a prominent user warning before use. This is dangerous because users may submit sensitive URLs or page content under the assumption of local processing, resulting in unintended transmission, storage, and possible retention of confidential data by third parties.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
86% confidence
Finding

This call transmits user-provided crawl targets and an API credential to an external service. External transmission is expected for this tool's functionality, but in the current form it still presents a real privacy and secret-handling risk because the destination is configurable and the transfer is not clearly disclosed or constrained.

Content

Scanner excerpt · tools/crawl.py (reported line 35)May include surrounding context.

python
if selector:
            body["selector"] = selector

        res = requests.post(
            f"{API_URL}/api/crawl",
            headers={
                "Authorization": f"Bearer {API_KEY}",

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

User-supplied crawl URLs and selectors are transmitted to a remote third-party service without any user-facing disclosure at execution time. In a skill context, this is dangerous because users may assume the tool operates locally while actually sending potentially sensitive internal URLs or private targets off-box to an external service.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest says the skill crawls JavaScript-rendered webpages through remote Chrome browsers, which aligns with the crawl path. However, this file also exposes separate search functionality plus API balance and platform status checks, which go beyond the stated crawling-only purpose and represent additional product capabilities.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
86% confidence
Finding

This request sends user search queries and the API token to an external service. That behavior matches the feature set, but it remains security-relevant because search terms may contain sensitive information and the endpoint is configurable, increasing the chance of accidental or malicious exfiltration.

Content

Scanner excerpt · tools/crawl.py (reported line 81)May include surrounding context.

python
error_exit("OPENCRAWL_API_KEY environment variable not set")

    try:
        res = requests.post(
            f"{API_URL}/api/search",
            headers={
                "Authorization": f"Bearer {API_KEY}",

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The dependency is specified as requests>=2.28.0, which allows installation of many future versions and does not guarantee reproducible builds. Unpinned dependencies increase supply-chain risk because different environments may resolve to different releases, including versions with breaking changes or newly introduced vulnerabilities.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.28.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

Because requests is not pinned, it is impossible to determine from this manifest whether the installed version is one affected by known advisories. In a skill that performs web crawling and likely makes outbound HTTP requests, using an unresolved dependency version can expose the runtime to request-handling flaws, credential leakage issues, or other known defects if a vulnerable release is installed.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.