Back to skill

Security audit

Decodo Web Scraper

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward Decodo web-scraping integration that sends user-selected targets to Decodo using a declared API token, with some normal privacy and dependency caution needed.

Install only if you are comfortable sending the requested URLs, search terms, video IDs, Reddit/Amazon targets, request metadata, and your Decodo API token to Decodo. Avoid using it on private/internal URLs or sensitive investigations unless your Decodo account and policies allow that data sharing, and prefer pinned dependency versions in controlled deployments.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Third-Party Dependencies Permit Unreviewed Future Releases

Content
View full analysis

Vulnerability Details

File Location: requirements.txt, lines 1–2
Vulnerability Type: Insecure dependency version constraints
Risk Level: Medium

Complete Code Snippet:

text
requests>=2.28.0
python-dotenv>=1.0.0

The documented installation command in README.md, lines 50–53, consumes these constraints directly:

bash
pip install -r requirements.txt

Technical Analysis

Both dependencies use open-ended minimum-version constraints. Consequently, installations may resolve to future package versions that were not part of this audit. The requirements file also supplies no package hashes, so package integrity is not verified against a reviewed artifact.

This is an avoidable supply-chain weakness rather than evidence that the currently named packages are malicious. Exploitation would require compromise of a permitted dependency release, its distribution channel, or the package-resolution environment.

Attack Path

  1. An attacker compromises a future release of requests or python-dotenv, their publishing credentials, or a package-distribution path used by the installer.
  2. The compromised release retains a version satisfying the applicable >= constraint.
  3. A user follows the documented pip install -r requirements.txt procedure.
  4. The resolver selects and downloads the compromised, unreviewed release.
  5. Malicious package installation or runtime code executes with the privileges of the user or service installing or invoking the Skill.

Impact Assessment

Successful exploitation could provide arbitrary code execution under the account that installs or runs the Skill. The resulting access could include files, environment variables, and network resources available to that account. In this Skill’s runtime context, that may include access to DECODO_AUTH_TOKEN. The issue does not independently provide privilege escalation beyond the installer or runtime account.

...[truncated 330 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace open-ended constraints with exact versions that have been reviewed, for example:
    text
    requests==<reviewed-version>
    python-dotenv==<reviewed-version>
    
  2. Generate and commit cryptographic hashes for all direct and transitive dependencies.
  3. Install with hash verification enabled:
    bash
    pip install --require-hashes -r requirements.txt
    
  4. Use a lock-generation workflow that records the complete transitive dependency graph for each supported Python environment.
  5. Update dependencies through reviewed pull requests with automated vulnerability, provenance, and integrity checks.
  6. Perform installation and execution under a dedicated, least-privileged account without unnecessary filesystem or secret access.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (12)

Tainted flow: 'headers' from os.environ.get (line 41, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · tools/scrape.py (reported line 60)May include surrounding context.

python
payload["locale"] = args.locale

    try:
        resp = requests.post(SCRAPE_URL, json=payload, headers=headers, timeout=120)
        resp.raise_for_status()
    except requests.RequestException as e:
        status_code = e.response.status_code if e.response is not None else None

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 55)May include surrounding context.

export DECODO_AUTH_TOKEN="your_base64_token"

text

.env file

DECODO_AUTH_TOKEN=your_base64_token

text
## OpenClaw agent integration

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · tools/scrape.py (reported line 12)May include surrounding context.

python
import requests
from dotenv import load_dotenv

load_dotenv(os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", ".env"))

SCRAPE_URL = "https://scraper-api.decodo.com/v2/scrape"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README clearly states this skill sends user-supplied URLs, search queries, and an authentication token to Decodo's external scraping service, but it does not warn users that their inputs and metadata leave the local agent environment. In an agent-tool context, this can lead to unintended disclosure of sensitive prompts, internal URLs, or private investigation targets to a third party, especially if an agent invokes the tool automatically.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill exposes network access and reads a credential from the environment, but the manifest does not declare any explicit tool scope such as permissions or allowed-tools. That weakens review-time and runtime transparency, making it easier for a caller or platform to underestimate the skill's ability to exfiltrate user-supplied data or use secrets over the network.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill encourages users to submit arbitrary URLs, search terms, Reddit links, Amazon links, and YouTube IDs to a third-party scraping API, but it does not prominently warn that this input is transmitted to Decodo. This can cause inadvertent disclosure of sensitive user queries, internal URLs, or proprietary targets to an external service, especially because the skill is explicitly designed around broad scraping of user-provided inputs.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The tool sends user-provided queries and URLs, along with an authorization credential, to a third-party scraping service without any explicit in-code notice, confirmation, or privacy guardrails. In an agent-skill context, this can cause users or calling systems to unknowingly transmit sensitive URLs, search terms, or internal targets to an external processor.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · tools/scrape.py (reported line 60)May include surrounding context.

python
payload["locale"] = args.locale

    try:
        resp = requests.post(SCRAPE_URL, json=payload, headers=headers, timeout=120)
        resp.raise_for_status()
    except requests.RequestException as e:
        status_code = e.response.status_code if e.response is not None else None

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency specifier requests>=2.28.0 is unpinned, so builds may resolve to different versions over time and undermine reproducibility and patch assurance. In a security-sensitive skill that performs network scraping, leaving a widely used HTTP library floating increases supply-chain and exposure risk because vulnerable or behavior-changing versions cannot be ruled out consistently.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.28.0
python-dotenv>=1.0.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

requests has multiple published advisories, and because the manifest does not pin a version, it is impossible to verify whether deployed environments will install a patched or affected release. In a web-scraping skill that makes outbound HTTP requests to potentially attacker-controlled URLs, this uncertainty is more dangerous than in an offline-only tool because request-handling flaws can be directly exposed.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The dependency specifier python-dotenv>=1.0.0 is also unpinned, which makes installations non-reproducible and prevents assurance about exactly which code will run in different environments. Because dotenv libraries interact with local configuration and files, version drift can introduce vulnerable behavior or unsafe parsing unexpectedly.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests>=2.28.0
python-dotenv>=1.0.0

Unverifiable Dependency: python-dotenv has 2 known advisory(ies) (CVE-2026-28684 (python-dotenv: Symlink following in set_key allows arbitrary file overwrite via ); CVE-2026-28684 (python-dotenv reads key-value pairs from a .env file and can set them as environ)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
86% confidence
Finding

python-dotenv has known advisories, and the unpinned manifest means the installed version may vary and could include a vulnerable release. While this package is not the core network attack surface, it can affect local file and environment-variable handling, so unverifiable versions still create avoidable supply-chain and configuration risk.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.