Back to skill

Security audit

Decodo Scraper

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent Decodo scraping integration with expected external API use, but users should understand that scraping targets and results are handled through Decodo and that dependency pinning is weak.

Install only if you are comfortable using Decodo as the external scraping provider. Do not submit secrets, private/internal URLs, personal data, or sensitive business research targets unless sharing them with Decodo is approved. Prefer a virtual environment with reviewed pinned dependencies, rotate the Decodo token if exposed, and make sure agents treat scraped pages, Reddit posts, product listings, and subtitles as untrusted content.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unbounded Third-Party Dependency Versions

Content
View full analysis
=2.28.0 python-dotenv>=1.0.0 ``` The documented installation procedure executes: ```bash pip install -r requirements.txt ``` ### Technical Analysis Both dependencies use minimum-version constraints without upper bounds, exact version pins, or package integrity hashes. Consequently, each installation can resolve to a different dependency version selected from the configured Python package index. This does not demonstrate that either current dependency is malicious. However, it creates a supply-chain trust weakness: a compromised future release, maliciously modified package-index response, or incompatible upstream update could be installed without any corresponding change to the reviewed Skill repository. Because Python packages may execute code during installation and are imported at runtime by `tools/scrape.py`, compromise of either dependency could result in arbitrary code execution under the account installing or running the Skill. ### Attack Path 1. An attacker compromises an upstream dependency release, its publishing credentials, or a package source trusted by the environment. 2. The attacker publishes a version satisfying `requests>=2.28.0` or `python-dotenv>=1.0.0`. 3. A user follows the documented `pip install -r requirements.txt` procedure. 4. The resolver selects and installs the compromised version because no exact version or hash is enforced. 5. Malicious package code executes during installation or when imported by `tools/scrape.py`. ### Impact Assessment Successful exploitation could execute arbitrary code with the privileges of the user installing or running the Skill. This could expose the `DECODO_AUTH_TOKEN`, environment variables, accessible files, agent data, and network resources avai ...[truncated 310 chars]
Remediation
View remediation
python-dotenv== ``` 2. Generate a lock file containing transitive dependencies rather than pinning only direct dependencies. 3. Record cryptographic hashes for every permitted distribution and install with: ```bash pip install --require-hashes -r requirements.txt ``` 4. Use a controlled package index or internal artifact mirror where practical. 5. Add automated dependency vulnerability and provenance scanning to the release process. 6. Review and deliberately update locked dependencies on a defined schedule. 7. Avoid installing or running the Skill with administrative privileges. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
tools/scrape.py:73
Finding

Untrusted Scraped Content Returned Without an Agent-Safety Boundary

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (12)

Tainted flow: 'headers' from os.environ.get (line 41, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · tools/scrape.py (reported line 60)May include surrounding context.

python
payload["locale"] = args.locale

    try:
        resp = requests.post(SCRAPE_URL, json=payload, headers=headers, timeout=120)
        resp.raise_for_status()
    except requests.RequestException as e:
        status_code = e.response.status_code if e.response is not None else None

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 55)May include surrounding context.

export DECODO_AUTH_TOKEN="your_base64_token"

text

.env file

DECODO_AUTH_TOKEN=your_base64_token

text
## OpenClaw agent integration

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · tools/scrape.py (reported line 12)May include surrounding context.

python
import requests
from dotenv import load_dotenv

load_dotenv(os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", ".env"))

SCRAPE_URL = "https://scraper-api.decodo.com/v2/scrape"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The README instructs users to configure a Decodo authentication token and use tools that send user-supplied URLs, search queries, and target identifiers to Decodo's external scraping service, but it does not clearly disclose the data-flow, third-party transmission, or privacy/security implications. In an agent context, this can cause operators to unknowingly route sensitive prompts, internal URLs, or proprietary research targets to an external provider, increasing data leakage and compliance risk.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill declares required credentials and clearly relies on network access, but it does not declare any explicit tool scope such as permissions or allowed-tools. That omission weakens least-privilege controls and makes it harder for a platform or reviewer to restrict what the skill is allowed to access, increasing the chance of overbroad execution or unintended data exposure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill invites users to submit arbitrary URLs, search queries, Amazon queries, Reddit URLs, and YouTube identifiers to a third-party scraping API, but it does not prominently warn that these inputs are transmitted off-platform to Decodo. This can lead to unintended disclosure of sensitive user-provided targets, research topics, or internal URLs, especially because the skill is explicitly designed around external scraping of user-supplied input.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill transmits user-supplied search terms, URLs, and a service authentication token to an external scraping provider, but the code contains no explicit notice, confirmation step, or minimization controls. In an agent-skill context, this matters because users may assume scraping happens locally, while sensitive URLs, identifiers, or research targets are actually disclosed to a third party.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · tools/scrape.py (reported line 60)May include surrounding context.

python
payload["locale"] = args.locale

    try:
        resp = requests.post(SCRAPE_URL, json=payload, headers=headers, timeout=120)
        resp.raise_for_status()
    except requests.RequestException as e:
        status_code = e.response.status_code if e.response is not None else None

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The dependency specification uses a lower-bound version constraint (requests>=2.28.0) rather than pinning to a specific known-good release. This makes builds non-reproducible and can allow installation of an unintended vulnerable or breaking version depending on when and where the skill is deployed.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.28.0
python-dotenv>=1.0.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

requests has multiple known advisories, and because the manifest does not pin a version, there is no way to verify from this file whether deployment will resolve to a patched or affected release. In a scraping skill that performs outbound HTTP requests, this uncertainty is more relevant because the library is likely core to runtime behavior and may process attacker-controlled URLs or redirects.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The dependency specification uses a lower-bound version constraint (python-dotenv>=1.0.0) instead of an exact version. That permits environment-dependent resolution and weakens supply-chain assurance because different installations may pull different releases, including versions later found to be vulnerable.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests>=2.28.0
python-dotenv>=1.0.0

Unverifiable Dependency: python-dotenv has 2 known advisory(ies) (CVE-2026-28684 (python-dotenv: Symlink following in set_key allows arbitrary file overwrite via ); CVE-2026-28684 (python-dotenv reads key-value pairs from a .env file and can set them as environ)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
84% confidence
Finding

python-dotenv has known advisories, and the unpinned requirement prevents confirming whether the installed version is safe. While this library is usually used during configuration loading rather than network processing, vulnerable versions could still introduce file-handling or environment-manipulation risks during local execution or deployment.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.