Back to skill

Security audit

Scrapling MCP

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed web-scraping helper with expected network and browser automation behavior, and I found no hidden data theft, backdoor, or automatic destructive behavior.

Install and run this only in a virtual environment or container, pin dependency versions where practical, and use the scraping, proxy, and Cloudflare-solving features only on sites you own or are explicitly authorized to test. Review any workflow that saves files or deletes checkpoint directories before letting an agent run it.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:15
Finding

Unpinned Third-Party Packages and Browser Components

Content
View full analysis
Remediation
View remediation
" ``` 2. Generate and maintain a lockfile containing exact transitive versions and cryptographic hashes. Install with hash enforcement where practical: ```bash python -m pip install --require-hashes -r requirements.lock ``` 3. Use an explicitly trusted package index and disable unintended fallback indexes to reduce dependency-confusion exposure. 4. Pin a compatible Playwright version so its expected Chromium revision is deterministic. Document the expected browser revision and validate downloaded artifacts through the verification mechanism supported by the distribution process. 5. Run dependency vulnerability and provenance checks in CI. Review and regenerate the lockfile deliberately when upgrading rather than automatically accepting current releases. 6. Install and run the MCP server in an isolated virtual environment or container under a non-privileged account. Restrict filesystem access, secrets, and network destinations to those required for authorized scraping. 7. Update both `SKILL.md` and `references/mcp-setup.md` so users are not directed to bypass the pinned, integrity-checked installation workflow. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (9)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill clearly enables network-capable scraping workflows through MCP and CLI examples, but it does not declare an explicit tool scope such as allowed tools or permissions. That creates an authorization gap where an agent may invoke networked capabilities without clear policy boundaries, increasing the chance of unintended web access or misuse in environments that rely on metadata for enforcement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill provides concrete anti-bot bypass guidance, including stealth fetching, TLS fingerprint impersonation, proxy rotation, and automated Cloudflare/Turnstile solving, without presenting a prominent warning about legal, privacy, account, and authorization risks near those instructions. In context, this makes operational misuse easier because the skill is not merely descriptive; it offers actionable escalation steps for evading defenses.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file describes tools that fetch remote pages, perform browser-based requests, and crawl sites, but it does not provide any user-facing warning about network activity, requests to third-party sites, or the possibility of transmitting headers and other request data. Under the markdown-specific SQP-2 criteria, descriptions of behaviors that can affect privacy or system/network integrity should disclose those effects.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/spider-recipes.md (reported line 72)May include surrounding context.

md
"""Scrape via API endpoints instead of HTML."""
    
    name = "api_products"
    api_base = "https://api.example.com/v1"
    
    def start_requests(self):
        for page in range(1, 100):

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The media downloader recipe goes beyond passive scraping guidance and demonstrates writing attacker-controlled remote content to the local filesystem. In an MCP/agent context, this expands the skill from extraction into local side effects, which can enable disk abuse, unsafe file persistence, or accidental handling of untrusted content without safeguards.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The pause/resume reset example includes unconditional recursive deletion of a local directory with shutil.rmtree. Even though the path is hardcoded in the example, it normalizes destructive filesystem operations inside a skill whose stated purpose is scraping strategy and increases risk that an agent or user will adapt and execute deletion without proper confirmation or path validation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The markdown recipe shows recursive directory deletion without any warning, confirmation step, or discussion of consequences. In agent-facing documentation, omission of safety guardrails can lead to unsafe automation patterns where destructive actions are copied directly into workflows.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The recipe writes downloaded media to disk without noting that it modifies the local filesystem or that the content is untrusted remote data. While common in scraping contexts, presenting this without warnings encourages silent side effects and may lead to storage of malicious or undesired files.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This code file exposes an --auto-save option whose help text says it will auto-save selector fingerprints, implying filesystem writes. There is no confirmation prompt, explicit runtime disclosure, or warning about where data will be written or what will be persisted.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.