Back to skill

Security audit

Scrapling Install

Security checks for vulnerabilities and agentic risk

Overview

This is a legitimate scraping skill, but it gives agents broad web-fetching, stealth/Cloudflare, proxy, login, and local file-write patterns without enough scoping or safeguards.

Install only in an isolated environment with outbound network restrictions. Use it only for sites you are authorized to access, avoid Cloudflare-solving/proxy modes unless explicitly permitted, do not hardcode proxy or login credentials, and review any recipe that writes downloads, exports, checkpoints, or deletes directories.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/scrapling_scrape.py:41
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis
None: p = argparse.ArgumentParser() p.add_argument("--url", required=True) p.add_argument("--mode", choices=["fetcher", "dynamic", "stealthy"], default="fetcher") p.add_argument("--css", help="CSS selector (supports ::text and ::attr())") p.add_argument("--xpath", help="XPath selector") p.add_argument("--first", action="store_true", help="Return only the first match") p.add_argument("--headless", action="store_true", help="Headless browser (dynamic/stealthy)") p.add_argument("--solve-cloudflare", action="store_true", help="Attempt to solve Cloudflare (stealthy session)") p.add_argument("--network-idle", action="store_true", help="Wait for network idle (dynamic session)") p.add_argument("--adaptive", action="store_true", help="Use adaptive selectors (if supported)") p.add_argument("--auto-save", action="store_true", help="Auto-save selector fingerprints (if supported)") p.add_argument("--pretty", action="store_true", help="Pretty-print JSON") args = p.parse_args() if not args.css and not args.xpath: _die("Provide --css or --xpath") url = args.url try: # Sessions are more reliable than one-shot fetchers for anything non-trivial. from scrapling.fetchers import FetcherSession, DynamicSession, StealthySession except Exception: _die( "Scrapling is not installed in this Python environment. Try:\n" " python3 -m pip install scrapling\n" "If you need browser-based fetching, you may also need:\n" " python3 -m playwright install chromium" ) if args.mode == "fe ...[truncated 4726 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
SKILL.md:13
Finding

Unpinned Third-Party Package and Browser Installation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (10)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill provides extensive guidance for network-capable scraping and MCP server invocation, but it does not declare any tool scope such as permissions or allowed-tools. In an agent environment, that omission can cause overbroad or ambiguous access to network operations, especially given the skill explicitly instructs use of external fetching, stealth fetching, proxy rotation, and anti-bot bypass features.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The documentation explicitly promotes a stealth/anti-bot fetch mode with 'Cloudflare bypass' and a flag to 'solve_cloudflare' but provides no warning about authorization, terms-of-service, or legal/ethical constraints. In an MCP skill context, this lowers friction for misuse by making evasive access patterns directly discoverable and easy to invoke, which can facilitate scraping of protected targets without informed user safeguards.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The documentation includes concrete proxy rotation examples, including an authenticated proxy URL with embedded credentials, without warning about credential leakage, logging exposure, provider trust, or legal/privacy implications. In a scraping-focused skill, this can normalize unsafe handling of proxy secrets and encourage anti-bot evasion practices without sufficient guardrails.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/spider-recipes.md (reported line 72)May include surrounding context.

md
"""Scrape via API endpoints instead of HTML."""
    
    name = "api_products"
    api_base = "https://api.example.com/v1"
    
    def start_requests(self):
        for page in range(1, 100):

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The media downloader example performs local file writes without clearly disclosing that it will persist arbitrary remote content to disk. In a skill marketed primarily as strategy and MCP-native scraping guidance, undisclosed local side effects can surprise users and increase risk of unwanted storage consumption or saving unsafe content.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest describes this skill as providing MCP-native guidance, recipes, and best practices for using the scrapling MCP server, not as a local file-writing downloader. This recipe demonstrates direct persistence of remote media to the local filesystem, which is not clearly justified by the stated guidance-focused purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The checkpoint reset example invokes recursive directory deletion without an explicit warning that the operation is destructive. Even though the sample path is fixed, users often adapt examples directly, so omitting a warning materially raises the chance of accidental data loss.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The script explicitly exposes stealth scraping and Cloudflare-solving behavior through a simple CLI flag, which operationalizes anti-bot bypass rather than merely documenting scraping concepts. In the context of an MCP skill intended to guide and enable web scraping, this lowers the barrier to unauthorized collection from protected sites and can facilitate terms-of-service evasion or abusive automation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The manifest frames the skill as providing guidance for using a scraping MCP server. The example goes beyond guidance on extraction/crawling by instructing destructive local filesystem deletion of a checkpoint directory, a capability not obviously required by the stated purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The script searches extracted content only for the specific Turkish terms "Yapay", "Üretken", "Atölye", and "Atolye". This creates a language-specific behavior in the skill logic without offering user configuration or documenting that the extraction is intentionally limited to that locale.

Content

No source excerpt is available for this finding.

Static analysis

Detected: malicious.stealth_browser_abuse

Browser automation advertises stealth/anti-detection behavior with bot-protection bypass and persistent sessions.

Critical
Code
malicious.stealth_browser_abuse
Location
SKILL.md:88