Back to skill

Security audit

Scrapling AI

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward website-scraping helper, but users should treat its network access, anti-bot bypass features, MCP server, and package install as sensitive.

Install only if you are comfortable giving the agent a general-purpose web scraping tool. Use it only on sites you are allowed to access, avoid scraping internal or localhost URLs unless intended, do not pass sensitive headers or credentials, pin or review the Scrapling package before installation where possible, and treat scraped files as untrusted remote content.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:12
Finding

Unpinned Third-Party Package Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:12-25 and SKILL.md:49-53
Vulnerability Type: T08: Insecure Dependencies
Risk Level: Medium

Vulnerable Code

yaml
"install":
  [
    {
      "id": "pipx",
      "kind": "pipx",
      "package": "scrapling",
      "bins": ["scrapling"],
      "label": "Install Scrapling CLI (pipx)",
    },
    {
      "id": "python3-pip",
      "kind": "pip",
      "package": "scrapling",
      "bins": ["scrapling"],
      "label": "Install Scrapling CLI (pip)",
    },
  ],
bash
# Install CLI
pipx install scrapling
scrapling --version

Technical Analysis

The Skill installs the third-party scrapling package by name without specifying an exact version, cryptographic hash, lockfile, verified publisher, or explicitly trusted package source. Consequently, installation resolves whichever package release is current at execution time rather than the specific artifact reviewed during the audit.

Python package installation can execute package build and installation logic. The installed CLI subsequently runs with the privileges of the user or agent invoking the Skill. Although the audited file contains no evidence that the current package is malicious, this installation pattern creates a supply-chain exposure: a compromised publisher account, malicious future release, compromised transitive dependency, or package-index resolution issue could introduce attacker-controlled code after the Skill itself has been reviewed.

Attack Path

  1. An attacker compromises the upstream package publisher, a transitive dependency, or the relevant package distribution channel.
  2. The attacker publishes a malicious release that remains compatible with the unqualified package name scrapling.
  3. A user or agent follows the documented pipx install scrapling command, or the Skill framework processes the unpinned installation metadata.
  4. The p ...[truncated 914 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin scrapling to a specific, reviewed version in both installation metadata and documentation, for example scrapling==<reviewed-version>.
  2. Verify the selected release against cryptographic hashes and install with hash enforcement where supported.
  3. Record the expected official package index, upstream repository, and publisher identity so package provenance can be validated.
  4. Use a lockfile or equivalent reproducible dependency manifest that also pins transitive dependencies.
  5. Install and execute the CLI in an isolated, least-privileged environment without unnecessary credentials or access to sensitive files.
  6. Review release notes and dependency changes before updating the pin, and subject each new artifact to security scanning.
  7. Configure package tooling to use an explicitly trusted index and avoid unintended fallback to untrusted or private indexes.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (2)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill encourages scraping remote websites, setting custom headers, bypassing anti-bot protections, and exposing an MCP scrape tool, but does not warn about privacy, legal, network, or system-boundary risks. In an agent setting, this omission is meaningful because it may cause users or downstream agents to send identifying headers, access sensitive/internal endpoints, or expose a scraping capability through MCP without understanding the security implications.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill shows examples that write scraped output directly to a local file but does not warn users that scraped content may contain sensitive, copyrighted, or untrusted data. In an agent context, this can lead to inadvertent persistence of remote content on disk, creating privacy, compliance, or downstream trust issues if that content is later consumed by other tools.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.