Back to skill

Security audit

Scrapling Safe

Security checks for vulnerabilities and agentic risk

Overview

This web-scraping skill is mostly purpose-aligned, but its safety claims are not enforced well enough for the network and file-writing authority it uses.

Install only if you are comfortable running a scraper that may contact any URL the agent is given and write output anywhere under your home directory. Use it in an isolated environment, avoid private/internal URLs, choose a dedicated output folder, and prefer pinned dependencies or a reviewed lockfile before use.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scrapling.py:98
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: scrapling.py:98-125
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: High

Vulnerable Code

python
clean_url = url.strip()
self.results = []
self.output_file = output_file

if output_file and not self.validate_output_path(output_file):
    return {"error": "invalid_output_path", "message": "路径不合法", "results": []}

try:
    if mode == "get":
        # HTTP 请求模式
        page = Fetcher.get(clean_url, timeout=DEFAULT_TIMEOUT)
    elif mode == "stealthy":
        # 隐身模式
        page = StealthyFetcher.fetch(
            clean_url,
            headless=headless,
            timeout=DEFAULT_TIMEOUT,
            solve_cloudflare=solve_cloudflare
        )
    elif mode == "dynamic":
        # 浏览器自动化
        page = DynamicFetcher.fetch(
            clean_url,
            headless=headless,
            timeout=DEFAULT_TIMEOUT
        )
    elif mode == "spider":
        # 爬虫模式
        spider = Spider(
            name="demo",
            start_urls=[clean_url],
            concurrent_requests=DEFAULT_CONCURRENT,
            delay=DEFAULT_DELAY
        )

Technical Analysis

The user-controlled URL is stripped of whitespace and passed directly to HTTP, stealth, browser, or spider fetchers. The code does not validate the URL scheme, hostname, resolved IP address, destination port, or redirect targets.

Consequently, the documented restriction to public websites is not enforced in code. An attacker who can control the URL can direct the fetcher toward loopback interfaces, private networks, link-local services, or cloud instance metadata endpoints. Non-HTTP schemes may also be reachable if an underlying fetcher supports them.

A timeout limits request duration but does not prevent access to unauthorized network locations. The same destination policy must be applied after DNS resolution and to every redire ...[truncated 1184 chars]

Remediation
View remediation

Remediation Suggestions

  • Permit only explicitly supported schemes, normally http and https.
  • Parse URLs with a strict URL parser and reject malformed URLs, embedded credentials, and ambiguous host representations.
  • Resolve the hostname before connecting and reject loopback, private, link-local, multicast, unspecified, and reserved IPv4 and IPv6 ranges.
  • Explicitly block cloud metadata destinations and other environment-specific sensitive endpoints.
  • Revalidate every redirect destination and every spider-discovered URL.
  • Protect against DNS rebinding by connecting only to the validated resolved address while preserving the intended HTTP host and TLS verification.
  • Consider an explicit domain allowlist or a controlled outbound proxy when the permitted targets are known.
  • Apply network-level egress controls as defense in depth.

T09 · Insecure Skill Coding Practices

Error
Location
scrapling.py:48
Finding

User-Controlled Output Path Can Overwrite Sensitive Files in the Home Directory

Content
View full analysis

Vulnerability Details

File Location: scrapling.py:48-68, 151-153
Vulnerability Type: Arbitrary File Overwrite and Symlink Race
Risk Level: High

Vulnerable Code

python
def validate_output_path(self, path: str) -> bool:
    """验证输出路径是否在允许范围内"""
    if not path:
        return False
    
    try:
        if path.startswith("~/"):
            path = os.path.expanduser(path)
        
        abs_path = os.path.abspath(path)
        real_path = os.path.realpath(abs_path)
        real_base = os.path.realpath(ALLOWED_BASE_DIR)
        
        if real_path.startswith(real_base + os.sep) or real_path == real_base:
            return True
        
        return False
    except Exception:
        return False
python
# 保存结果
if self.output_file:
    with open(self.output_file, 'w', encoding='utf-8') as f:
        json.dump(response, f, ensure_ascii=False, indent=2)

Technical Analysis

The validator only establishes that the canonicalized path is inside, or equal to, the user's home directory. It does not restrict output to a dedicated data directory, reject sensitive filenames, enforce an allowed extension, or prevent replacement of an existing file.

The subsequent use of open(..., 'w') truncates an existing target. A path such as ~/.bashrc therefore passes validation and can be overwritten. The validator also accepts the home directory itself, although opening it as a regular file will ordinarily fail.

Validation and file opening are separate filesystem operations. An attacker with the ability to modify the relevant path between those operations could replace it with a symbolic link, creating a time-of-check/time-of-use race. Although realpath() blocks a symlink that already exists during validation and points outside the home directory, it does not make the later open operation atomic or symlink-safe.

Attack Path

  1. The attacker causes the sk ...[truncated 1242 chars]
Remediation
View remediation

Remediation Suggestions

  • Store all generated files beneath a dedicated directory such as ~/.local/share/scrapling-safe/outputs.
  • Create the output directory with restrictive permissions and reject paths outside it.
  • Accept only simple filenames rather than arbitrary paths.
  • Allow only documented extensions such as .json, if JSON is the only implemented output format.
  • Reject existing targets by default and require an explicit, separately authorized overwrite option if replacement is necessary.
  • Open files atomically using operating-system flags equivalent to O_CREAT | O_EXCL | O_NOFOLLOW.
  • Verify with fstat() that the opened descriptor is a regular file and has the expected ownership.
  • Write to a securely created temporary file in the same output directory and atomically rename it when replacement is intentionally supported.
  • Do not treat the home directory itself as a valid output-file target.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:48
Finding

Documented Commands Bypass the Project's Local Safety Wrapper

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:48-70
Vulnerability Type: Security-Control Bypass Through Inconsistent Entrypoints
Risk Level: Medium

Vulnerable Code

bash
# HTTP 请求抓取
scrapling get 'https://example.com' --output ~/result.json

# 隐身模式抓取
scrapling stealthy 'https://example.com' --output ~/result.json

# 浏览器自动化(动态内容)
scrapling dynamic 'https://example.com' --output ~/result.json
bash
# 使用 CSS 选择器
scrapling get 'https://quotes.toscrape.com' --css-selector '.quote' --output ~/quotes.json

# 提取特定字段
scrapling get 'https://quotes.toscrape.com' --css-selector '.quote .text' --output ~/text.txt
bash
# 隐身模式 + 解决 Cloudflare
scrapling stealthy 'https://nopecha.com/demo/cloudflare' --solve-cloudflare --output ~/result.json

# 并发抓取(限制为 1)
scrapling spider 'https://example.com' --concurrent 1 --output ~/crawl.json

Equivalent direct upstream CLI examples also appear in README.md:39-47.

Technical Analysis

The project implements local controls in scrapling.py, including output-path validation, fixed timeout values, and a fixed spider concurrency value. However, the instructions consistently invoke the installed scrapling executable rather than this Python wrapper.

Based on the reviewed project, there is no packaging configuration or launcher demonstrating that the scrapling command resolves to scrapling.py. The installation instructions identify Scrapling as a third-party dependency, so these examples appear to invoke the upstream dependency's CLI. The local controls therefore cannot be assumed to apply to the documented execution path.

This discrepancy is security-relevant because users and Agents are told that path, timeout, and concurrency restrictions exist while being directed toward a different entrypoint whose behavior is outside this wrapper.

Attack Path

  1. A user or Agent loads the skill instructions and follows a docum ...[truncated 920 chars]
Remediation
View remediation

Remediation Suggestions

  • Define a unique project entrypoint, such as scrapling-safe, that unambiguously invokes this wrapper.
  • Update every command in SKILL.md and README.md to use that entrypoint.
  • Add packaging metadata or a launcher script that maps the documented command to scrapling.py.
  • Add integration tests that execute the documented command and verify that path, timeout, URL, and concurrency controls are enforced.
  • Avoid making safety claims for behavior delegated directly to an upstream executable unless those guarantees have been independently verified.
  • Align the wrapper's argument format with its documentation; the current wrapper expects --mode, while the examples use positional modes.

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unbounded Third-Party Dependency Versions Create Supply-Chain Risk

Content
View full analysis

Vulnerability Details

File Location: requirements.txt:1-2
Vulnerability Type: Unpinned Dependencies and Non-Reproducible Installation
Risk Level: Medium

Vulnerable Code

text
scrapling>=1.0.0
requests>=2.31.0

Related installation instructions in SKILL.md:77-78 are:

text
pip install scrapling[fetchers]
scrapling install

Technical Analysis

Both Python dependencies use open-ended lower bounds. A future installation may therefore select versions that were never reviewed with this project. No lockfile or package hashes are present to ensure reproducible resolution or artifact integrity.

The documentation additionally directs the user to install the fetchers extra and run scrapling install, which may retrieve browser components or other upstream-managed assets. The reviewed files do not pin or verify those components.

This does not prove that any currently named dependency is malicious. It is an insecure dependency-management practice that increases exposure to compromised releases, future regressions, or incompatible behavior.

Attack Path

  1. A user follows the installation instructions or installs from requirements.txt.
  2. The package resolver selects the newest version satisfying each open-ended constraint.
  3. A newer release, transitive dependency, or externally installed browser component differs from the audited version.
  4. Installation or runtime imports execute that third-party code with the user's privileges.
  5. If the selected artifact is compromised, it can access the same files, network, and environment available to the skill process.

Impact Assessment

A compromised dependency can execute arbitrary code with the privileges of the user performing installation or running the skill. This could expose user files, environment variables, network access, and Agent data. The finding represents potential supply-chain exposure rather than evidence that ...[truncated 71 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin every direct dependency to a reviewed exact version.
  • Generate and commit a lockfile that includes all transitive dependencies.
  • Require cryptographic hashes for downloaded Python artifacts where supported.
  • Pin and verify browser binaries or other components installed by scrapling install.
  • Use a trusted package index and prevent unexpected fallback to untrusted repositories.
  • Run automated dependency vulnerability and provenance checks in continuous integration.
  • Review upgrades explicitly rather than accepting arbitrary future versions.
  • Install and run the skill in an isolated, least-privileged environment.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill advertises and instructs use of network access and file output, but it does not declare any explicit tool scope such as allowed-tools or permissions. That creates a governance gap: an agent or platform may permit broader-than-intended capabilities, and users cannot clearly verify what resources the skill is supposed to access. In this context, the risk is amplified by the scraping and file-writing behavior, even though the markdown claims safety restrictions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The file consistently presents the skill name, description, safety guidance, and usage instructions in Chinese. This can violate a language/locale policy when the skill implicitly forces one language for all users and does not document any opt-in or alternative language support.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module documentation asserts that the tool only accesses public websites, but the implementation accepts any user-supplied URL and performs network requests without scheme, host, or address validation. In a skill context, this can mislead users or orchestrators into treating the tool as safer than it is, while still enabling access to internal services, localhost, or cloud metadata endpoints if invoked with attacker-controlled input.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

This code performs outbound HTTP/browser-based requests to user-supplied URLs and may transmit system/browser metadata during scraping, but the execution path contains no confirmation prompt or user-facing notice. Although the module docstring describes the tool generally, the runtime behavior does not disclose the network access when the operation occurs.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

The dependency specifier scrapling>=1.0.0 is unpinned, so installs may resolve to different future versions with unknown security properties or breaking behavior. In a scraping/automation skill, this increases supply-chain risk because the package may gain vulnerable or unsafe transitive behavior without review.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
scrapling>=1.0.0
requests>=2.31.0

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
98% confidence
Finding

The dependency specifier requests>=2.31.0 is unpinned, allowing installation of any later version, including versions that may introduce vulnerabilities or incompatible changes. Because this skill performs HTTP requests, any flaw in requests directly affects network communication, credential handling, redirect behavior, and TLS-related trust decisions.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
scrapling>=1.0.0
requests>=2.31.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
95% confidence
Finding

requests has multiple known advisories, and because the manifest does not pin an exact version, it is impossible to verify whether deployment will select a patched or vulnerable release. In this skill's context, requests is central to fetching remote content, so exploitable issues could affect confidentiality of credentials, request integrity, or handling of attacker-controlled URLs.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file's natural-language documentation and user-facing messages are written in Chinese, which effectively forces a specific language without any opt-in or locale selection. Under the stated policy, language constraints should be optional or explicitly justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.