Back to skill

Security audit

Scrapling Web Fetch

Security checks for vulnerabilities and agentic risk

Overview

This skill does what it says, but its arbitrary web-fetching is under-scoped enough that users should review it before installing.

Review before installing. Use this only for public URLs, avoid internal or tokenized links, run it in a network-restricted and isolated Python environment, and prefer pinned dependencies or a lock file. This is not malicious based on the inspected artifacts, but its URL-fetching scope should be tightened before broad use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/scrapling_fetch.py:156
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: scripts/scrapling_fetch.py, lines 156–163
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: High

python
def fetch_one(url, max_chars, as_json, overrides, Fetcher, HTML2Text):
    parts = urlparse(url)
    if parts.scheme not in ('http', 'https'):
        return {'ok': False, 'url': url, 'error': 'url must start with http:// or https://'}

    Fetcher.configure(auto_match=True)
    fetcher = Fetcher()
    page = fetcher.get(url)

Technical Analysis

The URL validation only verifies that the supplied scheme is HTTP or HTTPS. It does not prevent requests to:

  • Loopback addresses such as 127.0.0.1 or ::1
  • RFC 1918 private networks
  • Link-local addresses
  • Cloud metadata services such as 169.254.169.254
  • Internal hostnames
  • Public hostnames that resolve to restricted addresses
  • Redirect destinations that lead to restricted addresses

The requested response is subsequently converted to Markdown and returned through standard output or JSON. Therefore, an attacker who controls or influences the input URL may use the Skill as a network proxy to access resources visible from the execution environment.

Accepting arbitrary public webpage URLs is necessary for the declared content-extraction functionality. However, access to local, private, link-local, and metadata networks exceeds the minimum privileges required for ordinary public webpage extraction.

Attack Path

  1. An attacker supplies a URL referencing an internal service, loopback endpoint, cloud metadata endpoint, or attacker-controlled hostname that resolves to a restricted address.
  2. urlparse() accepts the input because its scheme is http or https.
  3. fetcher.get(url) sends the request from the environment running the Skill.
  4. Scrapling receives content from the otherwise inaccessible endpoint.
  5. The script extracts and converts t ...[truncated 1174 chars]
Remediation
View remediation

Remediation Suggestions

  • Resolve the destination hostname before making a request.
  • Reject loopback, private, link-local, multicast, reserved, and unspecified IPv4 and IPv6 addresses.
  • Explicitly deny known cloud metadata destinations, including 169.254.169.254.
  • Re-resolve and revalidate every redirect destination before following it.
  • Mitigate DNS rebinding by ensuring the validated address is the address actually used for the connection.
  • Prefer an explicit domain allowlist when the expected target sites are known.
  • Apply connection and response timeouts.
  • Limit response size before loading or converting the complete document.
  • Restrict the number of URLs accepted in batch mode.
  • Where possible, run the Skill in a network sandbox that cannot reach internal services or metadata endpoints.
  • Add automated tests covering IPv4, IPv6, alternate IP representations, redirects, internal DNS names, and DNS rebinding scenarios.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:27
Finding

Unpinned Third-Party Dependency Installation Creates Supply-Chain Risk

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 27–31
Vulnerability Type: Unpinned third-party dependencies
Risk Level: Medium

markdown
若缺失,可安装:
```bash
python3 -m pip install scrapling html2text
text

### Technical Analysis

The documented installation command retrieves mutable latest versions of `scrapling`, `html2text`, and their transitive dependencies. It provides neither version constraints nor package hashes.

No evidence of typosquatting, dependency confusion, a suspicious package index, or an already malicious package was identified in the supplied project. The risk arises because the reviewed Skill does not define which dependency versions will be installed in the future. A compromised, replaced, or unexpectedly changed upstream release could therefore execute code that was not included in this audit.

Python packages and their build systems may execute code during installation, and imported dependencies execute with the privileges of the Skill process at runtime. Consequently, following the installation instructions delegates trust to mutable upstream artifacts.

### Attack Path

1. A user follows the installation command in `SKILL.md`.
2. pip resolves the latest available package and transitive dependency versions from its configured package index.
3. A compromised or unexpectedly modified release is selected because no approved version or hash is enforced.
4. Malicious code may execute during package building, installation, import, or subsequent webpage fetching.
5. The code operates with the filesystem, process, user, and network privileges of the Python environment.

This path depends on upstream package or package-index compromise; the audit found no proof that the currently named packages are malicious.

### Impact Assessment

If an upstream dependency is compromised, potential impact includes arbitrary code execution under the installing or executing
...[truncated 381 chars]
Remediation
View remediation

Remediation Suggestions

  • Pin direct dependencies to versions that have been reviewed and tested.
  • Generate a lock file that includes all transitive dependencies.
  • Record cryptographic hashes and install with pip --require-hashes.
  • Use a trusted, explicitly configured package index.
  • Periodically scan and review both direct and transitive dependencies.
  • Install dependencies inside a dedicated virtual environment.
  • Avoid running pip with administrator or root privileges.
  • Consider maintaining a controlled internal package mirror for production deployment.
  • Update SKILL.md so its installation example references the reviewed lock or requirements file rather than unconstrained package names.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill invokes a local script that reads local reference files and performs external network requests, but the manifest does not declare any tool scope or permissions boundaries. This makes the skill's capabilities implicit rather than explicit, increasing the chance of unintended activation, unsafe execution in constrained environments, or user unawareness that URL fetching will occur.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description explains functionality but does not warn that using the skill will contact external websites and disclose the requested URL to those destinations. This can create privacy and security issues when users provide sensitive, internal, tokenized, or otherwise confidential links, because the fetch itself leaks metadata and may trigger server-side logging or tracking.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The activation guidance says to prioritize this skill whenever a user wants webpage content, text extraction, or article fetching, which is broad enough to trigger on many ordinary requests without confirming user intent or privacy expectations. In a skill that performs external fetching, overbroad routing can cause unnecessary network access to attacker-controlled URLs or accidental transmission of sensitive internal links.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The file presents core instructions primarily in Chinese while also including English text, but it does not state that the skill is intended for Chinese-speaking users or provide a language-choice mechanism. This can violate language or locale policy when a skill implicitly forces a language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The file begins with a Chinese-only functional description section and does not indicate that language selection is optional or user-controlled. Under the stated policy, locale or language constraints should either be optional or explicitly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

This file presents the selector strategy entirely in Chinese and does not indicate that users may choose another language or that the skill is intentionally limited to a Chinese-language audience. Under the policy for natural-language violations, forcing a specific language without user opt-in can be a locale-policy issue.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This markdown file documents commands that fetch content from arbitrary URLs, which implies outbound network requests and potential transmission of user-supplied targets. The usage guidance contains no warning about network activity, remote content retrieval, or related privacy/system implications.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

This code performs an HTTP request to a user-supplied URL via fetcher.get(url), which transmits network metadata and retrieves remote content. While the action is core to the script's purpose, the file itself provides no docstring, comment, or explicit user-facing notice that it will make outbound network requests when invoked.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.