Back to skill

Security audit

Scrapling Fetch Basic

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed web-scraping skill, but it needs review because it can fetch arbitrary URLs and promotes stealth anti-bot bypass without clear safeguards.

Install only if you are comfortable running a scraper that can access any URL reachable from your machine or agent runtime. Use it only on sites you are authorized to access, avoid stealth mode unless you have permission, and do not let untrusted content choose URLs for the tool.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/scrapling_fetch.py:162
Finding

Arbitrary URL Fetching Enables Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: scripts/scrapling_fetch.py:42, scripts/scrapling_fetch.py:80-86, and scripts/scrapling_fetch.py:162-171
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: High

The command-line URL is passed directly to Scrapling's HTTP and browser-based fetchers without validating the URL scheme, destination hostname, resolved IP address, port, or redirect targets.

Vulnerable Code

User-controlled URL declaration and dispatch:

python
parser.add_argument("url", help="目标网页 URL")
parser.add_argument("max_chars", nargs="?", type=int, default=30000,
                    help="最大输出字符数(默认: 30000)")
parser.add_argument("--mode", choices=["basic", "stealth"], default="basic",
                    help="抓取模式: basic(快速)/ stealth(隐身,绕过反爬)")
parser.add_argument("--json", action="store_true", help="JSON 格式输出")
parser.add_argument("--debug", action="store_true", help="显示调试信息")

args = parser.parse_args()

try:
    # 抓取页面
    if args.mode == "stealth":
        html, selector = fetch_stealth(args.url, args.debug)
    else:
        html, selector = fetch_basic(args.url, args.debug)

Direct request in basic mode:

python
page = Fetcher.get(url)

Direct request in stealth mode:

python
try:
    page = StealthyFetcher.fetch(
        url,
        headless=True,
        network_idle=True,
        wait=3,
    )

Technical Analysis

The url positional argument is entirely caller-controlled. Both fetching modes use that value as a network destination without applying an allowlist or rejecting loopback, private, link-local, reserved, or otherwise sensitive addresses.

Consequently, when this skill operates in a network environment more privileged than the caller, it can act as a request proxy into that environment. Potential targets include:

  • Services bound to the runtime's loopback interface.
  • Private RFC1918 ne ...[truncated 2313 chars]
Remediation
View remediation

Remediation Suggestions

Introduce a centralized URL validation and safe-fetch layer used by both fetching modes:

  1. Parse URLs with a standards-compliant parser and permit only explicitly required schemes, normally http and https.
  2. Reject malformed URLs, embedded credentials, missing hostnames, ambiguous numeric IP representations, and unsupported ports.
  3. Resolve the hostname before connecting and reject every address classified as loopback, private, link-local, multicast, reserved, unspecified, or otherwise non-global.
  4. Validate every returned DNS address rather than checking only the first address.
  5. Disable automatic redirects where possible. If redirects are required, apply the complete validation process to every redirect target and enforce a low redirect limit.
  6. Protect against DNS rebinding by ensuring the validated address is the address actually used for the connection. Prefer network-layer egress controls over relying exclusively on application checks.
  7. Use an explicit hostname allowlist when the expected target set is known.
  8. Block access to cloud metadata, localhost, and private networks through firewall or container egress policy as defense in depth.
  9. Apply request timeouts, response-size limits, and concurrency limits to reduce denial-of-service exposure.
  10. Add tests covering loopback addresses, IPv6 loopback, private ranges, link-local ranges, alternate IP encodings, attacker-controlled redirects, and DNS rebinding scenarios.
  11. Return a generic validation error without fetching or disclosing prohibited destination details.

Example validation should occur before either of these calls:

python
safe_url = validate_public_http_url(args.url)

if args.mode == "stealth":
    html, selector = fetch_stealth(safe_url, args.debug)
else:
    html, selector = fetch_basic(safe_url, args.debug)

The same checks must also be enforced inside the network layer for redirect destinations rather ...[truncated 48 chars]

Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (5)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill advertises and documents network-capable behavior but does not declare any tool scope such as permissions or allowed-tools. This weakens security boundaries and reviewability because an agent may invoke network access without an explicit, user-visible capability declaration.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill explicitly promotes stealth mode and anti-bot bypass behavior but does not warn users about legal, contractual, or abuse risks when targeting protected websites. In context, this increases the chance of misuse because the feature is framed as a normal capability rather than a sensitive one requiring user caution and authorization.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill description is written entirely in Chinese and presents the capability only in that language, with no indication that other languages are supported or that the user can opt in to a locale preference. Under the policy, language or locale constraints should either provide user choice or be explicitly documented as justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code file contains natural-language strings such as the module docstring usage examples and descriptions exclusively in Chinese. The policy forbids forcing a specific language without offering the user a choice or clearly documenting a justified locale constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The description and most of the user-facing documentation are presented in Chinese, which effectively enforces a specific language for users of the skill. The file does not offer an alternative language option or explain that the skill is intentionally region- or locale-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.