Back to skill

Security audit

Scrape Web

Security checks for vulnerabilities and agentic risk

Overview

This web scraping skill has a clear purpose, but it needs Review because it can fetch any URL reachable from the host and installs mutable third-party scraping components.

Review before installing in any sensitive environment. Use it only where outbound scraping from the host is acceptable, avoid internal/admin/cloud-metadata URLs, treat `--out` as a local file write that may overwrite data, and pin or review dependencies before relying on it.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/scrape_web.py:15
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: scripts/scrape_web.py, lines 15-16, 40, and 60
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: High

Vulnerable Code

python
def fetch_via_http(url: str) -> str:
    try:
        import httpx
    except Exception as e:
        raise RuntimeError("Missing dependency. install with: pip install httpx") from e
    r = httpx.get(
        url,
        follow_redirects=True,
        timeout=30,
        headers={
            # 尽量告诉服务器我们要文本
            "Accept": "text/plain,text/markdown,text/html,*/*",
            "User-Agent": "Mozilla/5.0",
        },
    )
    r.raise_for_status()
    # 自动按响应头编码解码;若为空,httpx 会做合理推断
    return r.text
python
page = StealthyFetcher.fetch(url, headless=True, network_idle=True)
python
parser.add_argument("--url", required=True, help="Target URL")

Technical Analysis

The user-controlled --url value is passed directly to both httpx.get() and StealthyFetcher.fetch() without validating the URL scheme, destination hostname, resolved IP address, or destination port.

The direct HTTP implementation also enables follow_redirects=True. Consequently, validating only an initial public URL would not be sufficient: an attacker-controlled public endpoint could redirect the request to a loopback, private, link-local, or cloud metadata address.

The affected process can therefore be used as a network proxy to access resources available from the host's network context but unavailable to the attacker directly.

Attack Path

  1. An attacker supplies a URL targeting an internal resource, such as a loopback service, private-network host, or cloud metadata endpoint.
  2. Alternatively, the attacker supplies a public URL that redirects to an internal destination.
  3. The URL is forwarded without destination validation to Scrapling or httpx.
  4. The request executes with the ...[truncated 752 chars]
Remediation
View remediation

Remediation Suggestions

  • Accept only explicitly permitted schemes, normally https and, if required, http.
  • Reject URLs containing embedded credentials or malformed hostnames.
  • Resolve the hostname before connecting and reject loopback, private, link-local, multicast, unspecified, and reserved addresses for both IPv4 and IPv6.
  • Protect against DNS rebinding by ensuring the address actually used for the connection remains within the validated address set.
  • Disable automatic redirects or validate every redirect destination using the same scheme, hostname, port, and resolved-address controls.
  • Restrict destination ports to those required for web scraping.
  • Prefer an explicit domain allowlist where the expected targets are known.
  • Run the scraper in an outbound network sandbox that cannot access cloud metadata endpoints, internal management networks, or local administrative services.
  • Apply the same validation policy to both StealthyFetcher.fetch() and the httpx fallback so that error handling cannot bypass the restrictions.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:19
Finding

Unpinned Executable Third-Party Dependencies Create Supply-Chain Risk

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 19-21
Vulnerability Type: Unpinned third-party dependencies and mutable installation behavior
Risk Level: Medium

Vulnerable Code

bash
pip install "scrapling[all]"
scrapling install
pip install httpx

Technical Analysis

The installation instructions install unconstrained versions of scrapling, its optional dependencies, and httpx. Package resolution can therefore produce different code over time, including changed transitive dependencies, without any corresponding change to the reviewed Skill.

The instructions also execute scrapling install, which may download and install additional runtime components. The versions, sources, and integrity values of those components are not documented or pinned in the project.

This prevents reproducible installation and expands the trusted supply chain beyond the source files included in the audit. The audit found no evidence that the currently named packages are malicious; the issue is the unsafe, mutable dependency acquisition process.

Attack Path

  1. A user follows the documented installation commands.
  2. pip resolves the latest package versions and their transitive dependencies from the configured package index.
  3. Package installation hooks or imported runtime code execute with the privileges of the installing or invoking user.
  4. The additional scrapling install command retrieves or installs further components not pinned by the project.
  5. If a future release, transitive dependency, package source, or downloaded component is compromised, malicious code can execute locally without changes to this Skill repository.

Impact Assessment

A compromised dependency or installer-delivered component could execute code with the privileges of the user performing installation or running the Skill. Depending on that user's permissions, this could permit access to local files, environment va ...[truncated 317 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin exact, reviewed versions of all direct dependencies.
  • Generate and commit a lock file that includes resolved transitive dependencies.
  • Require cryptographic hashes for downloaded Python distributions, such as through a hash-locked requirements file and pip --require-hashes.
  • Install packages only from explicitly trusted package indexes.
  • Document the exact components, versions, download sources, and integrity checks used by scrapling install.
  • Prefer a prebuilt, versioned runtime image whose browser and system components have already been verified.
  • Add automated dependency vulnerability and provenance scanning to the release process.
  • Review and deliberately update the lock file rather than resolving unrestricted releases during routine installation.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (6)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill requires code execution and documents both outbound web requests and writing scraped results to disk, but it does not declare any explicit tool scope such as permissions or allowed-tools. This weakens policy enforcement and can let the skill operate with broader capabilities than users or the platform may expect, increasing the chance of unintended network access or file creation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest description and usage text are presented in Chinese, which imposes a specific language on users without opt-in. Under the policy, language constraints should either offer user choice or be clearly documented as necessary for a region-specific purpose.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The manifest description scopes the skill to obtaining webpage content via Python + Scrapling with simple selectors. The implemented CLI adds a separate file-write capability through --out, which changes the skill from pure retrieval to local data export/persistence.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill instructions describe fetching arbitrary external URLs and optionally saving the scraped content to disk, but they do not warn users that this triggers outbound requests and local file writes. That omission can mislead users about privacy, data handling, or the side effects of running the skill, especially in environments where network activity and persistence are sensitive.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The function performs an outbound HTTP request to the user-supplied URL and sends request metadata such as headers, but the code provides no inline disclosure beyond the generic CLI description. Network transmission is one of the operations that should have some visible warning or explanation in code when not otherwise documented.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The code writes scraped content to an arbitrary path provided via --out, but there is no confirmation prompt, warning comment, or user-facing disclosure about overwriting or creating a file. For a code file, file writes are safety-relevant operations that should have at least some visible notice unless already documented elsewhere.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.