Back to skill

Security audit

E-commerce Data Scraper Pro

Security checks for vulnerabilities and agentic risk

Overview

This is a mostly coherent data-scraping skill, but it needs review because it can send supplied bearer tokens and fetch arbitrary internal or external URLs without clear safeguards.

Install only if you are comfortable running a scraper that can contact any URL you or an agent provides. Use HTTPS endpoints, avoid production tokens, do not target private/internal services, keep output paths explicit, and consider running it in a network-restricted sandbox with pinned dependencies.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/data-scraper.py:82
Finding

Unrestricted URL Fetching Enables Server-Side Requests to Internal Resources

Content
View full analysis

Vulnerability Details

File Location: scripts/data-scraper.py, lines 82 and 147
Vulnerability Type: Server-Side Request Forgery (SSRF)-like unrestricted network access
Risk Level: Medium

Vulnerable Code

python
response = requests.get(url, headers=headers, timeout=30)

The same unrestricted request behavior is present in the API retrieval function:

python
response = requests.get(endpoint, headers=headers, timeout=30)
response.raise_for_status()

result["status"] = "success"
result["data"] = response.json()

Technical Analysis

Both url and endpoint originate from command-line arguments or a user-provided URL list. They are passed directly to requests.get() without validating the URL scheme, destination hostname, resolved IP address, or redirect destination.

Arbitrary public URL retrieval is necessary for the declared scraping functionality. However, the implementation does not distinguish public websites from loopback, private, reserved, link-local, or cloud metadata addresses. Consequently, the process can issue requests to resources that are reachable from the Agent environment but unavailable to an external party.

The API command is particularly significant because it stores the parsed JSON response in result["data"], after which the result is printed or written to an output file. The scraping command does not currently return the fetched response body, but it can still interact with internal services and reveal limited status or error information.

Attack Path

  1. An attacker causes an Agent or user to invoke the Skill with a malicious --endpoint, --url, or entry in --urls-file.
  2. The supplied destination points to a local service, private network host, link-local service, or cloud metadata endpoint.
  3. requests.get() connects to that destination using the network privileges of the Skill runtime.
  4. For an API endpoint returning JSON, the response is ...[truncated 1038 chars]
Remediation
View remediation

Remediation Suggestions

  1. Parse every destination before making a request and permit only explicitly supported schemes, preferably HTTPS.
  2. Resolve the hostname and reject every resolved address belonging to loopback, private, link-local, reserved, multicast, or unspecified address ranges.
  3. Explicitly block cloud metadata destinations, including link-local metadata addresses and provider-specific metadata hostnames.
  4. Disable automatic redirects or validate the scheme, hostname, and resolved IP address of every redirect target before following it.
  5. Consider requiring an explicit hostname allowlist when the Skill runs in an environment with access to sensitive internal services.
  6. Apply the same validation to every entry loaded from --urls-file.
  7. Add tests covering IPv4, IPv6, alternative numeric IP representations, DNS rebinding scenarios, and redirects from public hosts to private addresses.
  8. Run the Skill in a network sandbox that blocks access to internal and metadata networks unless such access is explicitly required.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/data-scraper.py:137
Finding

Bearer Credentials May Be Transmitted to Arbitrary Plaintext HTTP Endpoints

Content
View full analysis

Vulnerability Details

File Location: scripts/data-scraper.py, lines 137–147
Vulnerability Type: Insecure transmission of authentication credentials
Risk Level: Medium

Vulnerable Code

python
# 处理认证
auth = kwargs.get("auth")
if auth:
    if auth.startswith("Bearer "):
        headers["Authorization"] = auth
    elif ":" in auth:
        from requests.auth import HTTPBasicAuth
        username, password = auth.split(":", 1)
        result["auth"] = "basic"

response = requests.get(endpoint, headers=headers, timeout=30)

Technical Analysis

When authentication begins with Bearer , the implementation places the complete credential in the HTTP Authorization header. The endpoint is independently controlled through the --endpoint argument, and no validation requires HTTPS or restricts credentials to a trusted origin.

As a result, a bearer token can be sent to an arbitrary http:// endpoint. Plaintext HTTP provides no transport confidentiality or server authentication, allowing credential interception or modification by parties positioned on the network path. A malicious endpoint operator can also directly collect the supplied token.

This network transmission is part of the documented API authentication feature rather than hidden exfiltration. The security issue is the absence of transport and destination safeguards around sensitive credentials.

The Basic authentication branch parses a username and password but does not attach them to the request. That behavior is a functional defect and does not itself transmit the parsed Basic credentials.

Attack Path

  1. An attacker supplies or recommends an API endpoint using plaintext HTTP or a server under the attacker's control.
  2. The user or Agent invokes the documented API command and supplies a valid bearer token through --auth.
  3. The code copies the token into the Authorization header without checking the destination or tra ...[truncated 840 chars]
Remediation
View remediation

Remediation Suggestions

  1. Reject authenticated requests unless the endpoint uses HTTPS.
  2. Validate the destination against an explicit trusted-host policy before attaching an Authorization header.
  3. Disable redirects for authenticated requests or permit them only when the redirect preserves HTTPS and the trusted origin.
  4. Never forward credentials to a different hostname, port, or scheme after a redirect.
  5. Prefer protected environment variables, restricted credential files, or a secret manager over command-line arguments, which may be exposed through shell history or process listings.
  6. Display the credential destination and require explicit confirmation when authentication is sent to a previously untrusted host.
  7. Document token scoping and recommend short-lived, least-privileged credentials.
  8. Ensure errors, diagnostics, and saved output never include the credential value.
  9. Correct the Basic authentication implementation separately by either safely passing an HTTPBasicAuth object over HTTPS or rejecting unsupported Basic authentication.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (21)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The documentation overstates implemented behavior and omits important capabilities like local file input/output and a test/self-check command. This kind of description-behavior mismatch can mislead users about what the skill actually does, reducing informed consent and increasing the chance that it is run with unsafe assumptions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file includes example commands that write scraped data to local files and send authenticated HTTP requests using a bearer token, but the description does not explicitly warn users about handling credentials safely or that collected data will be stored/exported. Although the legality section discusses compliance and privacy at a high level, it does not disclose these concrete operational effects near the usage guidance.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill advertises and demonstrates network access plus local file read/write behavior, but it does not declare any explicit tool scope such as permissions or allowed-tools. This weakens least-privilege controls and makes it harder for users or the platform to understand and constrain what the skill can access before execution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill encourages sending target URLs and API bearer tokens to external services but does not clearly warn users that sensitive endpoints and credentials will be transmitted over the network. This can lead to inadvertent exposure of secrets or sensitive target information, especially if users paste production tokens or scrape internal resources.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 37)May include surrounding context.

uv run scripts/data-scraper.py scrape --urls-file urls.txt --output data.json

从 API 获取数据

uv run scripts/data-scraper.py api --endpoint "https://api.example.com/data" --auth "Bearer TOKEN"

text

### 高级选项

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 34)May include surrounding context.

uv run scripts/data-scraper.py scrape --urls-file urls.txt --output data.json

从 API 获取数据

uv run scripts/data-scraper.py api --endpoint "https://api.example.com/data" --auth "Bearer TOKEN"

text

### 高级选项

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The manifest advertises web/API scraping and bulk data collection but provides no warning, scope limitation, or privacy/compliance guidance. In a scraping skill, that omission increases the risk of misuse against third-party services, collection of personal data, or operational impact on target systems because users are not informed about consent, rate limits, or acceptable-use boundaries.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module title and description are presented in Chinese, and the CLI help and user-facing messages throughout the file also assume Chinese output. This forces a specific language for users without any opt-in or explanation that the tool is intentionally region- or locale-specific.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The inline comment says the code is handling authentication, and the branch for user:pass parsing records result["auth"] = "basic", implying Basic Auth support. However, the parsed username/password are never used in requests.get, so the documented behavior contradicts the actual implementation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code accepts --auth credentials, places bearer tokens into request headers, and then sends them in an outbound HTTP request. Although the code has comments and status output, there is no explicit user-facing warning or disclosure that credentials provided on the command line will be transmitted to the specified endpoint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill documentation is entirely presented in Chinese, including headings, instructions, and safety guidance, with no indication that users may choose another language or that the skill is intended only for a Chinese-speaking audience. Per the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The description and instructional content prominently force Chinese-language usage, while the policy requires avoiding fixed language or locale constraints unless users can opt in or the restriction is justified. No language-selection option or reason for limiting the skill to Chinese is provided.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The display name, description, and feature text are presented in Chinese, but the manifest does not indicate that language selection is optional or that the skill is intended only for a Chinese-speaking audience. This can violate language/locale policy expectations when a skill implicitly fixes one language without user opt-in or documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language comments in this file are written only in Chinese, which can impose a language choice on maintainers or users without offering an alternative or documenting a locale-specific reason. The policy explicitly calls out language or locale constraints that are forced without opt-in.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
92% confidence
Finding

Using a lower-bound specifier for requests allows builds to resolve to different versions over time, including versions with known security defects or breaking behavior. In a data-scraping skill that makes outbound HTTP requests, dependency drift increases supply-chain risk and can expose the agent to request-handling vulnerabilities.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
# Data Scraper 依赖

# HTTP 请求
requests>=2.28.0

# HTML 解析
beautifulsoup4>=4.11.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
94% confidence
Finding

requests has multiple published advisories, and because the manifest does not pin a version, it is impossible to verify whether deployment will use a patched release. This is more relevant in a scraper because the library is central to processing untrusted remote URLs and HTTP interactions.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
84% confidence
Finding

An unpinned beautifulsoup4 dependency can resolve to different releases across environments, reducing reproducibility and increasing supply-chain uncertainty. While the direct security impact is usually lower than network-facing libraries, unexpected parser behavior or compromised upstream packages could still affect scraping output or reliability.

Content

Scanner excerpt · requirements.txt (reported line 7)May include surrounding context.

text
requests>=2.28.0

# HTML 解析
beautifulsoup4>=4.11.0

# Excel 输出(可选)
openpyxl>=3.0.0

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
90% confidence
Finding

openpyxl is unpinned despite a history of XML-related advisories, so installations may pull versions with security-relevant parsing behavior. Because this skill may export or process spreadsheet data, an unsafe or inconsistent library version could expose consumers to document-processing risks.

Content

Scanner excerpt · requirements.txt (reported line 10)May include surrounding context.

text
beautifulsoup4>=4.11.0

# Excel 输出(可选)
openpyxl>=3.0.0

# 数据处理(可选)
pandas>=1.5.0

Unverifiable Dependency: openpyxl has 2 known advisory(ies) (CVE-2017-5992 (Improper Restriction of XML External Entity Reference in Openpyxl); CVE-2017-5992 (Openpyxl 2.4.1 resolves external entities by default, which allows remote attack)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

openpyxl has known historical advisories related to XML entity handling, and the current requirement does not prove that a safe version will be installed. If the broader skill handles spreadsheet files from external sources, this uncertainty can translate into document-parsing risk.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
86% confidence
Finding

An unpinned pandas dependency creates non-reproducible builds and uncertainty about whether installed versions contain bug or security fixes. In a data-processing tool this can lead to inconsistent behavior and potential exposure if unsafe deserialization or parser paths are used elsewhere in the codebase.

Content

Scanner excerpt · requirements.txt (reported line 13)May include surrounding context.

text
openpyxl>=3.0.0

# 数据处理(可选)
pandas>=1.5.0

# JavaScript 渲染支持(可选,高级功能)
# playwright>=1.30.0

Unverifiable Dependency: pandas has 1 known advisory(ies) (CVE-2020-13091 (** DISPUTED ** pandas through 1.0.3 can unserialize and execute commands from an)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
78% confidence
Finding

pandas has at least one reported advisory, and without version pinning there is no assurance that installations avoid affected releases. The direct severity here is limited by the lack of code context showing unsafe deserialization, but dependency uncertainty remains a supply-chain weakness.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.