Back to skill

Security audit

E-commerce Data Scraper Pro

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent data-scraping skill, but users should review it because API tokens can be sent to any user-provided endpoint without built-in safeguards.

Install only if you are comfortable with a basic, partially implemented scraper. Do not pass valuable API tokens unless you fully trust the exact HTTPS endpoint, and avoid putting secrets directly on the command line. Treat local output files as potentially sensitive and check the target site's terms, robots.txt, and privacy rules before scraping.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/data-scraper.py:134
Finding

Bearer Credentials Can Be Sent to Arbitrary or Insecure Endpoints

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

整体方向与声明相近:代码的主要目的确实是网页/API 数据抓取,并支持批量处理与文件输出,没有发现与声明无关的高风险隐藏能力。但存在明显的描述-行为不一致:1)scrape_url 虽加载 BeautifulSoup 和预定义选择器模板,却没有执行任何实际字段抽取,result['data'] 始终为空,仅返回“完整功能需要安装依赖后实现”的说明,因此“从网页提取结构化数据”这一核心能力并未真正实现;2)输出格式参数宣称支持 json/csv/excel,但 save_output 仅真正实现 JSON,CSV 只是写入占位注释加 JSON,Excel 完全未实现而退回 JSON。相比之下,API 获取和批量处理能力基本存在。因此应判定为部分但实质性的能力夸大,属于 mismatch。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The markdown includes an example that passes an authorization token on the command line and writes fetched data to a local file, but it does not warn users that shell arguments may be exposed in shell history or process listings, nor that scraped/API data will be stored locally. The existing compliance section discusses legal scraping behavior, but not privacy or system-safety implications of credential use and saved outputs.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill advertises and instructs use of capabilities that imply network access plus local file read/write, but it declares no explicit tool scope or permissions. In an agent environment, missing scope boundaries increases the risk of over-privileged execution, unintended file access, or unrestricted outbound requests beyond what users expect.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest description and top-level documentation are written in Chinese, and the skill does not indicate that other languages are supported or that the user can choose a preferred locale. This creates a natural-language policy concern because the skill appears to impose a specific language without opt-in or justification.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 37)May include surrounding context.

uv run scripts/data-scraper.py scrape --urls-file urls.txt --output data.json

从 API 获取数据

uv run scripts/data-scraper.py api --endpoint "https://api.example.com/data" --auth "Bearer TOKEN"

text

### 高级选项

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 34)May include surrounding context.

uv run scripts/data-scraper.py scrape --urls-file urls.txt --output data.json

从 API 获取数据

uv run scripts/data-scraper.py api --endpoint "https://api.example.com/data" --auth "Bearer TOKEN"

text

### 高级选项

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest explicitly advertises web/API extraction, batch processing, and data collection capabilities but provides no privacy, authorization, or data-handling notice. In a scraping tool, that omission increases the risk that users deploy it against personal, restricted, or rate-limited sources without clear safeguards or expectations, which can lead to privacy breaches or policy violations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code file contains natural-language strings that define the tool and its interface entirely in Chinese, including the module docstring and later CLI help text. That effectively forces a single language for users without offering a choice or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code comment at L140 says it is processing authentication, and the branch at L145-L148 parses user:pass credentials for basic auth, but the resulting credentials are never attached to the requests.get call at L150. This contradicts the apparent intent of the code/documentation because Bearer auth is applied while basic auth is only recorded in the result metadata.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill documentation appears to force a specific language/locale for all user-facing instructions and examples, with no opt-in or alternative language guidance. Under the stated policy, a language restriction should be optional or explicitly justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The display name and description are presented in Chinese, but the manifest does not indicate that language is selectable or that the skill is intended only for a Chinese-speaking locale. This can violate language/locale policy when users are not given an opt-in or documented locale constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This requirements file includes human-readable comments entirely in Chinese (for example, dependency descriptions on L01, L03, L06, L09, L12, and L15). Under the policy, forcing a specific language without user opt-in or documented justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
89% confidence
Finding

Using an unpinned dependency range for requests allows future installs to resolve to different versions over time, including versions with newly introduced vulnerabilities or breaking security behavior. In a data-scraping skill that makes outbound HTTP requests, dependency drift increases supply-chain risk and can expose network-facing functionality to known flaws.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
# Data Scraper 依赖

# HTTP 请求
requests>=2.28.0

# HTML 解析
beautifulsoup4>=4.11.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

Requests has multiple known advisories, and because the manifest does not pin a version, there is no way to verify whether deployments avoid affected releases. Given this skill performs web/API access, the uncertainty is more dangerous than in an offline-only tool because the dependency sits directly on a network-facing path.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
87% confidence
Finding

An unpinned beautifulsoup4 dependency creates non-reproducible builds and supply-chain uncertainty, even if the package is not highly security-sensitive by itself. Different resolved versions may change parser behavior or include security defects not accounted for during testing.

Content

Scanner excerpt · requirements.txt (reported line 7)May include surrounding context.

text
requests>=2.28.0

# HTML 解析
beautifulsoup4>=4.11.0

# Excel 输出(可选)
openpyxl>=3.0.0

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
91% confidence
Finding

Leaving openpyxl unpinned can allow installation of versions affected by XML-related issues, including historical XXE-style problems. In a scraping/export tool that may process untrusted spreadsheet content or generate Excel outputs in shared environments, uncontrolled version selection raises the chance of deploying a vulnerable release.

Content

Scanner excerpt · requirements.txt (reported line 10)May include surrounding context.

text
beautifulsoup4>=4.11.0

# Excel 输出(可选)
openpyxl>=3.0.0

# 数据处理(可选)
pandas>=1.5.0

Unverifiable Dependency: openpyxl has 2 known advisory(ies) (CVE-2017-5992 (Improper Restriction of XML External Entity Reference in Openpyxl); CVE-2017-5992 (Openpyxl 2.4.1 resolves external entities by default, which allows remote attack)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

Openpyxl has known historical advisories, and the current manifest does not provide enough version precision to determine whether an affected release could be installed. Because spreadsheet parsers may process attacker-controlled files or XML content, version ambiguity can translate into real risk if unsafe releases are resolved.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
84% confidence
Finding

An unpinned pandas dependency introduces supply-chain and reproducibility risk by allowing installs to vary across environments. Although the direct exploitability depends on how pandas is used, unresolved version drift can expose the skill to known issues in data parsing or deserialization-related code paths.

Content

Scanner excerpt · requirements.txt (reported line 13)May include surrounding context.

text
openpyxl>=3.0.0

# 数据处理(可选)
pandas>=1.5.0

# JavaScript 渲染支持(可选,高级功能)
# playwright>=1.30.0

Unverifiable Dependency: pandas has 1 known advisory(ies) (CVE-2020-13091 (** DISPUTED ** pandas through 1.0.3 can unserialize and execute commands from an)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
78% confidence
Finding

Pandas has at least one cited advisory, and without a pinned version the installed release cannot be verified as safe. The practical danger depends on whether the skill uses risky deserialization or parsing features, but the manifest still leaves avoidable uncertainty in a data-processing component.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.