Back to skill

Security audit

智能网页爬虫

Security checks for vulnerabilities and agentic risk

Overview

This is a purpose-aligned web scraper, but it needs Review because it drives an unsandboxed browser to arbitrary user-supplied URLs without network or local-resource limits.

Install only if you will run it in an isolated environment, use it only on sites you are authorized to scrape, avoid internal or local URLs, and review or pin dependencies first. Treat scraped exports as potentially sensitive data.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scraper.js:25
Finding

Chromium Sandbox Disabled While Rendering Untrusted Websites

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
] [--format json|csv|excel] [--output <file>]'); process.exit(1); } (async () => { const scraper = new SmartScraper(); let config = { selector: 'body', fields: [{ name: 'content', selector: '*', type: 'text' }] }; if (configIdx !== -1) { const configFile = args[configIdx + 1]; if (fs.existsSync(configFile)) { config = JSON.parse(fs.readFileSync(configFile)); } } const url = urlIdx !== -1 ? args[urlIdx + 1] : config.target?.url; const format = formatIdx !== -1 ? args ...[truncated 2690 chars]:34
Finding

Unrestricted Browser Navigation Permits Access to Local and Internal Resources

Content
View full analysis
[--config ] [--format json|csv|excel] [--output ]'); process.exit(1); } (async () => { const scraper = new SmartScraper(); let config = { selector: 'body', fields: [{ name: 'content', selector: '*', type: 'text' }] }; if (configIdx !== -1) { const configFile = args[configIdx + 1]; if (fs.existsSync(configFile)) { config = JSON.parse(fs.readFileSync(configFile)); } } const url = urlIdx !== -1 ? args[urlIdx + 1] : config.target?.url; const format = formatIdx !== -1 ? args[formatIdx + 1] : 'json'; const output = outputIdx !== -1 ? args[outputIdx + 1] : 'output.' + format; const data = await scraper.scrape(url, config); scraper.export(data, format, output); })(); ``` ### Technical Analysis The destination supplied through `--url` or `config.target.url` reaches `page.goto()` without validation. The implementation does not restrict navigation to `http:` and `https:`, reject local-resource schemes, block loopback or private address ranges, or validate redirect destinations. Consequently, a caller who can influence the scraping request can direct the browser toward resources reachable from the scraper host rather than only public websites. Potential targets include loo ...[truncated 1863 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
82% confidence
Finding

The code does implement core web scraping and export to JSON/CSV/Excel, so it is related to the declared purpose. However, several prominent claims in the description overstate the actual behavior. The scraper does not automatically infer page structure; instead it requires a configuration file with selectors and field definitions. Although a scrapeMultiple method exists, the provided executable flow only processes a single URL, so batch scraping is not actually exposed as primary behavior. Likewise, anti-bot support is minimal and not meaningfully adaptive. Therefore the description does not accurately represent the code's actual capabilities.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill explicitly promotes anti-detection behavior such as User-Agent rotation, randomized delays, and optional proxy pools, but does not clearly warn users that it performs active network interactions designed to evade site defenses. In context, this makes the skill more dangerous because it facilitates stealthy scraping behavior that can violate terms of service, trigger abuse complaints, or be repurposed for large-scale unauthorized collection.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill advertises data export to JSON/CSV/Excel and direct database storage without clearly warning that scraped data may contain personal, confidential, or third-party proprietary information. This can lead users to collect and persist sensitive data unsafely, increasing legal, privacy, and downstream security exposure if exported files or databases are mishandled.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Multiple visible strings in the script, including the title and CLI/log messages, are presented only in Chinese. This imposes a language choice on users without offering localization, configuration, or opt-in, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code accepts a URL and performs outbound browser navigation to that address, which can transmit the user's target selection, IP, headers, and browser fingerprinting data. Although it logs that scraping is starting, there is no user-facing warning or comment disclosing that the tool will make live network requests to third-party sites and collect page content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The skill name and description are presented in Chinese, and the document does not indicate that other languages are supported or that Chinese is required for a region-specific purpose. Per the policy, a skill should not impose a specific language or locale without user opt-in or a clearly documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The package description is written entirely in Chinese, which indicates a language-specific user-facing presentation without any stated opt-in or explanation that the skill is intended only for a Chinese-speaking context. The policy requires avoiding forced language/locale constraints unless the skill offers a choice or clearly documents the justification.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
92% confidence
Finding

Using a caret range for puppeteer allows different minor/patch releases to be installed over time, which can make builds non-reproducible and may introduce vulnerable or breaking versions through the software supply chain. In a web-scraping skill that drives a browser engine, dependency drift increases exposure because the package is security-sensitive and interacts with untrusted web content.

Content

Scanner excerpt · package.json (reported line 13)May include surrounding context.

json
"keywords": ["scraper", "crawler", "data", "automation"],
  "license": "MIT",
  "dependencies": {
    "puppeteer": "^21.0.0",
    "cheerio": "^1.0.0-rc.12",
    "user-agents": "^2.2.1",
    "json2csv": "^6.0.0-alpha.2",

Unverifiable Dependency: puppeteer has 1 known advisory(ies) (CVE-2019-5786 (Use-After-Free in puppeteer)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
88% confidence
Finding

Using an unpinned version range for user-agents makes installations non-deterministic and can silently pull in changed code from the registry. While this package is less inherently sensitive than a browser automation library, it still creates avoidable supply-chain risk.

Content

Scanner excerpt · package.json (reported line 15)May include surrounding context.

json
"dependencies": {
    "puppeteer": "^21.0.0",
    "cheerio": "^1.0.0-rc.12",
    "user-agents": "^2.2.1",
    "json2csv": "^6.0.0-alpha.2",
    "xlsx": "^0.18.5"
  }

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

Using a caret range for xlsx can cause unreviewed versions to be installed, which is especially risky for a library that parses complex file formats and has a history of security advisories. In a scraper/export tool, spreadsheet generation and processing increase the relevance of this dependency to potential attack paths.

Content

Scanner excerpt · package.json (reported line 17)May include surrounding context.

json
"cheerio": "^1.0.0-rc.12",
    "user-agents": "^2.2.1",
    "json2csv": "^6.0.0-alpha.2",
    "xlsx": "^0.18.5"
  }
}

Unverifiable Dependency: xlsx has 5 known advisory(ies) (CVE-2021-32012 (Denial of Service in SheetJS Pro); CVE-2023-30533 (Prototype Pollution in sheetJS); CVE-2024-22363 (SheetJS Regular Expression Denial of Service (ReDoS)) +2 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
81% confidence
Finding

The manifest uses an unpinned xlsx version and the dependency has multiple known advisories, including prototype pollution and ReDoS classes that are relevant to data-processing libraries. Because the exact resolved version is not shown, the presence of a vulnerable release cannot be proven from this file alone, but in context this materially raises supply-chain risk beyond a purely theoretical concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.