T09 · Insecure Skill Coding Practices
- Location
scraper.js:25- Finding
Chromium Sandbox Disabled While Rendering Untrusted Websites
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a purpose-aligned web scraper, but it needs Review because it drives an unsandboxed browser to arbitrary user-supplied URLs without network or local-resource limits.
Install only if you will run it in an isolated environment, use it only on sites you are authorized to scrape, avoid internal or local URLs, and review or pin dependencies first. Treat scraped exports as potentially sensitive data.
scraper.js:25Chromium Sandbox Disabled While Rendering Untrusted Websites
] [--format json|csv|excel] [--output <file>]');
process.exit(1);
}
(async () => {
const scraper = new SmartScraper();
let config = { selector: 'body', fields: [{ name: 'content', selector: '*', type: 'text' }] };
if (configIdx !== -1) {
const configFile = args[configIdx + 1];
if (fs.existsSync(configFile)) {
config = JSON.parse(fs.readFileSync(configFile));
}
}
const url = urlIdx !== -1 ? args[urlIdx + 1] : config.target?.url;
const format = formatIdx !== -1 ? args
...[truncated 2690 chars]:34Unrestricted Browser Navigation Permits Access to Local and Internal Resources
The code does implement core web scraping and export to JSON/CSV/Excel, so it is related to the declared purpose. However, several prominent claims in the description overstate the actual behavior. The scraper does not automatically infer page structure; instead it requires a configuration file with selectors and field definitions. Although a scrapeMultiple method exists, the provided executable flow only processes a single URL, so batch scraping is not actually exposed as primary behavior. Likewise, anti-bot support is minimal and not meaningfully adaptive. Therefore the description does not accurately represent the code's actual capabilities.
The skill explicitly promotes anti-detection behavior such as User-Agent rotation, randomized delays, and optional proxy pools, but does not clearly warn users that it performs active network interactions designed to evade site defenses. In context, this makes the skill more dangerous because it facilitates stealthy scraping behavior that can violate terms of service, trigger abuse complaints, or be repurposed for large-scale unauthorized collection.
The skill advertises data export to JSON/CSV/Excel and direct database storage without clearly warning that scraped data may contain personal, confidential, or third-party proprietary information. This can lead users to collect and persist sensitive data unsafely, increasing legal, privacy, and downstream security exposure if exported files or databases are mishandled.
Multiple visible strings in the script, including the title and CLI/log messages, are presented only in Chinese. This imposes a language choice on users without offering localization, configuration, or opt-in, which matches the language/locale policy violation criteria.
This code accepts a URL and performs outbound browser navigation to that address, which can transmit the user's target selection, IP, headers, and browser fingerprinting data. Although it logs that scraping is starting, there is no user-facing warning or comment disclosing that the tool will make live network requests to third-party sites and collect page content.
The skill name and description are presented in Chinese, and the document does not indicate that other languages are supported or that Chinese is required for a region-specific purpose. Per the policy, a skill should not impose a specific language or locale without user opt-in or a clearly documented justification.
The package description is written entirely in Chinese, which indicates a language-specific user-facing presentation without any stated opt-in or explanation that the skill is intended only for a Chinese-speaking context. The policy requires avoiding forced language/locale constraints unless the skill offers a choice or clearly documents the justification.
Using a caret range for puppeteer allows different minor/patch releases to be installed over time, which can make builds non-reproducible and may introduce vulnerable or breaking versions through the software supply chain. In a web-scraping skill that drives a browser engine, dependency drift increases exposure because the package is security-sensitive and interacts with untrusted web content.
"keywords": ["scraper", "crawler", "data", "automation"],
"license": "MIT",
"dependencies": {
"puppeteer": "^21.0.0",
"cheerio": "^1.0.0-rc.12",
"user-agents": "^2.2.1",
"json2csv": "^6.0.0-alpha.2",
Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.
Using an unpinned version range for user-agents makes installations non-deterministic and can silently pull in changed code from the registry. While this package is less inherently sensitive than a browser automation library, it still creates avoidable supply-chain risk.
"dependencies": {
"puppeteer": "^21.0.0",
"cheerio": "^1.0.0-rc.12",
"user-agents": "^2.2.1",
"json2csv": "^6.0.0-alpha.2",
"xlsx": "^0.18.5"
}
Using a caret range for xlsx can cause unreviewed versions to be installed, which is especially risky for a library that parses complex file formats and has a history of security advisories. In a scraper/export tool, spreadsheet generation and processing increase the relevance of this dependency to potential attack paths.
"cheerio": "^1.0.0-rc.12",
"user-agents": "^2.2.1",
"json2csv": "^6.0.0-alpha.2",
"xlsx": "^0.18.5"
}
}
The manifest uses an unpinned xlsx version and the dependency has multiple known advisories, including prototype pollution and ReDoS classes that are relevant to data-processing libraries. Because the exact resolved version is not shown, the presence of a vulnerable release cannot be proven from this file alone, but in context this materially raises supply-chain risk beyond a purely theoretical concern.
No suspicious patterns detected.