T09 · Insecure Skill Coding Practices
- Location
scripts/extract_data.js:11- Finding
Chromium Sandbox Disabled While Processing Configurable Web Content
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a web-extraction helper, but it needs Review because it runs broad browser automation with weak containment and under-discloses setup, upload, and authenticated-page risks.
Install only in an isolated, unprivileged environment. Avoid using it on logged-in or sensitive pages unless you have explicitly reviewed what will be captured, where outputs are written, and whether any upload path is enabled. Pin dependencies, remove --no-sandbox where possible, and do not add the cron job or run sudo setup commands casually.
scripts/extract_data.js:11Chromium Sandbox Disabled While Processing Configurable Web Content
SKILL.md:34Unpinned Third-Party Dependencies and Privileged Installation Instructions
The repeated description-behavior mismatch suggests the skill materially overstates its capability and generality. In security review terms, that is dangerous because operators cannot reliably infer what permissions, network access, or data handling practices are appropriate from the documentation provided.
The repeated description-behavior mismatch suggests the skill materially overstates its capability and generality. In security review terms, that is dangerous because operators cannot reliably infer what permissions, network access, or data handling practices are appropriate from the documentation provided.
The repeated description-behavior mismatch suggests the skill materially overstates its capability and generality. In security review terms, that is dangerous because operators cannot reliably infer what permissions, network access, or data handling practices are appropriate from the documentation provided.
Referenced artifact was not completely inspected
node scripts/extract_data.js --url "https://example.com/chart" --output data.json
Referenced artifact was not completely inspected
node scripts/extract_data.js --url "https://example.com/chart" --output data.json
Referenced artifact was not completely inspected
node scripts/extract_data.js --url "https://example.com/chart" --output data.json
This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.
# Increase wait time for dynamic content
# Check selectors in extract_data.js
# Enable debug mode: export DEBUG=playwright:*
Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).
print(f"\n{'='*60}")
print(f"🚀 {description}")
print(f"{'='*60}")
result = subprocess.run(cmd, shell=True, capture_output=False)
return result.returncode == 0
def main():
The skill documents execution of shell commands, environment-variable use, and file output, but does not declare any tool scope or permission boundaries. This makes the operational surface implicit rather than reviewable, increasing the chance that an agent invokes browser automation, writes files, or uses secrets without clear user consent.
The invocation guidance is broad enough to trigger the skill for many generic web-data tasks, including authenticated pages and visual extraction, without clear boundaries. Over-broad trigger criteria increase the chance of the agent selecting this skill in contexts involving sensitive browsing sessions, credentials, or data exports that need stricter controls.
The skill explicitly contemplates authenticated scraping and exporting/uploading extracted data, but the description does not warn users about credential handling, sensitive-data capture, or exfiltration risks. In a browser automation skill, omission of those warnings materially increases the chance of unsafe use on private dashboards or accounts.
Using 'npx playwright' without a pinned version allows installation or execution of whatever version is current at runtime. That weakens reproducibility and creates supply-chain risk if a bad release or dependency compromise is pulled into the environment.
The troubleshooting step again relies on unpinned 'npx playwright', exposing the environment to non-deterministic dependency resolution during support operations. Troubleshooting paths are often copy-pasted directly, so they carry the same supply-chain exposure as primary install steps.
The documentation instructs users to run 'sudo apt-get install', normalizing privileged execution as part of skill setup. Encouraging root-level commands increases the blast radius of mistakes and supply-chain compromise during installation, especially in automation environments where commands may be run without careful review.
# Install additional language packs
sudo apt-get install tesseract-ocr-eng
# Use image pre-processing
python3 scripts/preprocess_image.py --input screenshot.png
The example configuration enables automatic Feishu upload and includes concrete third-party destination identifiers, while the skill description only frames the tool as local web extraction and OCR. This creates a data-flow mismatch that can cause users to unknowingly send scraped content to an external service, increasing exfiltration and privacy risk.
Third-party upload is not clearly necessary for Playwright/OCR extraction, so enabling it in configuration broadens the trust boundary without strong justification. In a scraping/OCR context, extracted page contents may include sensitive business, personal, or proprietary data, making silent export to Feishu materially risky.
The code hardcodes toLocaleDateString('en-US', ...), which imposes a specific locale regardless of user or environment preferences. This matches the language/locale policy concern because no opt-in, fallback, or justification is provided for forcing U.S. English output.
The script's human-facing description and console messages are written entirely in Chinese, which implicitly fixes the interaction language for users and operators. Under the policy, locale-specific behavior should offer user choice or clearly document a justified regional constraint.
The top-level comment states '补充缺失的周数据' ('fill missing weekly data'), which implies completing or repairing missing dataset entries. In practice, the script ignores the declared missingDates list, performs dense hover sampling across a chart, and saves collected tooltip text without actually identifying or filling the listed missing weeks.
The manifest describes a Playwright+OCR web extraction skill with batch processing, data validation, and error recovery, but this script only navigates to a rankings page, scrapes tooltip text via DOM inspection and mouse hover, and writes raw results to disk. It performs no OCR and does not implement the claimed batch processing, validation, or recovery behaviors, making the actual operation materially narrower and different from the stated capability.
This Python file contains docstrings and printed status messages in Chinese, which effectively forces a specific language for users and operators. The policy allows locale constraints only when they are optional or clearly justified, but this file does not provide opt-in, alternatives, or documentation explaining a Chinese-only requirement.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
print(f"\n{'='*60}")
print(f"🚀 {description}")
print(f"{'='*60}")
result = subprocess.run(cmd, shell=True, capture_output=False)
return result.returncode == 0
def main():
The script executes shell commands via subprocess.run with shell=True, which triggers browser automation and downstream processing. Although progress is printed, there is no explicit disclosure that external commands will be executed on the user's system, and this is safety-relevant for code files.
The changelog switches to Chinese for substantial user-facing documentation while the rest of the file is in English, without offering a language choice or documenting a locale-specific reason. This can violate language/locale policy where skills should not force a specific language without user opt-in.
This JSON config sets "ocr_language" to "eng", which is a natural-language locale constraint. Under the policy rule, forcing a specific language without offering a user choice or documenting justification is a language/locale policy concern.
No suspicious patterns detected.