T09 · Insecure Skill Coding Practices
- Location
scripts/scraper.py:144- Finding
Chromium Sandbox Disabled While Processing Remote Web Content
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This appears to be a real Douyin scraping skill, but it weakens browser isolation and uses mutable dependency/browser downloads, so users should review it before installing.
Install only if you are comfortable running browser automation against Douyin from an isolated, non-root environment. Prefer pinning dependencies, avoiding third-party browser mirrors unless you trust them, removing sandbox-disabling flags where possible, and confirming each scraping request before execution.
scripts/scraper.py:144Chromium Sandbox Disabled While Processing Remote Web Content
requirements.txt:1Mutable Dependencies and Third-Party Browser Binary Mirrors
The declared purpose says the skill scrapes Douyin content and hot榜 data via Playwright automation. However, this code chunk does not implement scraping, searching, hot榜 retrieval, video extraction, or any Douyin interaction at all. Its sole purpose is provisioning the runtime environment for Playwright, either in Docker or a native venv. While setup can be a supporting detail in a larger project, this specific code chunk's actual behavior is materially different from the declared primary purpose, so it should be flagged as a mismatch.
Referenced artifact was not completely inspected
node scripts/douyin_scraper.js search "搜索一下海鲜视频" 10
Referenced artifact was not completely inspected
node scripts/douyin_scraper.js search "搜索一下海鲜视频" 10
Referenced artifact was not completely inspected
node scripts/douyin_scraper.js search "搜索一下海鲜视频" 10
Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.
def mode_native() -> None:
env = os.environ.copy()
env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright")
env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing")
venv_python = Path("venv/bin/python")
The README forces a specific language/locale for all user-facing instructions and examples, with no indication that users can choose another language. Under the stated policy, this is a natural-language locale violation unless the locale restriction is explicitly optional or justified.
npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.
npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.
The skill declares no explicit tool restrictions even though its documented workflow requires shell execution, network access, environment access, and file output. Without scoped permissions, an agent may invoke the skill with broader-than-necessary capabilities, increasing the blast radius if the skill is misused or later modified.
The skill advertises activation from very broad natural-language requests such as generic search phrasing, which can cause unintended invocation. In an agent setting, ambiguous triggers can lead to unnecessary browsing, shell execution, or scraping actions when the user did not clearly request this specific skill.
Using npx playwright without a pinned version introduces supply-chain risk because the executed package version can change over time. An agent or user following these instructions may install and run unexpected code from a newer or compromised upstream release.
The agent integration guidance tells the agent to execute the scraper whenever a user makes a natural-language search request, but it does not require explicit user consent, safety checks, or boundary conditions. That increases the chance of accidental invocation and unreviewed execution of shell and network actions.
The example triggers are broad, natural-language phrases like '搜索一下海鲜视频' and '抖音热榜有什么' with no explicit scoping, confirmation, or activation boundaries. In an agent environment, this can cause the skill to be invoked unintentionally from ordinary conversation, leading to unexpected browser automation, scraping, or access to external content without clear user intent verification.
All user-facing messages and usage guidance in the script are hard-coded in Chinese, with no option for the user to select another language and no documented reason for a Chinese-only interface. This is a natural-language policy concern because the skill effectively enforces one language by default.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def run(cmd: list[str], cwd: str | None = None, env: dict[str, str] | None = None) -> None:
print("\n>>>", " ".join(cmd))
subprocess.run(cmd, cwd=cwd, env=env, check=True)
def mode_official() -> None:
Natural-language strings in comments, console output, and CLI help consistently require Chinese comprehension, but the script does not offer any language selection or indicate that it is intentionally limited to a Chinese-speaking audience. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.
The module docstring and all user-facing CLI/help strings are written only in Chinese, which imposes a specific language on users. The file does not provide any opt-in, alternative locale, or justification that this skill is intentionally restricted to a Chinese-speaking or region-specific audience.
All example utterances are Chinese-only, and the file does not state that the skill is region-specific or that users can choose their language. If these examples define expected interaction patterns, they may imply a fixed language requirement without user opt-in.
The script creates a virtual environment, writes a new run.sh file, and changes executable permissions on multiple files. These are safety-relevant filesystem modifications, and while status messages describe steps as they happen, there is no upfront disclosure that running the installer will alter files in the working directory.
This shell script upgrades pip, installs Python packages, installs a Playwright browser, and may run npm install, all of which fetch and execute external package installation logic. Although the script prints progress messages, it does not disclose that it will download third-party code from the network or modify the local environment before doing so.
This code file performs a filesystem write by creating or overwriting Dockerfile.generated. Although commands are printed before execution, there is no user-facing disclosure specifically indicating that the script will modify files, and the file itself has no docstring or comment warning about that behavior.
The dependency is specified with a lower-bound only (playwright>=1.40.0), which allows installation of any newer release, including future versions with breaking changes or newly introduced supply-chain risk. While this is not an immediate exploit by itself, it reduces build reproducibility and can unexpectedly pull vulnerable or incompatible versions into the environment.
playwright>=1.40.0
This code performs a filesystem write using a user-supplied path, which can modify local data. Although it logs after saving, there is no prior warning, confirmation, or comment/docstring near the operation to disclose that the command will write to disk and potentially overwrite an existing file.
The help text states a format option of "json (默认), csv", implying CSV export is supported. However, when format === 'csv', the code explicitly notes CSV is not implemented and still calls saveToJson, so the documentation actively misrepresents the behavior.
No suspicious patterns detected.