Back to skill

Security audit

抖音搜索爬虫

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real Douyin scraping skill, but it weakens browser isolation and uses mutable dependency/browser downloads, so users should review it before installing.

Install only if you are comfortable running browser automation against Douyin from an isolated, non-root environment. Prefer pinning dependencies, avoiding third-party browser mirrors unless you trust them, removing sandbox-disabling flags where possible, and confirming each scraping request before execution.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/scraper.py:144
Finding

Chromium Sandbox Disabled While Processing Remote Web Content

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Mutable Dependencies and Third-Party Browser Binary Mirrors

Content
View full analysis
=1.40.0 ``` The main installer independently installs the latest available package without a version or integrity lock: ```bash # 激活虚拟环境并安装依赖 echo "" echo "📦 正在安装 Python 依赖..." source venv/bin/activate pip install --upgrade pip pip install playwright ``` Native installation configures third-party mirrors as default browser artifact sources: ```python def mode_native() -> None: env = os.environ.copy() env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright") env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing") venv_python = Path("venv/bin/python") venv_pip = Path("venv/bin/pip") if not venv_python.exists(): run([sys.executable, "-m", "venv", "venv"]) run([str(venv_pip), "install", "-r", "requirements.txt"], env=env) run([str(venv_python), "-m", "playwright", "install", "chromium"], env=env) ``` ### Technical Analysis The project does not provide a reproducible, integrity-locked dependency set: - `playwright>=1.40.0` allows future versions that were not present during this audit. - `install.sh` ignores `requirements.txt` and installs the latest Playwright release directly. - No hashes are supplied for the Python package or transitive dependencies. - Native browser installation defaults to third-party mirror domains for executable browser artifacts. Playwright installation downloads and later runs a complete browser executable. Consequently, integrity and provenance controls for both Python packages and browser artifacts are security-critical. HTTPS protects transport but does no ...[truncated 1757 chars]
Remediation
View remediation
`. 2. Generate and commit a lock file that includes exact transitive versions and cryptographic hashes. Install with hash verification, such as `pip install --require-hashes -r requirements.lock`. 3. Make `install.sh` consume the locked dependency file instead of running the unversioned `pip install playwright`. 4. Avoid automatically upgrading `pip` to an unreviewed latest version during routine Skill installation. Pin or separately manage installer tooling where reproducibility is required. 5. Use Playwright's official browser artifact source by default. 6. If third-party mirrors are operationally necessary: - Make their use explicit and opt-in. - Document the trust implications. - Pin the exact expected browser revision. - Verify downloaded artifacts against independently obtained, committed checksums or signatures. 7. Pin Docker images by immutable digest rather than relying only on a mutable tag. 8. Perform dependency vulnerability and provenance checks whenever pinned versions are updated. 9. Run dependency and browser installation as an unprivileged account in an isolated build environment. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (24)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose says the skill scrapes Douyin content and hot榜 data via Playwright automation. However, this code chunk does not implement scraping, searching, hot榜 retrieval, video extraction, or any Douyin interaction at all. Its sole purpose is provisioning the runtime environment for Playwright, either in Docker or a native venv. While setup can be a supporting detail in a larger project, this specific code chunk's actual behavior is materially different from the declared primary purpose, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 66)May include surrounding context.

md
node scripts/douyin_scraper.js search "搜索一下海鲜视频" 10

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 69)May include surrounding context.

md
node scripts/douyin_scraper.js search "搜索一下海鲜视频" 10

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 72)May include surrounding context.

md
node scripts/douyin_scraper.js search "搜索一下海鲜视频" 10

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · install_playwright_docker.py (reported line 30)May include surrounding context.

python
def mode_native() -> None:
    env = os.environ.copy()
    env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright")
    env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing")
    venv_python = Path("venv/bin/python")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README forces a specific language/locale for all user-facing instructions and examples, with no indication that users can choose another language. Under the stated policy, this is a natural-language locale violation unless the locale restriction is explicitly optional or justified.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
86% confidence
Finding

The skill declares no explicit tool restrictions even though its documented workflow requires shell execution, network access, environment access, and file output. Without scoped permissions, an agent may invoke the skill with broader-than-necessary capabilities, increasing the blast radius if the skill is misused or later modified.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill advertises activation from very broad natural-language requests such as generic search phrasing, which can cause unintended invocation. In an agent setting, ambiguous triggers can lead to unnecessary browsing, shell execution, or scraping actions when the user did not clearly request this specific skill.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
78% confidence
Finding

Using npx playwright without a pinned version introduces supply-chain risk because the executed package version can change over time. An agent or user following these instructions may install and run unexpected code from a newer or compromised upstream release.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The agent integration guidance tells the agent to execute the scraper whenever a user makes a natural-language search request, but it does not require explicit user consent, safety checks, or boundary conditions. That increases the chance of accidental invocation and unreviewed execution of shell and network actions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The example triggers are broad, natural-language phrases like '搜索一下海鲜视频' and '抖音热榜有什么' with no explicit scoping, confirmation, or activation boundaries. In an agent environment, this can cause the skill to be invoked unintentionally from ordinary conversation, leading to unexpected browser automation, scraping, or access to external content without clear user intent verification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

All user-facing messages and usage guidance in the script are hard-coded in Chinese, with no option for the user to select another language and no documented reason for a Chinese-only interface. This is a natural-language policy concern because the skill effectively enforces one language by default.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · install_playwright_docker.py (reported line 14)May include surrounding context.

python
def run(cmd: list[str], cwd: str | None = None, env: dict[str, str] | None = None) -> None:
    print("\n>>>", " ".join(cmd))
    subprocess.run(cmd, cwd=cwd, env=env, check=True)


def mode_official() -> None:

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Natural-language strings in comments, console output, and CLI help consistently require Chinese comprehension, but the script does not offer any language selection or indicate that it is intentionally limited to a Chinese-speaking audience. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The module docstring and all user-facing CLI/help strings are written only in Chinese, which imposes a specific language on users. The file does not provide any opt-in, alternative locale, or justification that this skill is intentionally restricted to a Chinese-speaking or region-specific audience.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

All example utterances are Chinese-only, and the file does not state that the skill is region-specific or that users can choose their language. If these examples define expected interaction patterns, they may imply a fixed language requirement without user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The script creates a virtual environment, writes a new run.sh file, and changes executable permissions on multiple files. These are safety-relevant filesystem modifications, and while status messages describe steps as they happen, there is no upfront disclosure that running the installer will alter files in the working directory.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This shell script upgrades pip, installs Python packages, installs a Playwright browser, and may run npm install, all of which fetch and execute external package installation logic. Although the script prints progress messages, it does not disclose that it will download third-party code from the network or modify the local environment before doing so.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This code file performs a filesystem write by creating or overwriting Dockerfile.generated. Although commands are printed before execution, there is no user-facing disclosure specifically indicating that the script will modify files, and the file itself has no docstring or comment warning about that behavior.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The dependency is specified with a lower-bound only (playwright>=1.40.0), which allows installation of any newer release, including future versions with breaking changes or newly introduced supply-chain risk. While this is not an immediate exploit by itself, it reduces build reproducibility and can unexpectedly pull vulnerable or incompatible versions into the environment.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
playwright>=1.40.0

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This code performs a filesystem write using a user-supplied path, which can modify local data. Although it logs after saving, there is no prior warning, confirmation, or comment/docstring near the operation to disclose that the command will write to disk and potentially overwrite an existing file.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The help text states a format option of "json (默认), csv", implying CSV export is supported. However, when format === 'csv', the code explicitly notes CSV is not implemented and still calls saveToJson, so the documentation actively misrepresents the behavior.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.