Back to skill

Security audit

抖音爆款爬虫 v3

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a Douyin scraping helper, but it asks agents to turn user text into shell commands and recommends logged-in browser or cookie-based scraping, which needs careful review before use.

Install only if you are comfortable running a browser automation scraper with network access and local file writes. Do not pass cookies or use a logged-in Douyin session unless you understand the account and privacy risks, and invoke the scraper with structured arguments rather than letting an agent build shell strings from free-form user text. Treat output as potentially mock or incomplete unless you verify it came from live page extraction.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:14
Finding

Shell Command Injection Through Agent-Constructed Search Commands

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
install_playwright_docker.py:29
Finding

Unpinned Dependencies and Unverified Browser Downloads From Third-Party Mirrors

Content
View full analysis
None: env = os.environ.copy() env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright") env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing") venv_python = Path("venv/bin/python") venv_pip = Path("venv/bin/pip") if not venv_python.exists(): run([sys.executable, "-m", "venv", "venv"]) run([str(venv_pip), "install", "-r", "requirements.txt"], env=env) run([str(venv_python), "-m", "playwright", "install", "chromium"], env=env) ``` `requirements.txt:1`: ```text playwright>=1.40.0 ``` `install.sh:42-49`: ```bash source venv/bin/activate pip install --upgrade pip pip install playwright # Install browser echo "" echo "🌐 Installing Playwright browser..." playwright install chromium ``` The installer also invokes `npm install` when npm is available, even though the audited project structure contains no reviewed `package.json` or lockfile: ```bash if command -v npm &> /dev/null; then echo "" echo "📦 Installing Node.js dependencies..." npm install fi ``` ### Technical Analysis The Python dependency declaration uses the lower-bound constraint `playwright>=1.40.0`, while `install.sh` installs the latest available Playwright release without any version constraint. Neither installation path uses package hashes or a fully locked dependency graph. Consequently, two installations performed at different times can retrieve different code despite using the same audited project. This weakens reproducibility and makes the effective installed code depend on the curre ...[truncated 2530 chars]
Remediation
View remediation
``` 7. Remove `npm install` from `install.sh` unless the project includes a reviewed `package.json` and lockfile. 8. If the JavaScript implementation is supported, commit an exact lockfile and use a deterministic command such as `npm ci` with lifecycle-script policy reviewed. 9. Perform installation and browser execution as an unprivileged user in a restricted container or sandbox. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (20)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The documentation claims real Douyin scraping and extraction, but also states the script may return simulated or hardcoded data and supports export behavior not reflected in the manifest. This mismatch is dangerous because agents and users may trust outputs as real data, make decisions on fabricated results, or grant broader access based on inaccurate capability descriptions.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · install_playwright_docker.py (reported line 30)May include surrounding context.

python
def mode_native() -> None:
    env = os.environ.copy()
    env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright")
    env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing")
    venv_python = Path("venv/bin/python")

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
85% confidence
Finding

The skill instructs the agent to execute shell commands and implies file read/write and environment use, but it declares no explicit tool scope or permissions. That increases the chance of unintended command execution or broader-than-expected access because the agent cannot constrain itself to the minimum required capabilities.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language activation rules are broad and map common phrases directly into shell command execution with extracted user text as parameters. Overly broad triggers increase the risk of accidental activation, unintended scraping actions, and unsafe command construction paths if later implementations handle quoting or parsing incorrectly.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The example trigger phrases overlap with ordinary conversation, making unintentional invocation more likely in unrelated contexts. In an agent setting, that can cause the skill to launch browser or shell actions without sufficiently explicit user consent or awareness.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The skill says single-video parsing is unsupported, yet later documents authenticated cookie-based access and browser-assisted flows that suggest broader scraping capability than initially stated. This inconsistency can conceal expanded authenticated access paths and makes it harder to assess what data the skill may access or automate.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The documentation explicitly instructs use of logged-in browser sessions and imported cookies to access Douyin, even though the stated purpose is scraping public search and trending data. Handling authenticated sessions and cookies introduces account/session theft, unauthorized access, privacy exposure, and terms-of-service evasion risks that are materially more sensitive than public scraping.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This shell script presents user-facing installation status, errors, and usage instructions entirely in Chinese. The file does not offer any language opt-in or indicate that the skill is intentionally limited to a Chinese-speaking or region-specific audience, which can violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · install_playwright_docker.py (reported line 14)May include surrounding context.

python
def run(cmd: list[str], cwd: str | None = None, env: dict[str, str] | None = None) -> None:
    print("\n>>>", " ".join(cmd))
    subprocess.run(cmd, cwd=cwd, env=env, check=True)


def mode_official() -> None:

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The function is presented as performing real Douyin scraping, but it returns fabricated sample records derived from user input instead of page-extracted data. This is dangerous because downstream users or agents may trust the output as authentic market or content intelligence, leading to silent integrity failures, bad decisions, or misuse in automated pipelines.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The hot-list path similarly claims to fetch trending Douyin data but produces invented records rather than scraped content from the site. In a data-collection skill, this undermines output integrity and can poison analytics, reports, or decisions that assume the data reflects the actual platform state.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring is written as a Chinese-only description, and the rest of the user-facing CLI/help strings in the file are also Chinese-only. This creates a language/locale restriction without any user opt-in or explanation that the skill is intentionally region- or language-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The entire skill documentation is presented only in Chinese, with no indication that users may choose another language or that the locale restriction is required for a region-specific purpose. The policy for this audit flags language or locale constraints when they are imposed without user opt-in or explicit justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

All example invocations are written only in Chinese, which indicates a language-specific interaction pattern without offering users any language or locale choice. The file provides no justification that the skill is region-specific or limited to Chinese-speaking contexts.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

This code creates or overwrites a local file named Dockerfile.generated, which is a file-modifying operation. Although the script prints subprocess commands, there is no comparable user-facing disclosure, comment, or docstring warning that it will write a file into the working directory.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

These commands install Python packages and download Chromium using network resources, potentially transmitting environment-derived configuration such as custom download hosts. The script echoes commands, but it does not explicitly warn users in comments, docstrings, or messaging that it will fetch and install external software over the network.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
94% confidence
Finding

The dependency is specified with a lower-bound only (playwright>=1.40.0), which allows future major or minor releases to be installed without review. This can introduce breaking changes or, more importantly, pull in a compromised or vulnerable upstream version through the supply chain, reducing build reproducibility and making security posture harder to control.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
playwright>=1.40.0

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The manifest describes searching, fetching hot lists, and extracting video/caption data, but this file also implements persistent export of results to JSON/CSV files via helper functions and CLI options. Local file export is adjacent to scraping, but it is still additional behavior not mentioned in the stated description.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.