Back to skill

Security audit

抖音爆款爬虫

Security checks for vulnerabilities and agentic risk

Overview

The skill is a Douyin scraping tool, but its code mostly returns simulated data while presenting results as scraped, and its installer pulls unpinned executable browser dependencies.

Review before installing. This does not show a backdoor or credential theft, but users should treat the results as mock data unless real extraction is implemented, install only in an isolated non-privileged environment, pin and verify Playwright/Chromium dependencies, and avoid writing outputs to important existing files.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned executable dependencies and browser binaries

Content
View full analysis

Vulnerability Details

File Location: requirements.txt:1; related installation logic at install.sh:39-53
Vulnerability Type: Unpinned and non-reproducible third-party dependencies
Risk Level: Medium

Vulnerable Code

requirements.txt:1:

text
playwright>=1.40.0

install.sh:39-53:

bash
echo ""
echo "📦 正在安装 Python 依赖..."
source venv/bin/activate
pip install --upgrade pip
pip install playwright

# 安装浏览器
echo ""
echo "🌐 正在安装 Playwright 浏览器..."
playwright install chromium

# 安装 Node.js 依赖(如果有 Node.js)
if command -v npm &> /dev/null; then
    echo ""
    echo "📦 正在安装 Node.js 依赖..."
    npm install
fi

The same unpinned installation instructions also appear in SKILL.md:75-76 and README.md:30-33,43-46.

Technical Analysis

The installation process retrieves and installs Playwright without an exact version or cryptographic hash. The declared requirement only establishes a lower bound, allowing dependency resolution to select future releases that were not reviewed as part of this audit. The subsequent playwright install chromium command also downloads an executable browser build associated with the resolved Playwright package.

The installer additionally invokes npm install, although the audited project contains no package.json or reviewed lockfile. This installation path is therefore incomplete and non-reproducible. If a package manifest is introduced into the working directory before installation, npm could execute lifecycle scripts defined by that manifest.

This is a supply-chain weakness rather than evidence that the current Playwright package is malicious. Exploitation requires compromise or manipulation of a package registry, package artifact, dependency resolution process, browser download source, or local package manifest.

Attack Path

  1. An attacker compromises a dependency release, registry artifact, browser distribution cha ...[truncated 1327 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin Playwright to an exact reviewed version, for example playwright==<reviewed-version>.
  2. Generate and enforce cryptographic hashes for Python dependencies, such as with a hash-locked requirements file and pip install --require-hashes.
  3. Pin and document the expected Playwright browser revision rather than implicitly accepting whichever browser corresponds to a newly resolved package release.
  4. Commit a reviewed package.json and lockfile if the Node.js implementation is supported. Use npm ci rather than npm install to enforce the lockfile.
  5. Remove the npm installation step if no Node.js package manifest is intentionally shipped.
  6. Disable or tightly control package lifecycle scripts where operationally possible.
  7. Perform dependency installation in a non-privileged, isolated environment and scan resolved packages and browser artifacts before deployment.
  8. Add automated dependency review so version updates require explicit approval and security testing.

T08 · Insecure Dependencies

Warning
Location
install_playwright_docker.py:30
Finding

Executable Chromium is downloaded from third-party mirrors without project-level integrity verification

Content
View full analysis

Vulnerability Details

File Location: install_playwright_docker.py:30-38
Vulnerability Type: Unsafe executable download source
Risk Level: Medium

Vulnerable Code

python
def mode_native() -> None:
    env = os.environ.copy()
    env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright")
    env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing")
    venv_python = Path("venv/bin/python")
    venv_pip = Path("venv/bin/pip")
    if not venv_python.exists():
        run([sys.executable, "-m", "venv", "venv"])
    run([str(venv_pip), "install", "-r", "requirements.txt"], env=env)
    run([str(venv_python), "-m", "playwright", "install", "chromium"], env=env)

Technical Analysis

Native installation defaults Playwright and Chromium downloads to npmmirror.com and cdn.npmmirror.com, which are third-party distribution infrastructure rather than the default official Playwright download source. The Skill does not independently verify an expected checksum, signature, or immutable digest for the downloaded Chromium artifact.

HTTPS protects data in transit but does not eliminate the need to trust the mirror operator, its storage, its certificate and account security, and its upstream synchronization process. If the mirror or delivered artifact is compromised, the installed browser may differ from the artifact reviewed or expected by the project.

Because the browser is an executable component subsequently launched by the scraper, artifact substitution can cross directly from a supply-chain compromise into local code execution.

Attack Path

  1. An attacker compromises the configured mirror, its publishing account, artifact storage, or an upstream synchronization process.
  2. The attacker substitutes a malicious Chromium or Playwright browser archive for the requested artifact.
  3. A user runs `python install_ ...[truncated 1023 chars]
Remediation
View remediation

Remediation Suggestions

  1. Use Playwright's official browser distribution infrastructure unless a third-party mirror is an explicitly approved organizational requirement.
  2. Pin Playwright to an exact version and record the exact expected Chromium revision.
  3. Verify the downloaded archive against a trusted, pre-recorded SHA-256 or stronger digest before installation or execution.
  4. Where vendor signatures are available, validate them against a pinned and independently obtained signing key.
  5. If a mirror must be retained, proxy artifacts through an internally controlled repository that verifies upstream integrity and serves immutable, approved artifacts.
  6. Fail closed if checksum or signature verification cannot be completed; do not silently use an unverified browser.
  7. Run browser installation and execution as a dedicated non-privileged user in a sandboxed container with limited filesystem and network access.
  8. Document the mirror trust boundary and establish monitoring and periodic review of mirrored artifacts.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (23)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Claiming real Douyin scraping while only visiting pages or returning local simulated data is a material description-behavior mismatch. In a security setting, deceptive or inaccurate capability claims are risky because they can conceal unintended network/browser actions and undermine trust boundaries around what the skill is permitted to do.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Claiming real Douyin scraping while only visiting pages or returning local simulated data is a material description-behavior mismatch. In a security setting, deceptive or inaccurate capability claims are risky because they can conceal unintended network/browser actions and undermine trust boundaries around what the skill is permitted to do.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Claiming real Douyin scraping while only visiting pages or returning local simulated data is a material description-behavior mismatch. In a security setting, deceptive or inaccurate capability claims are risky because they can conceal unintended network/browser actions and undermine trust boundaries around what the skill is permitted to do.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 47)May include surrounding context.

md
node scripts/douyin_scraper.js search "海鲜" 10

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 48)May include surrounding context.

md
node scripts/douyin_scraper.js search "海鲜" 10

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · install_playwright_docker.py (reported line 30)May include surrounding context.

python
def mode_native() -> None:
    env = os.environ.copy()
    env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright")
    env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing")
    venv_python = Path("venv/bin/python")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The entire skill documentation is written only in Chinese and does not indicate that users may choose another language or locale. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill advertises executable workflows using Python/Node commands and optional Playwright installation, but it does not declare any explicit tool scope such as shell or file permissions. That creates an authorization and review gap: an agent may execute shell commands, write files, or access environment-dependent tooling without the skill metadata clearly constraining or disclosing those capabilities.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill uses broad natural-language trigger examples such as generic search requests without clear boundaries for when the agent should invoke it. That can lead to over-triggering, where unrelated user requests cause shell execution or scraping-related workflows to run unnecessarily, expanding attack surface and increasing the chance of unintended actions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The invocation guidance explicitly tells the agent to call the skill for natural-language requests without narrowing criteria or requiring user confirmation. In combination with shell-based commands and ambiguous behavior, this makes accidental or adversarial triggering more likely, potentially causing unnecessary command execution or data collection attempts.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This shell script contains user-facing natural-language strings only in Chinese, including status messages and usage instructions. The file does not offer any language/locale choice or explain that the skill is intentionally limited to Chinese-speaking users, which fits the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · install_playwright_docker.py (reported line 14)May include surrounding context.

python
def run(cmd: list[str], cwd: str | None = None, env: dict[str, str] | None = None) -> None:
    print("\n>>>", " ".join(cmd))
    subprocess.run(cmd, cwd=cwd, env=env, check=True)


def mode_official() -> None:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The script navigates Playwright to Douyin URLs and performs live network requests based on user input, which transmits browsing requests to an external service. While the code logs the URL being accessed, it does not clearly warn the user that running the command will contact Douyin and send the search term externally.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The function claims to search Douyin content but returns fabricated placeholder records even after making a real network request. This is a security-relevant integrity issue because downstream users or agents may trust the output as authentic external data and make decisions based on false information.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The hot-list feature similarly fabricates ranking data instead of scraping the advertised Douyin hot list. This misrepresents the provenance and accuracy of the data, creating an integrity risk for any automation, analytics, or content selection workflow consuming the results.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script emits and constructs natural-language content exclusively in Chinese, including mock data titles/descriptions and warning messages, with no opt-in or language selection. The policy for this category calls out forced language/locale behavior as a violation unless the tool explicitly offers a choice or clearly documents a justified regional constraint.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The dependency is specified with a lower-bound range (playwright>=1.40.0) instead of an exact pinned version, so installs may resolve to different future releases. This creates supply-chain and reproducibility risk: a later compromised or breaking upstream release could be pulled automatically into the skill without review.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
playwright>=1.40.0

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The script writes JSON output directly to the user-supplied filepath, which can overwrite existing files. Although it logs the save location after writing, there is no prior warning or confirmation that local files will be created or replaced.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The CSV save path is fully controlled by the caller and is written synchronously, potentially overwriting an existing file. The current message only reports success after the write and does not disclose the file-modifying behavior beforehand.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This file contains user-facing natural-language instructions and examples exclusively in Chinese, and the parser logic is also oriented around Chinese trigger phrases. Under the policy rule, forcing a specific language without opt-in can be a natural-language policy violation unless the locale constraint is explicitly justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This code conditionally writes JSON or CSV data to a user-supplied path, which is a file-write operation covered by the warning requirement for code files. While the CLI argument name suggests output, there is no confirmation, log message, or inline comment/docstring warning that a file will be created or overwritten.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.