Back to skill

Security audit

抖音爆款爬虫

Security checks for vulnerabilities and agentic risk

Overview

The skill is a Douyin scraping helper, but it recommends logged-in browser use and its advertised real scraping paths can return fabricated sample records as if they were collected data.

Review before installing. Use an isolated browser profile or test account instead of your primary logged-in Douyin session, treat outputs as sample data unless live extraction is fixed, avoid running the browser unsandboxed on a sensitive machine, and pin or verify dependencies before installation.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/douyin_scraper.js:36
Finding

Chromium Browser Sandbox Disabled During Remote Page Processing

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Mutable Dependency Installation and Unverified Third-Party Browser Mirrors

Content
View full analysis
=1.40.0 ``` ```bash source venv/bin/activate pip install --upgrade pip pip install playwright # Install browser echo "" echo "Installing Playwright browser..." playwright install chromium ``` ```python def mode_native() -> None: env = os.environ.copy() env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright") env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing") venv_python = Path("venv/bin/python") venv_pip = Path("venv/bin/pip") if not venv_python.exists(): run([sys.executable, "-m", "venv", "venv"]) run([str(venv_pip), "install", "-r", "requirements.txt"], env=env) run([str(venv_python), "-m", "playwright", "install", "chromium"], env=env) ``` ### Technical Analysis The Python dependency declaration accepts every Playwright release from version 1.40.0 onward. The installation script similarly installs the newest version available at installation time and upgrades `pip` without pinning it. Consequently, two installations of the same reviewed project can execute materially different third-party code. Native installation also configures non-vendor mirror domains as default download sources for Playwright and Chromium artifacts. The project contains no lock file, package hashes, checksum verification, or signature-verification step that would independently establish the integrity of downloaded components. Using a mirror is not proof that the mirror is malicious. The security issue is that mutable executable dependencies and browser binaries are trusted without reproducible version and integrity controls. ### Attack Path 1. A user executes `install.sh` ...[truncated 1289 chars]
Remediation
View remediation

other

Warning
Location
scripts/scraper.py:77
Finding

Real Scraping Mode Returns Fabricated Records Without Reliable Provenance Marking

Content
View full analysis
list[VideoData]: if mock or sync_playwright is None: return self._mock_search(keyword, limit) try: with sync_playwright() as p: browser = p.chromium.launch( headless=self.headless, args=["--disable-blink-features=AutomationControlled", "--no-sandbox"], ) page = browser.new_page( user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" ) page.goto(f"https://www.douyin.com/search/{keyword}", wait_until="domcontentloaded", timeout=30000) time.sleep(self.delay) browser.close() except Exception as exc: print(f"Browser scraping failed ({exc}); returning sample data") return self._mock_search(keyword, limit) def hot(self, category: str, limit: int, mock: bool = False) -> list[VideoData]: if mock or sync_playwright is None: return self._mock_hot(category, limit) try: with sync_playwright() as p: browser = p.chromium.launch( headless=self.headless, args=["--disable-blink-features=AutomationControlled", "--no-sandbox"], ) page = browser.new_page( user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" ) page.goto("https://www.douyin.com/hot", wait_until="domcontentloaded", timeout=30000) time.sleep(self.delay) browser.close() except Exception as exc: print(f"Bro ...[truncated 4126 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (22)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill claims natural-language understanding and real Douyin content extraction, but the described implementation appears limited to simple keyword/CLI handling and may not actually parse or extract live platform data. Security-relevant documentation mismatches are dangerous because they obscure the true operational boundaries of the skill and can cause users to trust outputs or permissions that are not warranted.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The skill claims natural-language understanding and real Douyin content extraction, but the described implementation appears limited to simple keyword/CLI handling and may not actually parse or extract live platform data. Security-relevant documentation mismatches are dangerous because they obscure the true operational boundaries of the skill and can cause users to trust outputs or permissions that are not warranted.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

md
node scripts/douyin_scraper.js search "海鲜" 10

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 100)May include surrounding context.

md
node scripts/douyin_scraper.js search "海鲜" 10

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · install_playwright_docker.py (reported line 30)May include surrounding context.

python
def mode_native() -> None:
    env = os.environ.copy()
    env.setdefault("PLAYWRIGHT_DOWNLOAD_HOST", "https://npmmirror.com/mirrors/playwright")
    env.setdefault("PLAYWRIGHT_CHROMIUM_DOWNLOAD_HOST", "https://cdn.npmmirror.com/binaries/chrome-for-testing")
    venv_python = Path("venv/bin/python")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README explicitly recommends using a logged-in profile="user" browser session to bypass Douyin anti-bot protections, which causes the automation to operate with the user's authenticated cookies, account context, and potentially broader personal data visibility. In an agent skill context, this is dangerous because the skill may access account-scoped content or trigger actions under the user's identity without a clear consent boundary, privacy warning, or least-privilege guidance.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill advertises browser and script-based scraping behavior but does not declare explicit tool scope or permissions, despite requiring shell, environment, and file-writing capabilities. Undeclared capabilities make it harder for users or the platform to understand what the skill may access or modify, increasing the risk of overbroad execution and unintended data exposure.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The auto-activation triggers are broad everyday phrases, which increases the chance the skill will run in contexts where the user did not intend browser automation or scraping activity. Because this skill may invoke logged-in browsing or scripts, overbroad triggering raises the risk of accidental access to account-scoped content and unnecessary execution of sensitive capabilities.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill encourages use of browser profile="user" to reuse an existing logged-in Douyin session, but it does not prominently warn that this may access account-scoped data or perform actions in the user's authenticated context. Reusing a live login session materially increases privacy and account-risk exposure, especially if activation is broad or the browser automation is not tightly constrained.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The example triggers are very broad natural-language requests such as searching trending content or arbitrary video links, without any stated guardrails, scope limits, or exclusions. In a browser-automation scraping skill, this can cause the agent to invoke scraping behavior on loosely specified requests, increasing the chance of unauthorized data collection, policy violations, or use beyond the operator’s intended scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This shell script presents all user-facing output and usage guidance in Chinese, which effectively forces a specific language on users. The file does not offer an opt-in language choice or explain that the skill is intentionally limited to a Chinese-language audience.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · install_playwright_docker.py (reported line 14)May include surrounding context.

python
def run(cmd: list[str], cwd: str | None = None, env: dict[str, str] | None = None) -> None:
    print("\n>>>", " ".join(cmd))
    subprocess.run(cmd, cwd=cwd, env=env, check=True)


def mode_official() -> None:

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest and file header describe a scraper that obtains Douyin hot-video and text data, but both scraping paths populate hard-coded example fields after merely visiting the page. This means the implemented behavior is generating mock results rather than actually collecting Douyin data, which materially differs from the stated skill purpose.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The getHotVideos path is documented and exposed as fetching Douyin hot-list data, but the code constructs generic records like "热榜视频 1" and invented metrics. That behavior does not satisfy the manifest's promise to scrape real hot-video and copy data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The script accepts an arbitrary output path from command-line arguments and writes to it with writeFileSync, which will create or overwrite files without confirmation. In an agent or automation context, untrusted or mistaken inputs could clobber important local files or place data in sensitive locations, increasing the risk beyond a normal manual CLI workflow.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill advertises scraping Douyin hot-video/copy data, but both live-code paths merely open a page, sleep, and then always return fabricated mock records. This is dangerous because downstream users or agents may trust the output as real collected data, leading to silent integrity failures, bad business decisions, or deceptive automation behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The README is entirely in Chinese and the natural-language examples are only Chinese phrases, with no indication that other languages are supported or that the Chinese-only behavior is a deliberate, documented locale restriction. This can amount to a language policy issue if the skill forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

All natural-language instructions and trigger examples are written in Chinese, and the file does not indicate that the skill is region-specific or that users may choose another language. This can amount to a language policy issue if the organization expects skills not to force a specific language without opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This code writes Dockerfile.generated to disk, which is a file-modifying operation covered by the missing-warning rule for code files. While subprocess commands are printed, there is no user-facing disclosure, prompt, or explanatory comment/docstring warning that the script will create or overwrite a local file.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The dependency is specified with a lower-bound only (playwright>=1.40.0), which allows future major or minor releases to be installed without review. This creates supply-chain and reliability risk because a compromised, vulnerable, or breaking upstream release could be pulled automatically into the skill environment.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
playwright>=1.40.0

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file contains natural-language comments and CLI help text that force a specific language/locale experience. Under the policy, language constraints should either offer user choice or be clearly justified as region-specific; this file provides neither.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The script generates user-visible text and sample data descriptions exclusively in Chinese, and later prints those values to the console. Because there is no option to select another language or explicit documentation that the skill is intentionally Chinese-only, this is a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.