Back to skill

Security audit

Douyin Search

Security checks for vulnerabilities and agentic risk

Overview

This Douyin scraping skill mostly matches its stated purpose, but it handles full login cookies and authenticated browser state in ways that need careful Review before installation.

Install only if you are comfortable giving the skill a logged-in Douyin session. Do not paste cookies into chat; create any cookie file locally yourself, store credentials outside the repository if possible, restrict permissions, and avoid committing data/. Treat generated CSVs as untrusted user content and review the source before use despite the skill text saying not to.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:44
Finding

Agent Instructions Discourage Source Review and Request Full Authentication Cookies Through Chat

Content
View full analysis
/data/cookies-raw.txt # (老路径兼容: /tmp/douyin/cookies-raw.txt 也行) # 2. 转 Netscape 格式 python3 $SKILL/keepalive.py inject # 3. 验证 cookie python3 $SKILL/keepalive.py check # 退出码 0 = 有效 # 4. (仅评论需要) 持久化 ab session python3 $SKILL/keepalive.py state save ``` ```markdown 如果 cookie 不存在/过期/无效,立刻停止并提示用户重新导 cookies,或把整段 cookies 直接发给你,你落盘到 `$SKILL/data/cookies-raw.txt` 然后跑 inject → check → state save。 ``` The final sentence instructs the Agent to accept a complete cookie string from the user and write it to disk. ### Technical Analysis The Skill text alters how an Agent is expected to operate in two security-relevant ways: 1. It explicitly tells the Agent not to inspect the executable source. 2. It permits the Agent to ask the user to send an entire authenticated Douyin cookie set through the conversation. Source inspection is not incompatible with the Skill's declared functionality. Discouraging it weakens the review boundary and makes it less likely that unsafe implementation details will be detected. Complete browser cookies can contain reusable session credentials. Transmitting them through a conversational channel unnecessarily exposes them to chat retention, logging, telemetry, memory, or access by operators and integrations. The Skill only needs a local cookie file; it does not need the credential value to be present in the conversation. ### Attack Path 1. The Skill is loaded into an Agent session. 2. The Agent follows the instruction not to inspect the source code. 3. The cookie check fails or no local cookie file exists. 4. Following `SKILL.md`, the Agent asks t ...[truncated 701 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
comments-harvest.py:316
Finding

Predictable JavaScript Files in Shared Temporary Directory Enable Symlink and Race Attacks

Content
View full analysis
{{ const ls = item.innerText.split('\\n').map(s => s.trim()).filter(Boolean); if (ls.length < 2) return null; const user = ls[0] || '匿名'; let timeStr = ''; for (const l of ls) {{ if (/\\d+(天|小时|分|秒)前|刚刚|周前|月前|年前/.test(l) && l.length < 30) timeStr = l; }} let text = '', like = 0; for (let i = 1; i < ls.length; i++) {{ const l = ls[i]; if (l === '...' || l === '回复' || l === '展开' || l === '收起' || l === '置顶' || l === '作者' || l === '热' || l === '分享' || l === '举报') continue; if (timeStr && l === timeStr) continue; if (/^\\d+(\\.\\d+)?$/.test(l) && l.length < 8) {{ like = parseFloat(l); continue; }} if (l === user) continue; if (/^展开\\d+条回复$/.test(l)) continue; if (!text && l.length >= 2) text = l; }} return {{user, text, digg_count: like, time: timeStr}}; }}) .filter(Boolean) )""") r = subprocess.run( f"agent-browser eval \"$(cat {eval_fi ...[truncated 2505 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
aggregate.py:137
Finding

Untrusted Douyin Content Is Exported to CSV Without Spreadsheet Formula Neutralization

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
README.md:199
Finding

Credential Files Can Be Accidentally Committed Because the Documented Git Exclusions Are Missing

Content
View full analysis
" git tag v1.0.1 git push --tags ``` ```python raw_path = Path(COOKIE_RAW) if not raw_path.exists(): err(f"找不到 {COOKIE_RAW}") print(f" → 先把浏览器导出的 cookie 粘到该文件(k=v; k=v; 格式)", file=sys.stderr) sys.exit(1) raw = raw_path.read_text().strip() out_lines = [ "# Netscape HTTP Cookie File", "# https://curl.haxx.se/rfc/cookie_spec.html", "# This is a generated file! Do not edit.", "", ] count = 0 for part in raw.split(";"): part = part.strip() if not part or "=" not in part or part.startswith("douyin.com"): continue name, _, value = part.partition("=") name = name.strip() value = value.strip() if not name or not value: continue out_lines.append(f".douyin.com\tTRUE\t/\tFALSE\t0\t{name}\t{value}") count += 1 Path(COOKIE_FILE).write_text("\n".join(out_lines) + "\n") os.chmod(COOKIE_FILE, 0o600) ``` The audited project inventory contains no `.gitignore`, despite the README's assertion that the credential files are excluded. The code applies mode `0600` only to the converted cookie file, not to `cookies-raw.txt` or the saved browser state. ### Technical Analysis The recommended workflow stores complete session cookies and persistent browser state beneath the project tree. The documentation then recommends `git add -A`, which stages all unignored files. Because the supplied project has no ...[truncated 1687 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
downloader.py:124
Finding

Media Downloader Follows API-Controlled URLs and Redirects Without Destination Validation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (55)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill claims a limited scraping role but also describes saving video files locally, constructing filesystem paths, skipping existing files, and exposing remote download URLs for debugging. Those are meaningful side effects beyond simple read-only metadata retrieval, and failing to disclose them can mislead users into granting access they would not otherwise approve.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill claims a limited scraping role but also describes saving video files locally, constructing filesystem paths, skipping existing files, and exposing remote download URLs for debugging. Those are meaningful side effects beyond simple read-only metadata retrieval, and failing to disclose them can mislead users into granting access they would not otherwise approve.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill claims a limited scraping role but also describes saving video files locally, constructing filesystem paths, skipping existing files, and exposing remote download URLs for debugging. Those are meaningful side effects beyond simple read-only metadata retrieval, and failing to disclose them can mislead users into granting access they would not otherwise approve.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill explicitly tells the agent to ask for or accept the user's full Douyin cookies and write them to disk. That is highly dangerous because full session cookies can enable immediate impersonation of the user, and transmitting them through the agent channel or persisting them locally broadens the attack surface to logs, prompt history, shared storage, and other tools.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes the skill as read-only and focused on fetching/searching videos, users, works, and comments without modifying Douyin data. However, the --download option adds a materially broader capability: retrieving full video media and persisting it locally, which goes beyond basic metadata/comment scraping described in the manifest.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest says the skill is a read-only Douyin data collection tool focused on searching videos/users, fetching works, and reading comments, and explicitly says it does not modify Douyin data. This changelog shows the skill also downloads full video files locally, which is a materially broader behavior than basic metadata/comment retrieval described in the manifest.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes the skill as a read-only Douyin content scraping tool focused on searching videos/users, fetching works, and reading comments. This README adds a broader file-acquisition capability—saving full video media locally—which goes beyond basic data retrieval and is not mentioned in the manifest description.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · README.md (reported line 211)May include surrounding context.

python3 -c "import sys; sys.path.insert(0, '.'); from paths import report; print(report())"

text

**安全注意**: `data/cookies-raw.txt` 和 `data/cookies.txt` 包含你的抖音登录凭证,不要分享、提交到 git、或上传到云端。`.gitignore` 已经排除,但本地仍需注意权限(默认 `chmod 600`)。

---

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill advertises executable behavior involving shell, network, file read/write, and environment variables but does not declare any explicit tool scope or permission boundaries. This increases the chance that an agent will invoke powerful capabilities more broadly than intended, especially since the workflow includes cookie handling, browser state manipulation, and local file output.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs users to copy raw Douyin login cookies into local files and use them for authenticated scraping, but it does not prominently warn that these cookies are equivalent to account credentials. Storing raw session material on disk without explicit sensitivity, retention, and access-control guidance creates a substantial risk of account takeover or privacy compromise if the file is exposed.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The module docstring and all user-facing CLI descriptions are written exclusively in Chinese, which imposes a language/locale choice on users without opt-in. Under the stated policy, forcing a specific language is a natural-language policy violation unless the file explicitly offers a language choice or documents a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The top-level documentation frames the script as comment harvesting only, yet the implementation can also download full videos. In a security review, this mismatch is significant because deceptive or incomplete documentation can conceal broader data access and make operator consent uninformed.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file is positioned as a comment harvester but conditionally includes video download functionality, which expands behavior beyond the least-privilege expectation for the skill. Hidden or under-disclosed capability expansion is dangerous because users may grant trust or credentials for one purpose while the code performs another.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The warmup logic intentionally simulates human browsing patterns to reduce detection, including visiting other videos before returning to the target. Anti-detection behavior is risky in a nominally read-only scraping skill because it is not necessary for basic functionality and signals an attempt to evade platform safeguards.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · comments-harvest.py (reported line 310)May include surrounding context.

python
def run_ab(args, timeout=30):
    """subprocess wrapper for agent-browser"""
    return subprocess.run(
        ["agent-browser"] + args,
        capture_output=True, text=True, timeout=timeout,
    )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

This call executes a shell command with shell=True and embeds a path into command substitution: agent-browser eval "$(cat {js_file})". If TMP_DIR/js_file can be influenced through environment or path configuration, this creates shell-injection risk; even without direct injection, it needlessly routes code through a shell while handling browser-executed script content.

Content

Scanner excerpt · comments-harvest.py (reported line 320)May include surrounding context.

python
"""通过临时文件执行 ab eval(避免 shell 转义)"""
    js_file = str(TMP_DIR / "_ab_eval_tmp.js")
    Path(js_file).write_text(js)
    return subprocess.run(
        f"agent-browser eval \"$(cat {js_file})\"",
        shell=True, capture_output=True, text=True, timeout=timeout,
    )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

This is the same unsafe shell execution pattern in the aggressive harvest path. Because the script is explicitly designed to operate with authenticated browser state and stored cookies, a successful shell injection would run arbitrary local commands in a high-value context.

Content

Scanner excerpt · comments-harvest.py (reported line 387)May include surrounding context.

python
js = HARVEST_JS_AGGRESSIVE
        js_file = str(TMP_DIR / "_harvest_oneshot.js")
        Path(js_file).write_text(js)
        r = subprocess.run(
            f"agent-browser eval \"$(cat {js_file})\"",
            shell=True, capture_output=True, text=True, timeout=120,
        )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

The initialization step for human-mode harvesting again invokes shell=True with command substitution over a temporary file path. Repetition of this pattern increases exposure because it is executed routinely and could be abused anywhere path configuration is attacker-controlled.

Content

Scanner excerpt · comments-harvest.py (reported line 408)May include surrounding context.

python
js_file = str(TMP_DIR / "_harvest_step.js")
        Path(js_file).write_text(js)
        # 初始化(第一次 eval 走初始化分支)
        r = subprocess.run(
            f"agent-browser eval \"$(cat {js_file})\"",
            shell=True, capture_output=True, text=True, timeout=15,
        )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

This repeated per-round eval call uses the same shell-based pattern, making exploitation opportunities frequent during normal runs. The combination of shell execution and browser automation tied to a logged-in session raises the impact beyond a generic reliability issue.

Content

Scanner excerpt · comments-harvest.py (reported line 420)May include surrounding context.

python
stalled = 0
        total_added = 0
        for rnd in range(max_rounds):
            r = subprocess.run(
                f"agent-browser eval \"$(cat {js_file})\"",
                shell=True, capture_output=True, text=True, timeout=20,
            )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

The retry path duplicates the vulnerable shell invocation, so even error handling preserves the injection surface. Attackers often benefit from rarely reviewed retry branches, especially when they run under the same privileged local environment.

Content

Scanner excerpt · comments-harvest.py (reported line 428)May include surrounding context.

python
err(f"  round {rnd+1} eval 失败: {r.stderr[:100]}")
                # 重试 1 次
                time.sleep(2)
                r = subprocess.run(
                    f"agent-browser eval \"$(cat {js_file})\"",
                    shell=True, capture_output=True, text=True, timeout=20,
                )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

The flush step uses shell=True with an interpolated temporary path, preserving the same injection primitive late in execution. In this skill context, the script works with authenticated session state and local files, so arbitrary command execution could expose cookies, state, or harvested data.

Content

Scanner excerpt · comments-harvest.py (reported line 463)May include surrounding context.

python
})()
        """
        Path(str(TMP_DIR / "_harvest_flush.js")).write_text(flush_js)
        r = subprocess.run(
            f"agent-browser eval \"$(cat {TMP_DIR}/_harvest_flush.js)\"",
            shell=True, capture_output=True, text=True, timeout=10,
        )

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The script silently reads a local cookie file, parses session cookies, and reuses them for download operations without a strong user-facing warning at the point of use. In this context, those cookies are authentication material; unauthorized or unexpected reuse could expose the user's account session and expand the blast radius of any compromise.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The document content is written in Chinese and explicitly frames itself as the developer reference the agent should ignore, but it does not offer any language or locale choice. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is clearly justified.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Manifest 将该 skill 描述为抖音内容抓取且强调“只读 skill,不修改 douyin 任何数据”,主要列举的是搜索、作品与评论抓取。但本文档明确列出 downloader.py 为“视频下载 helper”,且后文专章说明下载实现与输出路径,表明该 skill 实际还会获取并保存完整视频文件,这比描述中的基础数据抓取更宽。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code loads a cookie header from a local cookie file and uses it for authenticated requests, which is access to sensitive session credentials. Although the module docstring mentions cookie paths, there is no explicit warning to the user that the skill consumes login cookies/session state for network access.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
comments-harvest.py:317