Back to skill

Security audit

抖音评论采集·主页作品抓取

Security checks for vulnerabilities and agentic risk

Overview

This Douyin scraping skill is mostly coherent, but it stores and reuses live account cookies and includes platform-bypass techniques that need careful review before installation.

Install only in an isolated environment and only with a Douyin account you are authorized to use. Treat generated cookie files as full account credentials, delete them when finished, do not commit or share outputs, and avoid using this on untrusted networks because the authenticated HTTP path disables TLS verification. Review whether the platform-bypass behavior is acceptable for your use case before running browser or cookie modes.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/fetch_comments.py:267
Finding

Authenticated Douyin Requests Disable TLS Certificate Verification

Content
View full analysis

Vulnerability Details

File Location: scripts/fetch_comments.py:197-203, 259-267
Vulnerability Type: Improper TLS certificate validation during authenticated requests
Risk Level: Medium

Vulnerable Code

python
def _www_probe(cookies: dict, timeout: float = 8.0) -> str:
    """Probe the Douyin homepage using the same cookie."""
    try:
        h = dict(_HEADERS)
        h["Cookie"] = "; ".join(f"{k}={v}" for k, v in cookies.items())
        r = httpx.get("https://www.douyin.com/", headers=h, timeout=timeout,
                     verify=False, follow_redirects=True)
        return f"status={r.status_code} len={len(r.text)}"
    except Exception as exc:
        return f"exc={type(exc).__name__}:{exc}"
python
headers = dict(_HEADERS)
headers["user-agent"] = fp["ua"]
headers["referer"] = cfg["referer_tpl"].format(aweme_id)
headers["sec-fetch-site"] = cfg["sec_fetch_site"]
headers["Cookie"] = "; ".join(f"{k}={v}" for k, v in cookies.items())

resp = httpx.get(final_url, headers=headers, timeout=timeout, verify=False)

Technical Analysis

The HTTP comment collection path constructs a Cookie header containing the user-provided Douyin session cookies and sends it while explicitly setting verify=False. This disables certificate-chain and hostname validation, so the client cannot authenticate that it is communicating with the genuine creator.douyin.com or www.douyin.com server.

The main comment request at line 267 is reached whenever the user invokes authenticated HTTP comment collection. The _www_probe request at lines 197-203 is additionally reached when a comment response cannot be parsed. Both requests transmit the complete parsed cookie collection without valid TLS peer authentication.

This is a regular exploitable implementation vulnerability rather than evidence of intentional credential theft. The configured destinations are consistent with the Skill’s decla ...[truncated 1621 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove verify=False from the authenticated comment request:

    python
    resp = httpx.get(final_url, headers=headers, timeout=timeout)
    
  2. Restore certificate verification for the diagnostic probe:

    python
    r = httpx.get(
        "https://www.douyin.com/",
        headers=h,
        timeout=timeout,
        follow_redirects=True,
    )
    
  3. Remove the warning-suppression logic for insecure TLS requests so future certificate-validation regressions remain visible.

  4. If a private enterprise certificate authority must be supported, accept an explicit trusted CA bundle rather than disabling verification:

    python
    with httpx.Client(verify="/path/to/enterprise-ca.pem") as client:
        resp = client.get(final_url, headers=headers, timeout=timeout)
    
  5. Fail closed on certificate errors. Do not retry authenticated requests with validation disabled.

  6. Add automated tests that intercept the request with an untrusted certificate and verify that both the main request and diagnostic probe reject the connection before transmitting authenticated application data.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (45)

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The documented behavior claims broader capabilities than the provided implementation reportedly supports, including homepage post scraping, time-window filtering, daily distribution analysis, and a unified router. This mismatch is dangerous because users or orchestration systems may grant broader trust, provide credentials, or rely on controls that do not actually exist, which can lead to misuse, unsafe execution paths, or accidental disclosure when operators troubleshoot missing features with real cookies and accounts.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The documentation instructs users to copy the full Cookie header from creator.douyin.com and reuse it for requests, but does not warn that this exposes live authenticated session credentials. Full-session cookies can enable account takeover, unauthorized access to creator data, or misuse beyond the intended operation if stored, logged, or shared improperly.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The help text explicitly advertises using browser mode and login state to get past the visible comment wall and retrieve more or all comments. That directly contradicts the stated safety boundary of collecting only authorized public data without bypassing login restrictions, indicating the skill is designed to evade access controls rather than merely access public content.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The browser-assisted flow uses authenticated cookies to collect data beyond what is visible to a guest session, specifically to overcome the platform's comment wall. In context, this is a functional access-control bypass against the manifest's claimed limitation, and it increases legal, privacy, and platform-abuse risk.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · scripts/install.py (reported line 14)May include surrounding context.

python
python install.py --all           # 必装 + 词频(jieba) + 浏览器(playwright+chromium)
  python install.py --with-analyze  # 额外装 jieba
  python install.py --with-browser  # 额外装 playwright 并下载 chromium
  python install.py --no-check      # 装完不跑自检

退出码:0 = 必装装好且自检通过;1 = 必装安装/自检失败。
"""

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · scripts/install.py (reported line 98)May include surrounding context.

python
python install.py --all           # 必装 + 词频(jieba) + 浏览器(playwright+chromium)
  python install.py --with-analyze  # 额外装 jieba
  python install.py --with-browser  # 额外装 playwright 并下载 chromium
  python install.py --no-check      # 装完不跑自检

退出码:0 = 必装装好且自检通过;1 = 必装安装/自检失败。
"""

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill advertises and documents capabilities that read/write local files, invoke shell commands, access environment variables, and perform network operations, but it does not declare any explicit tool scope or permissions boundary. In an agent ecosystem, this creates a trust and containment gap: the platform or user cannot easily evaluate whether the skill should be allowed to install packages, persist cookies, or make outbound requests before execution.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The changelog explicitly states that reusing login cookies can 'bypass guest anti-abuse controls' to retrieve older or more complete post data. Even if this uses the account owner's cookie, advertising bypass of platform anti-bot/rate-limit controls expands collection beyond normal guest access and conflicts with the stated restriction of only collecting authorized public data without bypassing controls. In the context of a scraping skill, this makes the behavior more dangerous because it normalizes evasion of platform protections and could be repurposed for broader unauthorized harvesting.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

These sections document reverse-engineered request signing and token acquisition, including a_bogus generation and obtaining a 'real' msToken from an official endpoint, which are commonly used to emulate trusted clients and bypass normal request integrity checks. In a data-collection skill, this meaningfully expands capability from ordinary API use into platform defense evasion, increasing abuse potential and legal/privacy risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The browser workflow is designed to intercept comment/list network responses from a real browser session to extract data that the page fetches, without disclosing the privacy, consent, or collection implications to the user. Even if comments are public, covert interception of session traffic as a collection primitive can capture more data than expected and normalizes opaque scraping behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation explicitly recommends anti-headless detection evasion techniques such as disabling Blink automation features and overriding navigator.webdriver, which goes beyond ordinary collection of authorized public data and into bypassing platform defenses. In the context of a scraping skill, this increases the likelihood of stealthy data collection against a service's protective controls and can facilitate policy or access-control circumvention.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Line L77 instructs users to set the browser UA/viewport/locale/timezone to a desktop Chinese environment. This is a natural-language locale constraint presented as required behavior, but the file does not offer user opt-in or clearly justify the restriction as a region-specific tool policy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This file contains user-facing natural-language content in Chinese in the module docstring and usage guidance, and later raises a Chinese-only installation message. The policy requires flagging language or locale constraints when the skill forces a specific language without user opt-in or clear justification.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code comments explicitly frame authenticated cookie injection as a way to get past the platform's 'comment wall' and retrieve more or all comments. That is dangerous because it operationalizes bypass of anti-abuse or access-friction mechanisms using stored authenticated sessions, moving the tool from ordinary browsing into assisted access-control circumvention.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script automatically loads and injects a local cookie file into a browser session, which reuses sensitive authentication material without a strong consent boundary or safety friction at the point of use. In a scraping tool whose purpose is to collect more data after guest access is limited, this increases the risk of credential misuse, accidental use of another person's session, or silent expansion from public-data collection into authenticated account access.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The CLI help and warnings repeatedly direct the operator to log in and reuse stored cookies when guest access is blocked, effectively coaching users through defeating anti-abuse barriers. In context, this materially increases misuse risk because the skill is a data-collection utility specifically designed to escalate from limited public scraping to authenticated extraction when friction appears.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Auto-installing packages via pip during execution is unrelated to the core task of comment collection and expands the script's trust and attack surface. If package resolution or indexes are compromised, the tool may pull and run untrusted code in the operator's Python environment.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
90% confidence
Finding

The script executes a pip install automatically at runtime when httpx is missing. This creates an unnecessary code-execution and supply-chain risk boundary during normal use, because network-retrieved packages are installed and executed without explicit user approval or environment isolation.

Content

Scanner excerpt · scripts/fetch_comments.py (reported line 48)May include surrounding context.

python
f"({sys.executable}) ...\n"
    )
    try:
        subprocess.check_call(
            [sys.executable, "-m", "pip", "install", "--quiet", "httpx>=0.27"]
        )
    except Exception as exc:  # noqa: BLE001

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
96% confidence
Finding

The script globally suppresses TLS verification warnings and repeatedly uses verify=False for HTTPS requests. In a tool that sends authentication cookies and retrieves remote content, this makes man-in-the-middle interception substantially easier and can expose session tokens, response tampering, and incorrect aweme_id resolution.

Content

Scanner excerpt · scripts/fetch_comments.py (reported line 61)May include surrounding context.

python
sys.exit(1)
    import httpx  # noqa: E402

try:  # 关掉 verify=False 的告警噪音
    import urllib3
    urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
except Exception:  # noqa: BLE001

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
97% confidence
Finding

Disabling TLS verification while resolving shared URLs allows a network attacker to tamper with redirects or page contents used to extract the target aweme_id. In this skill, that can misdirect collection, poison results, or facilitate cookie/session exposure in adjacent flows.

Content

Scanner excerpt · scripts/fetch_comments.py (reported line 136)May include surrounding context.

python
headers={"user-agent": UA, "accept": "text/html,*/*"},
            timeout=timeout,
            follow_redirects=True,
            verify=False,
        )
        final = str(r.url)
        body = r.text

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
97% confidence
Finding

The probe request sends the assembled Cookie header to douyin.com with verify=False, so an active attacker on the network path could intercept or alter authenticated traffic. Because these cookies are central to the authenticated scraping mode, compromise can lead to account/session theft or spoofed diagnostics.

Content

Scanner excerpt · scripts/fetch_comments.py (reported line 203)May include surrounding context.

python
h = dict(_HEADERS)
        h["Cookie"] = "; ".join(f"{k}={v}" for k, v in cookies.items())
        r = httpx.get("https://www.douyin.com/", headers=h, timeout=timeout,
                     verify=False, follow_redirects=True)
        return f"status={r.status_code} len={len(r.text)}"
    except Exception as exc:  # noqa: BLE001
        return f"exc={type(exc).__name__}:{exc}"

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
98% confidence
Finding

The main comment-fetch request includes authentication cookies and signature parameters but explicitly disables TLS verification. This exposes the most sensitive transaction in the tool to interception and response manipulation, which is especially dangerous given the script's use of authenticated creator endpoints and reverse-engineered signing.

Content

Scanner excerpt · scripts/fetch_comments.py (reported line 267)May include surrounding context.

python
headers["sec-fetch-site"] = cfg["sec_fetch_site"]
    headers["Cookie"] = "; ".join(f"{k}={v}" for k, v in cookies.items())

    resp = httpx.get(final_url, headers=headers, timeout=timeout, verify=False)
    status = resp.status_code
    try:
        resp_json = resp.json()

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The code writes comment content, nicknames, user IDs, and related metadata to an arbitrary output path, and similarly writes analysis output to disk. Although file output is part of the CLI functionality, there is no explicit warning in comments, docstrings, or user-facing messages that these outputs may contain personal or sensitive data from third-party users.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The browser context is hard-coded to locale="zh-CN" and timezone_id="Asia/Shanghai", which enforces a specific language/locale behavior. Under the policy, locale constraints should either be user-selectable or clearly justified as region-specific; this file does not provide such opt-in or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code file contains natural-language documentation and CLI help text that force a specific language/locale for users. Under the policy, locale-specific behavior should either provide user opt-in/choice or be clearly justified as region-specific; neither is present here.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.insecure_tls_verification

HTTPS certificate verification is disabled.

Warn
Code
suspicious.insecure_tls_verification
Location
scripts/fetch_comments.py:61