Back to skill

Security audit

科研福音-抖音数据抓取神器

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent Douyin public-data scraper, but it handles an API token and saves scraped results in ways users should review before installing.

Install only if you are comfortable sending Douyin keywords, target URLs, limits, and your GUAIKEI_API_TOKEN to guaikei.com. Treat the token as sensitive, rotate it if exposed, and review or delete the generated logs directory because saved results can include public comments, nicknames, user IDs, profile identifiers, content metadata, and IP-region labels.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
src/utils/request.js:94
Finding

API Token Exposed in URL Query Strings

Content
View full analysis

Vulnerability Details

File Location: src/utils/request.js:94-120, src/utils/request.js:123-140; token parameters originate in src/api/comment.js:37-52,66-80, src/api/search.js:59-77,104-119, src/api/post.js:23-35,50-60, and src/api/hot.js:19-22
Vulnerability Type: Credential exposure through URL query parameters
Risk Level: Medium

Vulnerable Code

src/api/comment.js:37-52:

js
async function createCommentTask(token, url, limit) {
  const params = {
    _: Date.now(),
    token: token,
  };

  const data = {
    url,
    limit,
  };

  return await requestApi(
    "POST",
    "/api/douyin/comment/url",
    params,
    data,
    constants.CREATE_MAX_ATTEMPTS,
    "创建任务",
  );
}

src/utils/request.js:94-120:

js
async function postJson(path, params, data) {
  if (!path || typeof path !== "string") {
    throw new SkillError("PATH_INVALID", "path 必须是非空字符串");
  }
  if (!params || typeof params !== "object") {
    throw new SkillError("PARAM_INVALID", "params 必须是对象");
  }
  if (!data || typeof data !== "object") {
    throw new SkillError("DATA_INVALID", "data 必须是对象");
  }
  params.skill_name = skillName();

  const fullPath = `${path}?${querystring.stringify(params)}`;
  const jsonData = JSON.stringify(data);

  const options = {
    host: constants.BASE_URL,
    path: fullPath,
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      "Accept-Encoding": "identity",
      "Content-Length": Buffer.byteLength(jsonData),
    },
  };

  return await request(options, jsonData);
}

src/utils/request.js:123-140:

js
async function getJson(path, params) {
  if (!path || typeof path !== "string") {
    throw new SkillError("PATH_INVALID", "path 必须是非空字符串");
  }
  if (!params || typeof params !== "object") {
    throw new SkillError("PARAM_INVALID", "params 必须是对象");
  }
  par
...[truncated 2102 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove token from every query parameter object.

  2. Transmit credentials through a dedicated header, preferably:

    js
    headers: {
      Authorization: `Bearer ${token}`,
      "Content-Type": "application/json",
      "Accept-Encoding": "identity",
    }
    
  3. Refactor getJson and postJson to accept authentication separately from ordinary request parameters so callers cannot accidentally serialize secrets into URLs.

  4. Configure the API service, proxies, and monitoring infrastructure to redact Authorization and other credential-bearing headers.

  5. Avoid printing request options or authentication headers in errors and debug output.

  6. Rotate existing tokens because historical URL logs may already contain them.

  7. Add automated tests asserting that generated request paths never contain token, authorization, or the configured secret value.

T09 · Insecure Skill Coding Practices

Note
Location
src/utils/log.js:5
Finding

Fetched User Data Persisted in Plaintext by Default

Content
View full analysis

Vulnerability Details

File Location: src/utils/log.js:5-38, src/douyin/comment-cli.js:135-163, src/douyin/post-cli.js:136-162, and src/douyin/search-cli.js:233-260
Vulnerability Type: Unnecessary plaintext retention of user and content data
Risk Level: Low

Vulnerable Code

src/utils/log.js:5-38:

js
async function taskWrite(filename, content) {
  if (!filename || typeof filename !== "string") {
    utils.printError("日志文件名必须是非空字符串");
    return;
  }
  if (!content || typeof content !== "string") {
    utils.printError("日志内容必须是非空字符串");
    return;
  }
  let safeFilename = filename
    .replace(/[\\/:*?"<>|]/g, "_")
    .replace(/\.\.+/g, "_")
    .replace(/^\.+|\.+$/g, "");

  if (safeFilename.length > 200) {
    safeFilename = safeFilename.slice(0, 200);
  }
  if (!safeFilename) {
    safeFilename = `log_${Date.now()}`;
  }
  const outputFilename = path.join(
    path.dirname(__filename),
    "..",
    "..",
    "logs",
    safeFilename,
  );

  try {
    await fs.promises.mkdir(path.dirname(outputFilename), { recursive: true });
    await fs.promises.writeFile(outputFilename, content);
    utils.printSuccess(`  → 已保存到 ${outputFilename}`);
  } catch (error) {
    utils.printError(`日志写入失败: ${error.message}`);
  }
}

src/douyin/comment-cli.js:135-163:

js
const finalOutput = {
  status: "success",
  error_code: "OK",
  message: "获取评论任务完成",
  timestamp: new Date().toLocaleString(),
  request: {
    command: "comment",
    url: url,
    limit: limit,
  },
  metadata: {
    skill_version: constants.VERSION,
    runtime_version: process.versions.node,
    execution_time: Date.now() - startTime,
  },
  results: commentTask,
};
console.log(JSON.stringify(finalOutput, null, 2));
utils.printSuccess(
  `获取评论任务完成, 共返回 ${finalOutput.results.length} 条结果`,
);

url = url.replace(/[^a-zA-Z0-9_-]/g, "")
...[truncated 2758 chars]
Remediation
View remediation

Remediation Suggestions

  1. Disable persistent result logging by default.

  2. Add an explicit option such as --output PATH or --save-results and require informed user consent before writing fetched data.

  3. Create output files with restrictive permissions:

    js
    await fs.promises.writeFile(outputFilename, content, {
      encoding: "utf8",
      mode: 0o600,
      flag: "wx",
    });
    
  4. Validate that a user-selected output path remains within an approved directory, and continue preventing traversal and unsafe filenames.

  5. Minimize retained data by omitting unnecessary stable identifiers and ip_label unless the user explicitly requests them.

  6. Document the exact storage location, fields retained, retention period, and deletion procedure.

  7. Implement expiration or cleanup controls and a maximum aggregate storage size.

  8. Add logs/ to .gitignore and packaging exclusion rules to reduce accidental publication.

  9. For environments requiring retention, consider encrypted storage with keys managed separately from the project directory.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (43)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Reading local files such as package.json is not necessarily severe by itself, but when it is undeclared and unrelated to the stated purpose, it indicates hidden or poorly governed behavior. In an agent context, any undisclosed local file access increases uncertainty about what other files may be read and whether local metadata could be exfiltrated or misused.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Reading local files such as package.json is not necessarily severe by itself, but when it is undeclared and unrelated to the stated purpose, it indicates hidden or poorly governed behavior. In an agent context, any undisclosed local file access increases uncertainty about what other files may be read and whether local metadata could be exfiltrated or misused.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Reading local files such as package.json is not necessarily severe by itself, but when it is undeclared and unrelated to the stated purpose, it indicates hidden or poorly governed behavior. In an agent context, any undisclosed local file access increases uncertainty about what other files may be read and whether local metadata could be exfiltrated or misused.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Reading local files such as package.json is not necessarily severe by itself, but when it is undeclared and unrelated to the stated purpose, it indicates hidden or poorly governed behavior. In an agent context, any undisclosed local file access increases uncertainty about what other files may be read and whether local metadata could be exfiltrated or misused.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 72)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill declares access to an environment variable (GUAIKEI_API_TOKEN) but does not define an explicit tool/permission scope. In agent environments, undeclared secret access weakens reviewability and least-privilege controls, making it harder to verify what sensitive inputs the skill is allowed to consume.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

Broad trigger phrases can cause the skill to activate on loosely related requests, leading to over-broad routing and unnecessary use of external APIs or secrets. In this case, a data-scraping skill with external token use should be narrowly invoked; ambiguous triggers increase the chance of accidental data externalization or execution in the wrong context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JSON schema contains natural-language titles and descriptions exclusively in Chinese, which can impose a specific language on users or downstream maintainers without opt-in. The file does not indicate that the skill is region-specific or provide any alternative locale, so it conflicts with the language/locale policy criterion.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README explicitly promotes automatic JSON log export for search, post, and comment collection, but it does not clearly warn that scraped public data and potentially sensitive research datasets will be persisted to local files. This creates a real privacy and data-handling risk because operators may unknowingly retain large volumes of comments, author identifiers, and content metadata on disk, increasing exposure through accidental sharing, insecure backups, or multi-user systems.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The README instructs users to obtain and set a GUAIKEI_API_TOKEN but does not plainly disclose that the token will be sent to a third-party external service during operation. This is a security transparency issue: users may assume the tool is purely local because it is described as lightweight and Node.js-based, while in reality it depends on remote API access and transmits credentials off-host.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This documentation explicitly enables scraping Douyin search results, creator posts, comments, and trending data at scale, including up to 10,000 records, but provides no guidance on privacy, platform terms, rate limiting, or lawful use. In the context of a data-scraping skill, omission of these safeguards increases the risk of misuse for bulk collection, profiling, or policy-violating monitoring of user-generated content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This function sends the provided token and Douyin video URL to an external API, which is a safety-relevant data transmission operation. While the code has developer docstrings, there is no confirmation prompt, user-facing warning, or visible disclosure in this file that user data and credentials are being transmitted.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

This request sends the token, target URL, and limit parameters to an external endpoint to retrieve comment task results. In this code file there is no confirmation, user-visible logging, or explicit warning that these inputs are transmitted over the network.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code persists the fetched comment data to disk via log.taskWrite, which is a file-write operation affecting user data handling. Although the tool logs status messages, this file does not clearly disclose beforehand that results will be saved locally, nor does it ask for confirmation before creating the output file.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The CLI persists scraped Douyin post results to a local JSON file after completing the request, even though the stated skill purpose is query/analysis rather than local retention. Silent local storage increases data exposure on shared hosts, leaves behind recoverable artifacts, and may violate user expectations or data-handling constraints, especially when results include creator metadata or large-scale scraped content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The program writes fetched Douyin results to a local file without obtaining user confirmation or giving a clear user-facing warning in this CLI path. This can surprise users, create unintended data retention, and expose scraped data to other local users, backup systems, or later compromise of the machine.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This code builds and sends outbound HTTPS GET and POST requests, including URL parameters and JSON request bodies, but provides no user-facing warning, confirmation, or disclosure that data may be transmitted to a remote service. The existing comments and retry logging describe mechanics and errors, not the fact that user or system data may leave the local environment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The user-facing warning and status strings are all hard-coded in Chinese, and the file provides no indication that the skill is intended only for a Chinese-speaking or region-specific audience. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This JavaScript file contains natural-language comments and user-facing validation/error/output strings entirely in Chinese, including all messages shown to users. Under the stated policy, forcing a specific language without user opt-in or documented locale justification is a language/locale policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The manifest description, usage guidance, and trigger definitions are written entirely in Chinese and implicitly assume Chinese-language interaction, while no opt-in or alternative locale is offered. This can be a natural-language locale policy issue when a skill constrains user interaction language without documenting a justified region-specific limitation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

This JSON schema contains user-facing title and description text entirely in Chinese, including field descriptions, but does not indicate that the skill is region-specific or offer any language/locale opt-in. Under the policy for natural-language violations, forcing a specific language without user choice is reportable across all file types.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

This JSON schema includes the title and all field descriptions exclusively in Chinese, such as the title at L04 and descriptions through L41. That creates a natural-language locale constraint without any opt-in, fallback, or documentation that the skill is intentionally region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

This JSON schema uses Chinese natural-language strings for the title and field descriptions throughout, but the file does not document that the skill is China-specific or that language selection is optional. Under the stated policy, forcing a specific language without opt-in can be a locale-policy issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

This JSON schema uses Chinese-only natural-language title and description fields, including parameter descriptions, with no indication that the skill is China-region-specific or that other language options are supported. Under the policy check for natural-language constraints, this can be treated as a locale/language restriction without documented opt-in or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The human-readable title and descriptions are written only in Chinese, which imposes a specific language/locale in the skill metadata without offering any user choice or documenting a region-specific requirement. This matches the policy concern for language/locale constraints in natural-language content.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:24