Back to skill

Security audit

抖音信息源

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does the advertised Douyin data lookup, but it needs review because it sends its API token in URL parameters and automatically saves retrieved public user/comment data to local files.

Install only if you are comfortable sending Douyin queries, URLs, and your GUAIKEI_API_TOKEN to guaikei.com, and be aware that successful searches, creator lookups, and comment pulls are saved as plaintext JSON logs in the skill directory. Use a limited/rotatable token, keep the workspace private, and delete logs when no longer needed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
src/utils/request.js:96
Finding

API Token Exposed in URL Query Parameters

Content
View full analysis

Vulnerability Details

File Location: src/api/search.js:59-63, src/api/search.js:104-107, src/api/comment.js:38-41, src/api/comment.js:67-71, src/api/post.js:23-28, src/api/hot.js:19-21, and src/utils/request.js:96-120
Vulnerability Type: Sensitive credential exposure through URL query parameters
Risk Level: Medium

Vulnerable Code

The API modules place the authentication token in the request parameter object:

js
const params = {
  _: Date.now(),
  token: token,
};

Search and comment result queries similarly include the token:

js
const params = {
  _: Date.now(),
  token: token,
  keyword: keyword,
  sort_type: sort,
  publish_time: time,
  filter_duration: duration,
  content_type: content,
  limit: limit,
};

The shared request function serializes these parameters directly into the URL:

js
params.skill_name = skillName();

const fullPath = `${path}?${querystring.stringify(params)}`;
const jsonData = JSON.stringify(data);

const options = {
  host: constants.BASE_URL,
  path: fullPath,
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Accept-Encoding": "identity",
    "Content-Length": Buffer.byteLength(jsonData),
  },
};

return await request(options, jsonData);

Technical Analysis

GUAIKEI_API_TOKEN is an authentication secret, but the application sends it as a token query parameter for search, comment, post, and hot-list requests. Although the connection uses HTTPS, TLS only protects the request while it is in transit. It does not prevent the complete URL from being retained by the destination service, API gateways, reverse proxies, access logs, application-performance monitoring systems, diagnostic traces, or infrastructure telemetry.

Query strings are commonly logged by default. Consequently, users and operators who can read URL logs may obtain a reusable API token even if they ...[truncated 1346 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove token from all URL parameter objects.

  2. Send the secret through an authorization header, for example:

    js
    const options = {
      host: constants.BASE_URL,
      path: fullPath,
      method: "POST",
      headers: {
        Authorization: `Bearer ${token}`,
        "Content-Type": "application/json",
        "Accept-Encoding": "identity",
        "Content-Length": Buffer.byteLength(jsonData),
      },
    };
    
  3. If the service does not support bearer authentication, use a dedicated secret header such as X-API-Key and update the server accordingly.

  4. Configure API gateways, reverse proxies, monitoring systems, and server logs to redact authentication headers and any legacy token query parameter.

  5. Avoid including secrets in error messages, tracing attributes, or request diagnostics.

  6. Rotate tokens that have previously been transmitted in URLs because historical infrastructure logs may retain them.

  7. Add automated tests that reject outbound request paths containing token=.

T09 · Insecure Skill Coding Practices

Note
Location
src/utils/log.js:27
Finding

Automatic Plaintext Persistence of Retrieved Personal Data

Content
View full analysis

Vulnerability Details

File Location: src/douyin/search-cli.js:258-261, src/douyin/post-cli.js:156-162, src/douyin/comment-cli.js:155-163, and src/utils/log.js:27-35
Vulnerability Type: Insecure local storage and excessive retention of retrieved data
Risk Level: Low

Vulnerable Code

Successful comment results are automatically serialized and saved:

js
url = url.replace(/[^a-zA-Z0-9_-]/g, "");
url = url.replace("httpswwwdouyincomvideo", "");
url = url.replace("httpswwwdouyincomnote", "");
await log.taskWrite(
  `${startTime}_${url}_comment.json`,
  JSON.stringify(finalOutput, null, 2),
);

Search results are also automatically persisted:

js
await log.taskWrite(
  `${startTime}_${keyword}_${sort}_${time}_${duration}_${content}_search.json`,
  JSON.stringify(finalOutput, null, 2),
);

The shared logging function creates the destination directory and writes the complete content without explicitly restrictive permissions:

js
const outputFilename = path.join(
  path.dirname(__filename),
  "..",
  "..",
  "logs",
  safeFilename,
);

try {
  await fs.promises.mkdir(path.dirname(outputFilename), { recursive: true });
  await fs.promises.writeFile(outputFilename, content);
  utils.printSuccess(`  → 已保存到 ${outputFilename}`);
} catch (error) {
  utils.printError(`日志写入失败: ${error.message}`);
}

Technical Analysis

Every successful search, post, or comment operation writes the complete structured response to the project-level logs/ directory. The saved results may contain public usernames, stable user identifiers, profile URLs, creator metadata, and comment content. No explicit user opt-in is required, and the implementation does not provide a no-save option, retention period, automatic cleanup, field minimization, pseudonymization, or explicitly restrictive file mode.

The behavior is documented in readme.md:97-103, so it is not hidden. ...[truncated 1779 chars]

Remediation
View remediation

Remediation Suggestions

  1. Make result persistence opt-in rather than automatic, such as through a --save flag.

  2. Provide a --no-save option if backward compatibility requires saving by default.

  3. Write files with restrictive permissions and create the directory with an appropriate mode:

    js
    await fs.promises.mkdir(path.dirname(outputFilename), {
      recursive: true,
      mode: 0o700,
    });
    
    await fs.promises.writeFile(outputFilename, content, {
      mode: 0o600,
      flag: "w",
    });
    
  4. Store only fields required for the user's stated purpose. Redact or pseudonymize stable user identifiers and profile URLs where they are unnecessary.

  5. Define a configurable retention period and automatically delete expired result files.

  6. Add logs/ to .gitignore and exclude it from packaging, support bundles, and cloud synchronization by default.

  7. Clearly notify users before saving data and document the data categories, destination, permissions, and retention behavior.

  8. For sensitive environments, support encrypted output using an operating-system credential store or a user-supplied encryption key.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (44)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

该代码片段仅是一个与业务无关的通用参数解析模块(parseArgs、readValueAfterFlag、buildHelp)。它不包含任何抖音平台相关处理,例如网络请求、热榜查询、视频搜索、作品抓取、评论获取、数据聚类或日报生成。因此,代码实际行为与声明的技能用途存在明显不一致。这不是单纯的底层支撑细节,因为当前提供的代码片段本身完全无法体现所声明的核心能力。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

该代码片段的实际功能是一个通用的本地日志写入辅助模块,不包含任何抖音平台请求、搜索、抓取、评论分析、热榜查询或聚类日报生成逻辑。虽然日志模块可以作为大型技能的配套实现细节,但当前提供的代码片段本身与声明的主要用途没有直接对应关系;其核心行为是文件系统写入,属于与声明能力明显不同的实现。因此就“该代码块实际做什么”而言,描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是一个功能完整的抖音数据采集与分析技能,但给出的代码片段只是一个辅助函数 skillName(),通过 fs 读取本地 package.json 并返回名称。就该代码片段本身而言,其行为与声明的主要用途完全不一致,也没有体现任何抖音搜索、热榜、博主抓取或评论分析功能。因此应判定为描述与代码行为不匹配。

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 72)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The trigger phrases include broad intents such as competitor analysis, sentiment monitoring, and data analysis that can overlap with many generic marketing requests. This can cause the agent to invoke the skill in contexts where the user did not clearly request Douyin-specific data access, leading to over-collection, unintended third-party data transfer, or privacy/compliance issues.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JSON schema is a manifest-type file, so natural-language policy checks apply. The title and description force a specific language/locale without any opt-in or documentation that the skill is intentionally Chinese-only, which can violate language-choice policy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Natural-language text throughout the schema uses only Chinese labels and descriptions for user-facing fields. Because no locale choice or explicit China-specific justification is provided in this file, it presents a language-policy concern under the natural-language policy rule.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The README advertises automatic JSON log export and later states outputs are saved under a logs directory, but it does not prominently warn that collected public video, author, and comment data will persist on disk. Persistent storage of scraped comments and engagement data can create privacy, retention, and data-handling risks, especially on shared machines or in team environments where logs may be copied or exposed unintentionally.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger mapping includes very broad, conversational phrases such as '搜索 / 搜一下 / 找 …' that can match common user requests without enough domain scoping. In an agent environment, this can cause unintended invocation of a data-collection skill, leading to surprise external requests, unnecessary token usage, or retrieval of platform data the user did not explicitly intend to query.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The intent recognition priority rules use generic phrases like '热搜 / 热点 / 榜单 / 今天什么火' and '搜索 / 搜一下 / 找 + 关键词' without sufficiently constraining them to Douyin. This raises the risk that unrelated requests are routed into this skill, which is especially problematic because the skill performs external data retrieval and stores results to disk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

This markdown file is written entirely in Chinese, beginning with the title at L01, and does not indicate that the skill is region-specific or that language is configurable. Under the policy rule, forcing a specific language without user opt-in or documented justification can be a locale-policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file documents commands for retrieving a creator's published works and comment data, which can affect user privacy and involve collection of third-party content. Under the markdown-specific warning rule, the description should disclose that the skill accesses platform data and may implicate privacy, data-use, or terms-of-service considerations, but no such warning appears in these sections.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This code sends a skill token and user-supplied video URL to remote API endpoints via requestApi, which is a network operation involving potentially sensitive data. Although the functions have developer-oriented docstrings, there is no user-facing disclosure, confirmation, or warning in this file about transmitting that data off-box.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The CLI writes the full comment results to a local JSON file after successful execution, and those results can contain user-generated content and engagement metadata that may be sensitive or regulated depending on deployment context. Because this persistence happens automatically and there is no visible consent, opt-in, retention control, or redaction in this file, it increases the risk of unintended local data exposure, especially on shared hosts, CI runners, or analyst workstations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This code uses Chinese-only user-facing strings throughout comments, error messages, and success output, such as at L21, L45, L49, L71, and L82. Under the policy, forcing a specific language without user opt-in or a documented justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill's descriptions, help text, and runtime messages are presented only in Chinese, and the file provides no mechanism for users to choose another language or explicitly opt into this locale. This can violate language/locale policy when a skill forces a specific language without user choice.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The CLI persists search queries and full results to a local file using a filename derived from user input. In a data-query skill, this creates unnecessary data retention and can expose potentially sensitive research terms, monitored topics, or returned content to other local users, backup systems, or later compromise of the host. The skill context makes this more relevant because analytics/search tooling often handles commercially sensitive monitoring queries even if it is not processing classic secrets.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The authentication error string is entirely in Chinese and includes remediation guidance, which imposes a specific language on users. Under the policy, language constraints should offer user choice or be clearly justified as region-specific; neither is evident in this file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The user-facing messages in this file are all written in Chinese, including warnings and status text, with no indication that the skill is China-specific or that users can opt into another language. This creates a natural-language policy issue because it forces a specific language on users without choice or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JavaScript file contains natural-language comments and runtime error/output strings exclusively in Chinese, including validation errors and formatted result messages. Because the file does not indicate that the skill is region-specific or provide any user opt-in for language selection, it risks violating the language/locale policy for natural-language content.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The natural-language examples include a broad phrase like '帮我做抖音竞品分析', which may be interpreted as sufficient to run the skill even when the user has not supplied necessary identifiers or confirmed a data-query intent. In a skill that sends requests to a third-party API, premature invocation increases the risk of unnecessary data processing and confusing or misleading automation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The schema title and description are written only in Chinese, which can constitute a language/locale policy issue when no user opt-in or documented locale limitation is provided. Because this is natural-language metadata in a JSON file, it falls under policy review for all file types.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

This JSON schema uses Chinese-only title and description text, which can constitute a language/locale policy issue when no opt-in, alternative locale, or region-specific justification is provided. Because the file is a schema/manifest-type artifact consumed by developers or tooling, the fixed language choice may conflict with an organizational requirement to avoid forcing a specific language by default.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This JSON schema contains human-readable title and description fields exclusively in Chinese, including the top-level title/description and property descriptions. Because this file provides natural-language interface metadata and does not indicate that the skill is China-specific or that language is configurable, it may violate a language/locale policy requiring user choice or documented locale constraints.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:24