Back to skill

Security audit

抖音选题分析

Security checks for vulnerabilities and agentic risk

Overview

This Douyin analytics skill is mostly purpose-aligned, but it should be reviewed because it passes its API token in URL query strings and automatically saves fetched datasets locally.

Install only if you are comfortable sending Douyin query targets and your Guaikei token to www.guaikei.com, and expect search/post/comment outputs to be saved under the skill's logs directory. Treat the token like a password, avoid shared terminals/workspaces, rotate it if exposed, and delete retained logs when they contain sensitive research targets or personal data.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
src/utils/request.js:98
Finding

API Token Exposed in URL Query Strings

Content
View full analysis

Vulnerability Details

File Location: src/utils/request.js:98-104, 128-135; token-bearing parameter construction also occurs in src/api/search.js:55-62, 102-109, src/api/post.js:23-28, 50-55, src/api/comment.js:38-44, 67-72, and src/api/hot.js:18-22
Vulnerability Type: Credential exposure through URL query parameters
Risk Level: Medium

Vulnerable Code

src/utils/request.js:98-104:

js
params.skill_name = skillName();

const fullPath = `${path}?${querystring.stringify(params)}`;
const jsonData = JSON.stringify(data);

const options = {
  host: constants.BASE_URL,

src/utils/request.js:128-135:

js
params._ = Date.now();

const fullPath = `${path}?${querystring.stringify(params)}`;
const options = {
  host: constants.BASE_URL,
  path: fullPath,
  method: "GET",
  headers: { "Accept-Encoding": "identity" },

Representative token construction in src/api/search.js:55-62:

js
const params = {
  _: Date.now(),
  token: token,
};

const data = {
  keyword,
  sort_type: sort,

Technical Analysis

The API token is inserted into the params object and serialized directly into the request URL by querystring.stringify(params). This occurs for both GET requests and POST requests. Consequently, requests contain URLs such as:

text
/api/douyin/general-search/info?...&token=API_TOKEN

HTTPS protects the request while it is in transit, but it does not prevent the complete URL from being recorded by the destination server, reverse proxies, API gateways, load balancers, observability platforms, debugging tools, or error telemetry. URL query strings are routinely retained in access logs, whereas authorization headers can be handled using established credential-redaction controls.

The API hostname is fixed to www.guaikei.com, so this is not an SSRF issue. The vulnerability is the unnecessary exposure and propagation of a reusable c ...[truncated 1248 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the token from all URL parameter objects.

  2. Send it through a standard authorization header, for example:

    js
    headers: {
      Authorization: `Bearer ${token}`,
      "Content-Type": "application/json",
      "Accept-Encoding": "identity",
    }
    
  3. Refactor getJson, postJson, and requestApi so credentials are accepted separately from ordinary query parameters and cannot be accidentally serialized into URLs.

  4. Ensure server, proxy, gateway, and telemetry configurations redact authorization headers and any legacy token query parameter.

  5. Rotate tokens that may previously have appeared in access or observability logs.

  6. Apply server-side expiration, least-privilege scopes, usage limits, and anomaly detection to reduce the effect of credential theft.

  7. Add automated tests asserting that generated request paths never contain token, api_key, or other credential fields.

T09 · Insecure Skill Coding Practices

Note
Location
src/utils/log.js:29
Finding

Automatic Persistent Storage of Potentially Sensitive Result Data

Content
View full analysis

Vulnerability Details

File Location: src/utils/log.js:29-35; invocation sites include src/douyin/search-cli.js:254-257, src/douyin/post-cli.js:157-162, and src/douyin/comment-cli.js:156-161
Vulnerability Type: Insecure local retention of collected data
Risk Level: Low

Vulnerable Code

src/utils/log.js:29-35:

js
const outputFilename = path.join(
  path.dirname(__filename),
  "..",
  "..",
  "logs",
  safeFilename,
);

try {
  await fs.promises.mkdir(path.dirname(outputFilename), { recursive: true });
  await fs.promises.writeFile(outputFilename, content);

Representative automatic invocation in src/douyin/search-cli.js:254-257:

js
await log.taskWrite(
  `${startTime}_${keyword}_${sort}_${time}_${duration}_${content}_search.json`,
  JSON.stringify(finalOutput, null, 2),
);

Automatic invocation in src/douyin/post-cli.js:157-162:

js
url = url.replace(/[^a-zA-Z0-9_-]/g, "");
url = url.replace("httpswwwdouyincomuser", "");
url = url.replace("httpsvdouyincom", "");
await log.taskWrite(
  `${startTime}_${url}_post.json`,
  JSON.stringify(finalOutput, null, 2),
);

Automatic invocation in src/douyin/comment-cli.js:156-161:

js
url = url.replace(/[^a-zA-Z0-9_-]/g, "");
url = url.replace("httpswwwdouyincomvideo", "");
url = url.replace("httpswwwdouyincomnote", "");
await log.taskWrite(
  `${startTime}_${url}_comment.json`,
  JSON.stringify(finalOutput, null, 2),
);

Technical Analysis

Every successful search, post, or comment operation serializes the complete output and writes it to the project-level logs/ directory. The retained output can contain the user's search intent, requested profile or content identifiers, public account information, content metadata, and comment text.

The write operation does not specify restrictive file permissions, expiration, cleanup, encryption, or an opt-out mechanism. File ...[truncated 1755 chars]

Remediation
View remediation

Remediation Suggestions

  1. Make persistent file output opt-in through an explicit option such as --output or --save.
  2. Clearly disclose what data is retained, where it is stored, and how long it remains.
  3. If storage is enabled, create files with restrictive permissions such as mode 0o600 and directories with mode 0o700.
  4. Avoid storing complete result objects by default; retain only fields required for the user's stated purpose.
  5. Provide a cleanup command and configurable retention period.
  6. Add the logs/ directory to .gitignore and exclude it from build, package, CI artifact, and backup workflows unless explicitly requested.
  7. Warn users before storing datasets containing usernames, comments, or other personal information.
  8. For higher-sensitivity deployments, encrypt retained results using a user-managed key and ensure secure deletion when retention expires.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (45)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是一个面向抖音数据检索与分析的技能,但提供的代码片段只是底层通用参数解析器(args parser),处理命令行输入和帮助信息生成。该代码没有网络请求、没有抖音相关接口调用、没有搜索/抓取/评论分析/热榜查询逻辑,也没有任何与抖音资源交互的行为。虽然参数解析器可作为此类技能的辅助实现细节,但单独看该代码块,其实际行为与声明的核心能力严重不符,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

代码片段仅实现了一个本地日志写入辅助函数 taskWrite,使用 fs 和 path 在项目 logs 目录下创建并写入文件,并对文件名进行基本安全清洗。这属于通用的本地文件操作/日志记录能力。相比之下,声明描述的是一个面向抖音的数据查询与分析技能,应包含网络请求、抖音搜索、热榜、作品列表或评论抓取等行为。当前代码没有体现任何与抖音平台交互或数据分析相关的逻辑,因此描述与实际行为存在明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a full-featured Douyin analytics and content research skill. However, the supplied code chunk does not implement any network calls, Douyin API access, scraping, search, trend lookup, creator content retrieval, or comment analysis. It merely reads package.json from the local filesystem to return the package name. This is materially different from the declared primary purpose, so the description does not accurately represent the behavior of the supplied code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

根据提供的代码片段,其功能局限于 TOKEN 管理:验证 token 格式、在无效时输出配置警告和推广信息、在有效时返回 token。这与声明中描述的抖音内容搜索、热榜查询、博主作品抓取、评论分析等核心能力明显不一致。虽然 token 管理可作为某些技能的辅助实现细节,但当前代码片段没有体现任何与抖音数据访问或分析直接相关的行为,因此就“描述是否准确代表该代码实际作用”而言,存在明显不匹配。

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 72)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger list is very broad and includes generic phrases like data analysis, competitor analysis, trend tracking, and short-video analytics. In an agent environment, overly broad activation can cause the skill to be selected in contexts where the user did not clearly request Douyin-specific external data retrieval, increasing the chance of unintended data disclosure to the third-party API or incorrect tool invocation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README states that scraped Douyin data is automatically exported to local JSON logs, but it does not prominently warn users that collected video, author, and comment data will be stored on disk. In a scraping/analytics skill, silent local persistence increases the risk of retaining personal data, sensitive research targets, or regulated content longer than intended, which can lead to accidental disclosure or mishandling.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The main setup instructions ask users to obtain and configure a persistent API token in an environment variable, but the security guidance is minimal and not prominent at the point of use. Because the token enables access to a third-party scraping service, weak handling, shell history exposure, shared-machine leakage, or accidental log/output disclosure could allow unauthorized use of the account or quota.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This file contains natural-language documentation and runtime messages exclusively in Chinese, including the module description, parameter docs, and retry log text. Under the language/locale policy, forcing a specific language without user opt-in or documenting a justified locale restriction can be a policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The natural-language content in comments and docstrings is entirely in Chinese, including the module description and parameter documentation. This imposes a specific language context without any visible opt-in or justification, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The file's human-readable comments and descriptions are entirely in Chinese, which imposes a specific language context without any indication of user choice or documented locale limitation. Under the language/locale policy, a forced language can be a policy issue unless the skill explicitly offers language selection or clearly states that it is intended only for a Chinese-language or region-specific use case.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code reads process.env.GUAIKEI_API_TOKEN, which is a credential access operation covered by the warning requirement for code files. Although the file logs URL handling and task progress, it provides no user-facing notice, prompt, or comment explaining that an API token from the environment will be used.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The CLI persists fetched comment data to a local JSON file after printing results, but there is no clear indication in this file that user data will be stored on disk. Comment contents and interaction data can contain personal or sensitive information, so silent persistence increases the risk of unintended retention, exposure to other local users/processes, and misuse beyond the user’s immediate query.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The program writes the complete result set to a local file without a clear warning or consent mechanism. In a skill focused on comment analysis, stored outputs may include user-generated content and metadata, so undisclosed persistence creates privacy, compliance, and local data exposure risks that are greater than what users may expect from a query command.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The CLI writes the full fetched results to a local JSON file automatically after successful execution, with no user confirmation, opt-in flag, or visible warning in this file. Because this skill handles scraped Douyin account and content data, the behavior can unintentionally persist third-party content metadata to disk, increasing exposure through local file disclosure, shared environments, backups, or later unintended reuse.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The CLI writes the full search output and request metadata to a local file automatically after every successful run. In this skill context, search terms, result sets, and timestamps may reveal user research interests, monitoring targets, or business intelligence, and the write behavior is not clearly disclosed in the skill description, creating unnecessary data persistence and exposure risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The file emits natural-language strings such as error messages in Chinese, and later success/failure messages also remain fixed to that locale. This can violate language/locale policy when the skill does not offer opt-in, fallback, or documentation that it is intentionally Chinese-only.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Several user-facing error messages in this file are hard-coded in Chinese, which imposes a language choice on users without any opt-in, fallback, or documented locale constraint. The policy forbids forcing a specific language unless the skill offers a language choice or clearly justifies the locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The parse/network/auth/validation error strings are presented only in Chinese, making the skill's user-visible behavior language-specific without user opt-in. This is a natural-language policy issue because the file does not offer locale selection or explain why Chinese-only output is required.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This code emits all user-facing warnings and status messages exclusively in Chinese, including error and informational text. The file does not provide any language selection, fallback, or documented region-specific justification, which creates a natural-language locale policy violation under the stated rules.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The natural-language fields in this schema are entirely in Chinese, including the title and property descriptions, with no indication that the skill is region-specific or that users may choose another language. Under the policy, forcing a specific language without opt-in can be a locale-policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

This JSON schema contains user-facing natural-language strings such as the title and descriptions entirely in Chinese. Because the file does not indicate that the skill is China-specific or provide any language/locale opt-in, it may violate a language/locale policy that requires neutral or user-selectable language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

This JSON schema contains natural-language titles and descriptions exclusively in Chinese, which can impose a specific language on downstream users or developers. The file does not indicate that the skill is China-specific or provide any opt-in or alternative locale, so it may violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The schema title and description are written only in Chinese, which imposes a specific language in user-facing metadata without any opt-in or documented locale limitation. Under the policy, language constraints should either offer user choice or be explicitly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:24