Back to skill

Security audit

抖音数据助手

Security checks for vulnerabilities and agentic risk

Overview

The skill appears purpose-built for Douyin data retrieval, but it handles an API token insecurely and automatically stores collected social data locally.

Review before installing. Use a dedicated, revocable GUAIKEI_API_TOKEN with limited quota if possible, assume keywords and target URLs are sent to guaikei.com, and clean up the generated logs directory after use. Avoid running it in shared, backed-up, or CI workspaces when collecting comments or creator datasets unless local retention is acceptable.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
src/api/search.js:59
Finding

API Token Transmitted in URL Query Parameters

Content
View full analysis

Vulnerability Details

File Location: src/api/search.js:59-78
Additional Affected Locations: src/api/comment.js:38-52, src/api/comment.js:67-79, src/api/post.js:23-36, src/api/post.js:50-61, src/api/hot.js:19-22, src/utils/request.js:91-105, src/utils/request.js:126-136
Vulnerability Type: Sensitive credential exposure through URL query strings
Risk Level: Medium

Vulnerable Code

javascript
const params = {
  _: Date.now(),
  token: token,
};

const data = {
  keyword,
  sort_type: sort,
  publish_time: time,
  filter_duration: duration,
  content_type: content,
  limit: limit,
};

return await requestApi(
  "POST",
  "/api/douyin/general-search/keyword",
  params,
  data,
  constants.CREATE_MAX_ATTEMPTS,
  "创建任务",
);

The shared request utility converts these parameters, including the token, into the URL:

javascript
params.skill_name = skillName();

const fullPath = `${path}?${querystring.stringify(params)}`;
const jsonData = JSON.stringify(data);

const options = {
  host: constants.BASE_URL,
  path: fullPath,
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Accept-Encoding": "identity",
    "Content-Length": Buffer.byteLength(jsonData),
  },
};

Technical Analysis

The GUAIKEI_API_TOKEN credential is copied into params and serialized into the request URL. This pattern is used by search, post, comment, and hot-list API operations.

HTTPS protects the URL while it is in transit, but it does not prevent the URL from being recorded at endpoints or within operational infrastructure. Query strings are commonly captured by:

  • Reverse-proxy and web-server access logs
  • API gateways, load balancers, and monitoring systems
  • Error reports and distributed tracing systems
  • Server-side analytics and request histories
  • Debugging tools that record complete request paths

Unlike a ...[truncated 1876 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove token from every query-parameter object.
  2. Send the credential in an HTTP authorization header, preferably:
    javascript
    headers: {
      Authorization: `Bearer ${token}`,
      "Content-Type": "application/json",
      "Accept-Encoding": "identity"
    }
    
  3. Refactor request(), postJson(), and getJson() to accept the token separately from ordinary request parameters.
  4. Ensure retry and error messages never include request headers or complete request URLs.
  5. Configure server, proxy, API-gateway, tracing, and monitoring systems to redact Authorization, token, and equivalent credential fields.
  6. Rotate existing API tokens because historical server or monitoring logs may already contain them.
  7. Use short-lived, scoped tokens where the service supports them, and enforce revocation and quota-alerting controls.
  8. Add an automated test asserting that generated request paths never contain token, api_key, or other credential parameters.

T09 · Insecure Skill Coding Practices

Note
Location
src/douyin/comment-cli.js:135
Finding

Retrieved Douyin Data Is Automatically Persisted in Plaintext Files

Content
View full analysis

Vulnerability Details

File Location: src/douyin/comment-cli.js:135-163
Additional Affected Locations: src/douyin/search-cli.js:235-261, src/douyin/post-cli.js:136-163, src/utils/log.js:5-35
Vulnerability Type: Unprotected local storage and indefinite retention of collected data
Risk Level: Low

Vulnerable Code

javascript
const finalOutput = {
  status: "success",
  error_code: "OK",
  message: "获取评论任务完成",
  timestamp: new Date().toLocaleString(),
  request: {
    command: "comment",
    url: url,
    limit: limit,
  },
  metadata: {
    skill_version: constants.VERSION,
    runtime_version: process.versions.node,
    execution_time: Date.now() - startTime,
  },
  results: commentTask,
};
console.log(JSON.stringify(finalOutput, null, 2));
utils.printSuccess(
  `获取评论任务完成, 共返回 ${finalOutput.results.length} 条结果`,
);

url = url.replace(/[^a-zA-Z0-9_-]/g, "");
url = url.replace("httpswwwdouyincomvideo", "");
url = url.replace("httpswwwdouyincomnote", "");
await log.taskWrite(
  `${startTime}_${url}_comment.json`,
  JSON.stringify(finalOutput, null, 2),
);

The file-writing utility stores the content directly beneath the project directory:

javascript
const outputFilename = path.join(
  path.dirname(__filename),
  "..",
  "..",
  "logs",
  safeFilename,
);

try {
  await fs.promises.mkdir(path.dirname(outputFilename), { recursive: true });
  await fs.promises.writeFile(outputFilename, content);

Technical Analysis

Successful comment, post, and search operations automatically serialize their complete structured results and write them as plaintext JSON beneath the project-level logs directory. The stored content can include requested URLs, search keywords, public user identifiers, user profile URLs, comments, author information, and other returned metadata.

The feature is disclosed in readme.md, so it is not cover ...[truncated 2523 chars]

Remediation
View remediation

Remediation Suggestions

  1. Make result persistence opt-in through an explicit option such as --output FILE or --save.
  2. Do not create a log file when the user only requests stdout output.
  3. Separate diagnostic logs from collected datasets; diagnostic logs should not contain complete API responses.
  4. When persistence is enabled, create files with owner-only permissions:
    javascript
    await fs.promises.writeFile(outputFilename, content, { mode: 0o600 });
    
  5. Create the logs directory with restrictive permissions and verify existing permissions before writing.
  6. Add configurable retention and automatic cleanup, such as deletion after a user-selected number of days.
  7. Minimize stored fields and provide a redaction option for usernames, stable user IDs, profile URLs, and comment text.
  8. Consider encryption at rest when datasets must be retained in shared or backed-up environments.
  9. Add logs/ to .gitignore and exclude it from package publication, CI artifacts, cloud synchronization, and routine backups unless explicitly required.
  10. Clearly notify the user before saving data and provide a documented deletion command.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (55)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是面向抖音平台的数据检索与分析技能,核心能力应涉及抖音内容搜索、热榜数据获取、博主作品列表抓取、评论内容分析等。给定代码却仅实现了通用命令行参数解析与帮助文本生成:识别 flags、校验参数值、处理布尔选项、位置参数、默认值、必填项,并生成 help 文本。它既没有抖音相关 API 调用、网页抓取、数据处理,也没有任何与热榜、作品、评论、达人数据相关的逻辑。因此该代码块的实际行为与声明用途存在明显且实质性的不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个面向抖音数据抓取与分析的技能,但所给代码片段仅是一个通用日志模块,用于将字符串内容写入本地 logs 目录。虽然日志记录可作为辅助实现细节存在,但当前代码片段本身没有体现任何与抖音数据查询、热榜、作品抓取、评论分析相关的行为。由于评估的是描述与该代码片段实际行为是否一致,这里存在明显不匹配:代码的直接功能是本地文件写入,而非声明的抖音数据能力。

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 72)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger list is broad enough to match general analytics, monitoring, or content-analysis requests without a clearly bounded Douyin-specific intent. Over-broad activation can cause the wrong skill to run, sending user-supplied text, links, or identifiers to a third-party API unnecessarily and increasing data-exposure and consent risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The README promotes automatic JSON logging/export of scraped videos, author data, and comments without clearly warning about retention, sensitivity, downstream sharing, or deletion practices. Given that comment text and interaction data may contain personal information or sensitive opinions, this increases the risk of privacy misuse, accidental disclosure, and non-compliant secondary use.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases include very broad everyday language such as 搜索/找/热点/评论/视频, which can cause this skill to activate outside clearly scoped Douyin-only requests. In an agent environment, overbroad activation can lead to unintended scraping actions, wrong-tool invocation, and collection or export of third-party platform data when the user did not specifically intend that behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The intent recognition priority rules are ambiguous and rely on loose keyword matching rather than strict scope constraints. Because the skill performs scraping, ambiguous routing can cause the agent to execute data-collection commands for unrelated or underspecified user requests, increasing the chance of unintended external access and unnecessary data handling.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The document operationalizes large-scale Douyin data collection, including keyword search, creator post scraping, and comment harvesting up to 10,000 records, but provides no privacy, lawful-use, rate-limit, or data-handling guidance. In this skill context, that omission matters because the skill is explicitly intended for monitoring, competitor analysis, and public-opinion analysis, which increases the likelihood of bulk profiling, excessive scraping, or non-compliant downstream use.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This file contains natural-language strings in Chinese, including the module description and operation labels, but provides no indication that users can choose their preferred language. Under the policy criteria, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The string "创建任务" appears to be a user-visible action label passed into the request helper, and it is fixed to Chinese with no evidence of locale negotiation or opt-in. This may violate language policy unless the skill is explicitly documented as region- or language-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The string "查询任务" is another fixed Chinese user-facing label with no sign of user language preference handling. Absent explicit regional scoping, this constitutes a language/locale policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The CLI persists fetched Douyin comment data to a local JSON file by default after printing results, without explicit user consent, opt-in, or disclosure in the normal execution flow. Comment data can contain personal data, usernames, or sensitive speech content, so silent local retention increases privacy, compliance, and data exposure risk if the host is shared, backed up, or later accessed by another process or user.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code uses Chinese-only natural-language strings in comments, errors, and success messages, such as the function description and user-facing output. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly justified.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The CLI writes the full fetched Douyin post results to a local JSON file after printing them, creating an undisclosed secondary data sink. Even if the data is nominally public, results may include scraped metadata or analysis outputs that users do not expect to be retained on disk, increasing privacy, compliance, and local data exposure risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code saves the complete API response to a local file without warning the user or requesting consent. This can expose scraped content, metadata, or derived analytics to other local users, backups, or log collection systems, especially in shared or managed environments.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The code retrieves a credential from an environment variable and uses it to create and fetch a remote search task, which implies transmission of the user's keyword and filters over the network. While network activity is part of a search tool's purpose, this file lacks a visible disclosure in help text or runtime output that user-provided queries are sent to an external service.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Search results are written to a local JSON file without any explicit notice, consent, or retention control. In this skill context, results may contain user query terms and scraped content metadata, which can create unintended local data persistence and privacy exposure on shared systems or managed environments.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This file contains multiple user-facing error messages in Chinese only, such as the fallback message at this line, without offering a language choice or documenting a justified locale restriction. The policy requires avoiding forced language/locale behavior unless the skill explicitly offers opt-in or is clearly region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The authentication failure text is entirely in Chinese and includes support instructions, making the user experience language-specific with no indication of alternate locales or user choice. This is a natural-language policy issue unless the skill is clearly documented as region-specific or provides localization options.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language example includes underspecified requests like broad competitor analysis, which may encourage premature invocation before collecting the required keyword, link, or account identifier. In this skill context, that can lead to incorrect routing or unnecessary third-party data handling, though the impact is moderated by the documented instruction to ask follow-up questions first.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The title and descriptions are written entirely in Chinese, and the file provides no indication that this schema is intended only for a Chinese-language or region-specific audience. Under the policy, forcing a specific language without user opt-in or documented locale justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This JSON schema contains user-facing title and description fields entirely in Chinese, including field descriptions, but does not indicate that the skill is intentionally limited to Chinese-speaking users or provide any language/locale opt-in. Under the policy, forcing a specific language without user choice can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

This JSON schema includes natural-language title and description strings exclusively in Chinese, which imposes a specific language/locale on users of the skill metadata. The file does not indicate that Chinese is optional, user-selected, or required for a region-specific purpose, so it may violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The human-readable title and descriptions are entirely in Chinese, and the file provides no indication that this skill is region-specific or that users can opt into another language. Per the policy, hard-coding a specific language without user opt-in or justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:24