Back to skill

Security audit

抖音蓝海爆品采集助手

Security checks for vulnerabilities and agentic risk

Overview

This Douyin analytics skill is mostly coherent and read-only, but it needs review because it sends API tokens in URL query parameters and automatically saves collected data as plaintext logs.

Review before installing. Use a dedicated, revocable GUAIKEI token with limited quota, avoid running it in shared or synced workspaces, and delete or protect the logs directory because result files may contain collected public user and comment data. Do not use it for private or login-gated Douyin data.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
src/utils/request.js:91
Finding

API Credential Exposed in URL Query Parameters

Content
View full analysis

Vulnerability Details

File Location: src/utils/request.js, lines 91-120; credential parameters originate from API modules such as src/api/search.js, lines 59-61
Vulnerability Type: Sensitive credential exposure through request URLs
Risk Level: Medium

Vulnerable Code

javascript
// src/api/search.js
const params = {
  _: Date.now(),
  token: token,
};
javascript
// src/utils/request.js
params.skill_name = skillName();

const fullPath = `${path}?${querystring.stringify(params)}`;
const jsonData = JSON.stringify(data);

const options = {
  host: constants.BASE_URL,
  path: fullPath,
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Accept-Encoding": "identity",
    "Content-Length": Buffer.byteLength(jsonData),
  },
};

return await request(options, jsonData);

The same pattern is used by the search, post, comment, and hot API operations.

Technical Analysis

The value of GUAIKEI_API_TOKEN is inserted into the params object and serialized into the URL query string. Consequently, requests use URLs resembling:

text
https://www.guaikei.com/api/douyin/general-search/keyword?token=SECRET&...

HTTPS protects the query string while it is in transit, but it does not prevent the complete URL from being recorded after TLS termination. Query strings are commonly captured by:

  • Reverse-proxy and load-balancer access logs
  • API gateway and web server logs
  • Application performance monitoring systems
  • Network debugging tools
  • Error reports and request tracing systems
  • Upstream analytics or observability services

The credential is repeatedly included in both task-creation and polling requests. Polling may run up to 20 times, increasing the number of records containing the token.

Authentication secrets should be transmitted in an HTTP authorization header rather than in a URL. The source review found no evi ...[truncated 1471 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the token from every query-parameter object.

  2. Send the credential in an HTTP authorization header:

    javascript
    const options = {
      host: constants.BASE_URL,
      path: fullPath,
      method: "POST",
      headers: {
        Authorization: `Bearer ${token}`,
        "Content-Type": "application/json",
        "Accept-Encoding": "identity",
        "Content-Length": Buffer.byteLength(jsonData),
      },
    };
    
  3. Refactor postJson and getJson to accept the token separately from ordinary parameters, preventing accidental serialization.

  4. Ensure client-side error messages, tracing, and debug logs redact Authorization, token, and equivalent secret fields.

  5. Configure the server, reverse proxies, and API gateways to redact historical query parameters.

  6. Rotate existing tokens because prior invocations may already have placed them in access logs.

  7. Prefer short-lived, narrowly scoped tokens with revocation and usage-monitoring support.

  8. Add automated tests asserting that generated request paths never contain the token.

T09 · Insecure Skill Coding Practices

Warning
Location
src/douyin/comment-cli.js:151
Finding

Automatic Plaintext Persistence of Collected User Data

Content
View full analysis

Vulnerability Details

File Location: src/douyin/comment-cli.js, lines 151-162; shared writer in src/utils/log.js, lines 24-35
Vulnerability Type: Insecure local storage and excessive retention of collected data
Risk Level: Medium

Vulnerable Code

javascript
// src/douyin/comment-cli.js
console.log(JSON.stringify(finalOutput, null, 2));
utils.printSuccess(
  `获取评论任务完成, 共返回 ${finalOutput.results.length} 条结果`,
);

url = url.replace(/[^a-zA-Z0-9_-]/g, "");
url = url.replace("httpswwwdouyincomvideo", "");
url = url.replace("httpswwwdouyincomnote", "");
await log.taskWrite(
  `${startTime}_${url}_comment.json`,
  JSON.stringify(finalOutput, null, 2),
);
javascript
// src/utils/log.js
const outputFilename = path.join(
  path.dirname(__filename),
  "..",
  "..",
  "logs",
  safeFilename,
);

try {
  await fs.promises.mkdir(path.dirname(outputFilename), { recursive: true });
  await fs.promises.writeFile(outputFilename, content);
  utils.printSuccess(`  → 已保存到 ${outputFilename}`);

Equivalent automatic result writes occur in:

  • src/douyin/post-cli.js, lines 152-162
  • src/douyin/search-cli.js, lines 253-261

The comment response schema confirms that results can contain comment text, stable user identifiers, nicknames, profile identifiers, and IP-region labels.

Technical Analysis

Every successful comment, post, or search operation automatically writes the complete structured response to the project-level logs directory. There is no command-line opt-out, retention period, data minimization, encryption, or explicit restrictive file mode.

fs.promises.writeFile and mkdir rely on the process umask and ambient filesystem configuration. On systems with permissive defaults, other local users, container processes, backup agents, synchronization software, or later application components may be able to read the files.

The filename sanitizer prevent ...[truncated 1935 chars]

Remediation
View remediation

Remediation Suggestions

  1. Make persistence opt-in rather than automatic, for example through an explicit --output or --save option.

  2. Clearly disclose before execution that saving results will create a persistent local copy.

  3. Create directories and files with owner-only permissions:

    javascript
    await fs.promises.mkdir(logDirectory, {
      recursive: true,
      mode: 0o700,
    });
    
    await fs.promises.writeFile(outputFilename, content, {
      encoding: "utf8",
      mode: 0o600,
      flag: "wx",
    });
    
  4. Provide --no-save, retention-period, and secure cleanup controls.

  5. Minimize stored data by default. Exclude stable user identifiers, nicknames, profile URLs, and IP-region labels unless explicitly requested.

  6. Support pseudonymization or redaction when results are retained for analysis.

  7. Consider authenticated encryption for result files containing identifying fields.

  8. Prevent logs from entering source control by adding logs/ to .gitignore.

  9. Warn users not to place the project directory in shared or automatically synchronized locations.

  10. Document the storage location, permissions, retention behavior, deletion procedure, and privacy implications in SKILL.md.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (44)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The presence of token-validation and promotional messaging for obtaining a GUAIKEI token, without corresponding implemented Douyin features, raises concern that the skill may solicit credentials or drive users to a third-party service under misleading pretenses. In context, this is more dangerous because the skill explicitly requests an API token and presents itself as a data-access tool, increasing the chance of unnecessary secret exposure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The presence of token-validation and promotional messaging for obtaining a GUAIKEI token, without corresponding implemented Douyin features, raises concern that the skill may solicit credentials or drive users to a third-party service under misleading pretenses. In context, this is more dangerous because the skill explicitly requests an API token and presents itself as a data-access tool, increasing the chance of unnecessary secret exposure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The presence of token-validation and promotional messaging for obtaining a GUAIKEI token, without corresponding implemented Douyin features, raises concern that the skill may solicit credentials or drive users to a third-party service under misleading pretenses. In context, this is more dangerous because the skill explicitly requests an API token and presents itself as a data-access tool, increasing the chance of unnecessary secret exposure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The presence of token-validation and promotional messaging for obtaining a GUAIKEI token, without corresponding implemented Douyin features, raises concern that the skill may solicit credentials or drive users to a third-party service under misleading pretenses. In context, this is more dangerous because the skill explicitly requests an API token and presents itself as a data-access tool, increasing the chance of unnecessary secret exposure.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 72)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

L004 将触发条件表述为“当用户要在抖音找蓝海品、避开大词竞争、从细分场景选品、看抖音热搜、按关键词看视频热度时必须触发”,其中多项是高层业务意图而非明确指令,边界较宽,容易与一般性的内容策划或营销讨论重叠。虽然文件后文提供了边界说明,但这一顶层 manifest 描述本身仍可能让路由器在缺少明确抖音数据查询意图时误调用技能。

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The package metadata uses a very broad description and generic trigger terms such as 抖音、关键词、热搜、选品 and English analytics/search phrases, which can cause the skill to be invoked for loosely related requests rather than only for narrowly scoped Douyin data collection tasks. In an agent environment, overbroad activation increases the chance of unintended tool use, exposing user queries to an external capability and producing irrelevant or policy-bypassing data access behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README encourages fetching public videos, author data, and especially comments, but does not clearly warn that collected content and exported logs may contain personal or sensitive information. In a data-collection skill focused on analytics and monitoring, this omission increases the risk of users collecting, storing, or reusing personal data without appropriate handling or consent awareness.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file presents all user-facing natural-language content in Chinese and does not indicate that the language is optional or user-selectable. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code persists the fetched comment data to disk via log.taskWrite, which can affect user data handling and leave local artifacts. Although the operation is visible in code, there is no confirmation prompt, user-facing warning, or explanatory comment/docstring near the write indicating that results will be saved to a file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JavaScript file contains multiple user-facing strings entirely in Chinese, including status messages and errors, but provides no indication that the skill is region-specific or that users can opt into a language/locale. That creates a natural-language policy concern because the skill effectively forces a specific language for all users.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The command writes the full result payload to a JSON file via log.taskWrite, which may persist user-requested data locally. In this file there is no explicit warning, confirmation prompt, or prior disclosure that execution will save output to disk rather than only print it.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The code reads GUAIKEI_API_TOKEN from the environment and performs remote search task creation and retrieval using the user's keyword and filters. Although remote search is part of the skill's purpose, there is no explicit user-facing disclosure that their query parameters will be sent to an external service authenticated with an API token.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code emits user-facing error messages in Chinese only (for example at L07 and L11), which imposes a specific language on users without any opt-in or fallback. This is a natural-language policy concern because the file contains no indication that the skill is region-specific or that users can choose their preferred language.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

This code performs outbound HTTPS GET and POST requests, including URL parameters and JSON request bodies, but only includes internal validation and retry logging. There is no confirmation prompt, user-facing notice, or warning comment/docstring explaining that user or system data may be sent to a remote service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The user-facing strings in this file are entirely in Chinese and are shown during token validation failures, but there is no indication that the skill is China-specific or that users can opt into this language. That creates a natural-language policy concern because the skill imposes a specific locale on users without documented choice or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file's user-facing error strings and formatted output are entirely in Chinese, and the comments indicate the skill is designed around Chinese-language keywords and results. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly justified, which is not present here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The schema title and field descriptions are written entirely in Chinese, which can impose a language expectation on users or downstream tooling without any stated opt-in or justification. Under the policy rule for natural-language violations, this is a locale/language constraint that is not documented as optional or region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The schema title and description are written only in Chinese, and the field descriptions continue that pattern throughout the file. Because this is natural-language content in a globally consumable schema file, it imposes a specific language without any visible opt-in or justification for a China-specific audience.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

This JSON schema contains natural-language titles and descriptions exclusively in Chinese, such as the title and multiple property descriptions. Because the file does not document that the skill is China-specific or offer any language choice, it may violate a language/locale policy requiring user opt-in or justification for a forced locale.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This manifest-style JSON schema contains user-facing title and description fields entirely in Chinese, including parameter descriptions. Because the file provides no indication that the skill is region-specific or that users can choose another language, it appears to impose a language/locale choice without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This JSON schema uses Chinese-only natural-language metadata in the title and description, but the file does not document that the skill is region-specific or that users can opt into this locale. Under the policy, forcing a specific language without user choice or a justified locale constraint is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This manifest-style JSON schema contains user-facing title and description fields exclusively in Chinese, which can amount to a language/locale restriction without explicit opt-in. The file does not indicate that the skill is region-specific or provide an alternative language option.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This JSON schema uses Chinese titles and descriptions throughout, which effectively forces a specific language for users or downstream tooling consuming these natural-language fields. The file does not indicate that the skill is region-specific or provide any user opt-in or multilingual alternative, so it may violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:25