Back to skill

Security audit

抖音内容引擎

Security checks for vulnerabilities and agentic risk

Overview

This Douyin public-data skill is coherent and disclosed, with the main risk being that it sends queries or links to a third-party API and saves some results as local plaintext logs.

Install only if you are comfortable sending Douyin keywords or links and your Guaikei token to www.guaikei.com. Treat the generated logs directory as sensitive because it can contain research keywords, account/video identifiers, comments, and public user metadata; delete or protect those files if the workspace is shared or backed up.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
src/utils/log.js:25
Finding

Automatic Plaintext Retention of Queries and Collected Douyin Data

Content
View full analysis

Vulnerability Details

File Location: src/utils/log.js:25-35; invoked from src/douyin/search-cli.js:280-283, src/douyin/comment-cli.js:198-201, and src/douyin/post-cli.js:209-212
Vulnerability Type: Plaintext storage of potentially sensitive query and collected user data
Risk Level: Medium

Vulnerable Code

src/utils/log.js:25-35:

js
const outputFilename = path.join(
  path.dirname(__filename),
  "..",
  "..",
  "logs",
  safeFilename,
);

try {
  await fs.promises.mkdir(path.dirname(outputFilename), { recursive: true });
  await fs.promises.writeFile(outputFilename, content);

src/douyin/search-cli.js:280-283:

js
await log.taskWrite(
  `${startTime}_${keyword}_${sort}_${time}_${duration}_${content}_search.json`,
  JSON.stringify(finalOutput, null, 2),
);

src/douyin/comment-cli.js:198-201:

js
await log.taskWrite(
  `${startTime}_${url}_comment.json`,
  JSON.stringify(finalOutput, null, 2),
);

src/douyin/post-cli.js:209-212:

js
await log.taskWrite(
  `${startTime}_${url}_post.json`,
  JSON.stringify(finalOutput, null, 2),
);

Technical Analysis

Every successful search, comment retrieval, and profile-post retrieval operation automatically serializes the complete output object and writes it to the project’s logs directory. The stored object includes the original request parameters and all returned records. Depending on the command, this can retain search keywords, profile or video identifiers, comments, public user identifiers, and other collected content.

Although the README discloses that results are archived, the implementation does not provide an opt-out, expiration policy, cleanup mechanism, encryption, or explicit restrictive filesystem permissions. mkdir and writeFile inherit permissions from the process umask. In a shared workspace, permissive umask configuration can therefore make these files acce ...[truncated 1717 chars]

Remediation
View remediation

Remediation Suggestions

  1. Make local persistence opt-in through an explicit option such as --output or --save.
  2. Provide a --no-log option if backward compatibility requires logging to remain enabled by default.
  3. Create the log directory with mode 0700 and result files with mode 0600, while documenting platform-specific permission limitations.
  4. Avoid placing raw search keywords in filenames; use a random identifier or cryptographic hash instead.
  5. Implement configurable retention and automatic cleanup, such as deleting files after a defined number of days.
  6. Clearly warn users before saving large comment or profile datasets and describe exactly which fields are retained.
  7. Consider data minimization by omitting unnecessary request metadata and user identifiers from persisted output.
  8. For environments requiring stronger confidentiality, support encrypted storage or instruct users to save results only in an access-controlled directory.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (40)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

代码片段仅包含通用参数解析与帮助文本生成功能(parseArgs、readValueAfterFlag、buildHelp),属于底层 CLI 工具组件。它不访问抖音、网络、API、页面或任何内容数据源,也没有实现搜索、抓取、评论读取、热榜查询、内容分析、舆情监控等声明中的核心业务能力。虽然这类参数解析器可能作为更大系统的辅助模块存在,但就所给代码片段本身而言,其实际行为与“抖音内容引擎”描述的主要目的存在明显且实质性的偏差。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

该代码片段的实际作用仅是将字符串内容安全写入本地日志文件,属于基础日志/持久化辅助功能。声明描述的则是一个面向抖音内容分析的数据抓取与分析引擎,涉及搜索、抓取评论、博主主页作品、热榜等能力。两者在主要目的和能力上明显不一致。虽然日志写入可作为支持性实现细节,但当前提供的代码片段本身完全没有体现任何与抖音平台交互、数据抓取、分析或 JSON 输出相关的行为,因此就这段代码与声明的对应关系而言,应判定为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a full-featured Douyin data collection and analysis skill. However, the provided code chunk does not implement any Douyin-facing functionality, network access, scraping, ranking retrieval, comment fetching, or structured content analysis. It merely accesses a local filesystem resource (package.json) and returns the package name. This is materially different from the declared primary purpose, so the description does not accurately represent the behavior of the supplied code chunk.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 77)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The trigger list includes broad phrases such as '舆情监控', '热点追踪', and '爆款挖掘' that can match user requests not specific to Douyin. In an agent environment, overly broad routing can cause unintended transmission of user queries or links to a third-party API, creating data leakage and consent/scope issues even if the underlying API use is documented.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This JSON schema contains natural-language titles and descriptions exclusively in Chinese, which imposes a specific language/locale on consumers without any visible opt-in or justification. The policy for this review flags language-forcing behavior across all file types when no language choice or documented regional scope is provided.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The package description contains broad trigger phrases such as '获取抖音作品', '查看抖音视频', '抖音内容采集', and '抓取抖音作品时使用', which can cause the skill to activate for loosely related user requests. Over-broad activation increases the chance of unintended invocation, causing unnecessary access to scraping or monitoring capabilities and potentially exposing users to data collection they did not explicitly request.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This file contains natural-language comments and user-visible retry/error text exclusively in Chinese, and there is no indication that the skill offers a language choice or that the locale restriction is documented as intentional. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

This code file contains natural-language comments and parameter descriptions exclusively in Chinese, which can amount to a language-policy issue when no opt-in or documented locale restriction is provided. The rule specifically flags language or locale constraints that are imposed without user choice or justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The CLI persists fetched comment results to a local JSON file automatically, without explicit user consent, visibility, or a way to disable storage. Because comments and related metadata may contain sensitive or regulated content, this creates an unintended local data-retention risk: data can remain on disk, be collected by other local users/processes, or leak via backups and artifact uploads.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code contains multiple user-facing strings in Chinese, including error messages and completion notices, but provides no user opt-in or alternative locale. That creates a natural-language policy concern if the skill is expected to be generally usable rather than explicitly region- or language-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This code creates and fetches a remote post task via external API calls, using the normalized URL and an authentication token. Although the CLI prints progress messages, it does not clearly disclose that user-supplied data is being sent to a third-party service, which matches the missing-warning criterion for network transmission in code files.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code persists the final output to disk through log.taskWrite, which is a file-write operation covered by the warning requirement. The current messages describe task success but do not clearly tell the user that a local file will be created, and no inline comment/docstring near the write explains this side effect.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file contains user-facing error strings in Chinese only, which imposes a specific language on users without any opt-in or fallback. Under the policy, language forcing is a natural-language violation unless the skill offers a language choice or clearly documents a justified locale constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The success and failure messages are emitted only in Chinese, which means the skill enforces a language choice at runtime. There is no indication in this file of user opt-in, multilingual support, or a documented region-specific reason for the restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file contains multiple user-visible error messages in Chinese, including authentication guidance and request failures, without any indication that the skill is China-specific or that users can opt into this locale. This creates a natural-language policy issue because the skill appears to enforce a specific language regardless of user preference.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The file's user-facing and descriptive text is entirely in Chinese, indicating the skill is designed around a fixed language locale. Under the policy, forcing a specific language without user opt-in or documented justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The error and informational messages shown to users are hard-coded in Chinese, and the file provides no indication that users can choose another language or that Chinese-only behavior is required. This violates the language/locale policy for all file types when a specific language is imposed without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JavaScript file contains multiple user-facing error strings in Chinese, such as the messages printed during keyword validation and option validation. Because the skill forces a specific language in its outputs without any visible opt-in or justification that it is a China-specific or Chinese-only tool, it matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
70% confidence
Finding

The manifest description, examples, and user-facing guidance are entirely Chinese-centric and targeted to Douyin use cases, but there is no explicit opt-in or statement that responses can adapt to the user's preferred language. This can conflict with language/locale policy expectations when a skill effectively forces one language without user choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The natural-language strings in the schema consistently force a specific language for user-facing metadata. Under the language/locale policy, this should either provide user opt-in or clearly document that the skill is intentionally limited to a Chinese-language or region-specific context.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This is a manifest-style JSON file, so vague-trigger checks apply. The title and descriptions explain what parameters are accepted, but they do not define specific activation phrases, scope limits, or exclusion conditions, which can make invocation criteria overly broad or unclear.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The schema title and description are written only in Chinese, which imposes a specific language in user-facing metadata without any indication of opt-in, alternative locale support, or region-specific justification. This matches the policy category for language/locale constraints expressed in natural-language content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The schema's human-readable fields such as title and multiple description values are written only in Chinese, which imposes a specific language on consumers of the skill metadata. The file does not indicate that the skill is region-specific or provide any user opt-in or alternative locale, which matches the language/locale policy concern.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:16