Back to skill

Security audit

抖音每日点赞飙升榜 抖音自定义垂类视频点赞榜 抖音自定义关键词点赞榜

Security checks for vulnerabilities and agentic risk

Overview

This Douyin analytics skill is mostly coherent and not malicious, but it automatically saves full scraped/search results to local JSON files without clear opt-out or retention controls.

Review the default logging behavior before installing. Use this only in a workspace where saved Douyin results, search terms, creator IDs, comments, nicknames, and IP-region labels can be retained safely; clean artifact/logs when no longer needed and protect GUAIKEI_API_TOKEN as a third-party API credential.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
src/douyin/search-cli.js:253
Finding

Automatic Plaintext Retention of Query Results and Public Personal Data

Content
View full analysis

Vulnerability Details

File Location: src/douyin/search-cli.js:253-260, src/douyin/comment-cli.js:151-162, and src/utils/log.js:25-35
Vulnerability Type: Automatic plaintext storage of potentially sensitive query and response data
Risk Level: Medium

Vulnerable Code

src/douyin/search-cli.js:253-260:

js
console.log(JSON.stringify(finalOutput, null, 2));
utils.printSuccess(
  `搜索任务完成, 共返回 ${finalOutput.results.length} 条结果`,
);

await log.taskWrite(
  `${startTime}_${keyword}_${sort}_${time}_${duration}_${content}_search.json`,
  JSON.stringify(finalOutput, null, 2),
);

src/douyin/comment-cli.js:151-162:

js
console.log(JSON.stringify(finalOutput, null, 2));
utils.printSuccess(
  `获取评论任务完成, 共返回 ${finalOutput.results.length} 条结果`,
);

url = url.replace(/[^a-zA-Z0-9_-]/g, "");
url = url.replace("httpswwwdouyincomvideo", "");
url = url.replace("httpswwwdouyincomnote", "");
await log.taskWrite(
  `${startTime}_${url}_comment.json`,
  JSON.stringify(finalOutput, null, 2),
);

src/utils/log.js:25-35:

js
const outputFilename = path.join(
  path.dirname(__filename),
  "..",
  "..",
  "logs",
  safeFilename,
);

try {
  await fs.promises.mkdir(path.dirname(outputFilename), { recursive: true });
  await fs.promises.writeFile(outputFilename, content);

Technical Analysis

Successful search and comment operations are automatically serialized and written to the project’s logs/ directory. Persistence is not opt-in, no retention policy is enforced, and no redaction is applied before writing the complete response.

Comment response records may contain comment text, user IDs, nicknames, stable SEC_UID identifiers, and IP-region labels, as documented in assets/comment_cli_resp.schema.json:23-41. Search logs also retain the user’s keyword, filtering parameters, and complete API results. The files are created using the process and ...[truncated 2091 chars]

Remediation
View remediation

Remediation Suggestions

  1. Disable persistence by default. Return results through standard output without creating files unless the user explicitly requests storage.
  2. Add an explicit output option. Require a flag such as --save or --output <path> before writing any result data.
  3. Apply data minimization. Exclude fields that are unnecessary for the requested analysis, particularly stable user identifiers and IP-region labels.
  4. Provide redaction controls. Support an anonymized output mode that replaces user IDs and nicknames with non-reversible aliases and removes location-related fields.
  5. Use restrictive file permissions. Create the log directory with mode 0700 and output files with mode 0600, subject to platform support.
  6. Implement retention controls. Add configurable expiration, maximum file count, and secure cleanup functionality.
  7. Document local storage behavior. Clearly notify users before saving data and identify the storage location, retained fields, and deletion procedure.
  8. Separate operational logs from datasets. Operational logs should contain only status and diagnostic metadata, while full API responses should be treated as explicitly exported datasets.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (52)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

If the actual code is mainly generic logging/terminal tooling rather than Douyin data retrieval, the skill is misclassified and can be incorrectly selected by trigger-based routing. Misrouting is especially risky here because the skill requests an API token and may handle public-comment or profile data through a third-party service, amplifying consequences of misplaced trust.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the actual code is mainly generic logging/terminal tooling rather than Douyin data retrieval, the skill is misclassified and can be incorrectly selected by trigger-based routing. Misrouting is especially risky here because the skill requests an API token and may handle public-comment or profile data through a third-party service, amplifying consequences of misplaced trust.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

If the actual code is mainly generic logging/terminal tooling rather than Douyin data retrieval, the skill is misclassified and can be incorrectly selected by trigger-based routing. Misrouting is especially risky here because the skill requests an API token and may handle public-comment or profile data through a third-party service, amplifying consequences of misplaced trust.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the actual code is mainly generic logging/terminal tooling rather than Douyin data retrieval, the skill is misclassified and can be incorrectly selected by trigger-based routing. Misrouting is especially risky here because the skill requests an API token and may handle public-comment or profile data through a third-party service, amplifying consequences of misplaced trust.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the actual code is mainly generic logging/terminal tooling rather than Douyin data retrieval, the skill is misclassified and can be incorrectly selected by trigger-based routing. Misrouting is especially risky here because the skill requests an API token and may handle public-comment or profile data through a third-party service, amplifying consequences of misplaced trust.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the actual code is mainly generic logging/terminal tooling rather than Douyin data retrieval, the skill is misclassified and can be incorrectly selected by trigger-based routing. Misrouting is especially risky here because the skill requests an API token and may handle public-comment or profile data through a third-party service, amplifying consequences of misplaced trust.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the actual code is mainly generic logging/terminal tooling rather than Douyin data retrieval, the skill is misclassified and can be incorrectly selected by trigger-based routing. Misrouting is especially risky here because the skill requests an API token and may handle public-comment or profile data through a third-party service, amplifying consequences of misplaced trust.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

If the actual code is mainly generic logging/terminal tooling rather than Douyin data retrieval, the skill is misclassified and can be incorrectly selected by trigger-based routing. Misrouting is especially risky here because the skill requests an API token and may handle public-comment or profile data through a third-party service, amplifying consequences of misplaced trust.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 72)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
85% confidence
Finding

The skill declares access to an environment variable (GUAIKEI_API_TOKEN) but does not define an explicit tool/permission scope such as allowed tools or permission boundaries. This weakens least-privilege guarantees and makes it harder for a host to reason about what sensitive inputs the skill is permitted to access, especially when combined with third-party API usage.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrases include broad ranking-related terms such as generic '榜' or '排名' concepts, which can cause the skill to be selected for unrelated requests. In an agent environment, over-broad triggers are dangerous because they increase unintended invocation, potentially exposing user prompts, links, or identifiers to an external service without a sufficiently specific match.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

User-facing title and descriptions are written entirely in Chinese, which may impose a specific language experience without any opt-in or documented locale limitation. The file does not indicate that the skill is region-specific or that users can choose another language, so this appears to violate the language/locale policy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JSON schema uses Chinese-only natural-language titles and descriptions throughout, which imposes a specific language/locale with no indication that the skill is region-specific or that users can opt into another language. Under the stated policy, forcing a specific language without user choice is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest markets the skill as a narrowly scoped 'daily like-surge leaderboard' tool, but the exposed scripts clearly enable broader capabilities including hot-search retrieval, post scraping, and comment analysis. This scope mismatch is dangerous because it can cause users, reviewers, or orchestration systems to grant trust and permissions under a narrower understanding than the code actually supports, increasing the chance of unintended data collection or misuse.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrases are broad and generic, such as terms equivalent to 'ranking' and 'surge list', which are likely to match common user requests beyond the intended narrow use case. In an agent environment, overbroad triggers can cause the skill to activate unexpectedly and handle requests involving scraping, hot-search data, or comment analysis when the user did not clearly request those actions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The README explicitly advertises automatic JSON log export of scraped videos, comments, and analysis results to a local logs directory, but it does not clearly warn users that potentially sensitive third-party content and derived analysis data will be persisted on disk. In this skill’s context, the tool collects large volumes of public social-media data for monitoring and analysis, so silent local retention increases the risk of unintended disclosure, over-retention, or mishandling of scraped content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The README instructs users to obtain and set GUAIKEI_API_TOKEN, but it does not clearly state that the skill will use this credential for outbound requests to a third-party service to retrieve Douyin data. Because the skill is built around remote data collection and includes support/contact channels for token issuance, insufficient disclosure can lead users to supply credentials without fully understanding their use, scope, or exposure risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This markdown file contains user-facing natural-language content exclusively in Chinese, including the title and all changelog entries. Under the policy rule, forcing a specific language without user opt-in or clear justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The CLI writes full comment results to a local JSON file by default, and those results can contain user-generated content, usernames, IDs, and other potentially sensitive or regulated data. Silent persistence increases the risk of unintended retention, exposure to other local users/processes, backup systems, or later exfiltration, especially in automation environments.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This code contains user-facing strings such as error messages and status text entirely in Chinese, and there is no indication that the skill is restricted to a Chinese-speaking audience or that users can opt into another language. Under the language/locale policy rule, forcing a specific language without opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest name and description emphasize querying a 抖音每日点赞飙升榜 / TOP1000 surge ranking with filters, but this file is a CLI for fetching a specific blogger's public works list from a homepage URL or sec_uid. Although the broader manifest later mentions '博主作品抓取' as a supported capability, this file's behavior does not align with the headline purpose of a daily surge ranking query tool and instead performs a distinct post-list retrieval workflow.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The CLI writes the full scraped results to a local JSON file automatically, without an explicit opt-in, warning, or controls around where sensitive scraped data is stored. In a multi-user or shared runtime, this can leave behind persistent data artifacts containing URLs, post metadata, and potentially other collected content that may be accessed later by unintended parties.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill's descriptions, help text, validation errors, and status messages are all presented in Chinese, and there is no indication that the language is configurable or intentionally restricted to a China-specific audience. The policy flags language/locale constraints when a skill forces a specific language without user opt-in or documented justification.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest describes a Douyin ranking/query tool with search, hotlist, creator works, and comment analysis capabilities, but does not mention credential handling or environment-based secret access. This file explicitly reads process.env.GUAIKEI_API_TOKEN to obtain an API credential, which is a privileged capability beyond the user-facing query behavior described.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The CLI writes full query output to a local file named with the search keyword, which creates undisclosed local persistence of user activity and retrieved content. This can expose sensitive search terms or collected data to other local users, backup systems, or downstream tooling, especially because neither consent nor retention controls are visible here.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:16