Back to skill

Security audit

抖音爆款爬虫

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed Douyin public-data scraper that uses a third-party API and local JSON logs, with no evidence of hidden execution, credential theft, or destructive behavior.

Install only if you are comfortable sending Douyin search terms, public URLs or IDs, and your GUAIKEI_API_TOKEN to guaikei.com. Review and delete the generated local logs when they contain comments, nicknames, account IDs, or business research, and avoid using the tool for private data or public redistribution without authorization.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (45)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

从该代码块看,实际功能局限于搜索模块:向 /api/douyin/general-search/keyword 创建关键词搜索任务,向 /api/douyin/general-search/info 查询搜索结果,并对结果数组做简单加工。它支持声明中的一部分能力,即“关键词搜索视频/图文,可按点赞数、发布时间、视频时长、内容类型筛选排序”。但声明还包含热榜查询、博主作品抓取、评论分析、自然语言搜索请求支持、以及更广义的视频信息提取,这些在当前代码中均未体现。因此就该代码块与整体声明对比,描述范围明显大于实际实现,存在不匹配。未发现额外的未声明敏感能力;问题主要是声明的能力超出了该代码实际行为。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

代码的主函数和帮助信息都明确表明其用途是“获取抖音博主作品列表”。参数只有 url/sec_uid 和 limit,调用的是 post.createPostTask 与 post.getPostTask,围绕公开作品列表抓取展开;没有任何与关键词搜索、自然语言解析、热榜接口、评论抓取或评论互动数据处理相关的参数、调用或输出结构。因此,这个代码片段并不能支撑声明中的四大能力,只匹配其中第(3)项“博主作品抓取”。虽然不存在明显的额外越权或未声明资源访问,但声明对该代码片段的功能描述显著过宽,属于描述与实际行为不一致。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是一个面向抖音数据采集与分析的技能,而给出的代码片段仅是通用参数解析模块(parseArgs、readValueAfterFlag、buildHelp)。它处理命令行参数校验、布尔/字符串选项、重复参数检查、位置参数和帮助信息生成。该代码既不访问抖音,也不执行网络请求、数据抓取、搜索、热榜读取、用户作品提取或评论分析。因此这不是单纯的底层支撑细节与描述一致,而是代码片段的实际功能与声明的核心用途明显不符,构成显著不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个面向抖音数据抓取与分析的技能,而提供的代码片段只是一个本地日志写入模块。它使用 fs/path 在本地 logs 目录创建并写入文件,并进行文件名清洗与错误提示。虽然日志记录可视为配套实现细节,但该片段本身没有体现任何与抖音搜索、热榜、博主作品抓取、评论分析相关的行为。基于“描述是否准确代表该代码块实际作用”这一标准,这里存在明显不一致:代码实际功能是文件日志写入,而非声明中的核心业务能力。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a network-enabled Douyin data collection and analysis skill with multiple end-user capabilities. However, the supplied code chunk contains only a small local utility that accesses the filesystem to read package.json and return its name. It does not implement any Douyin-related logic, search, scraping, API access, comment analysis, or trend retrieval. This is a clear description-behavior mismatch because the actual code shown is unrelated to the declared primary purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

声明描述的是一个抖音数据抓取与分析技能,核心能力应涉及网络请求、抖音内容检索、热榜获取、用户作品列表抓取或评论提取。但给出的代码片段只处理 GUAIKEI_API_TOKEN 的合法性校验,并在 token 缺失/无效时打印警告和营销信息。它没有展示任何与抖音数据访问、搜索、抓取、解析或评论分析相关的实现。因此,就该代码片段本身而言,其实际行为与声明用途存在明显不一致。尽管 token 管理可能是完整技能的辅助模块,但此片段单独体现的功能与所宣称的业务能力 materially different,且还包含未声明的推广性输出行为。

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 72)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill declares use of an environment variable (GUAIKEI_API_TOKEN) but does not define an explicit tool scope such as permissions or allowed-tools. That weakens least-privilege controls and makes the runtime’s secret and capability boundaries less explicit, which can increase the chance of accidental overexposure or misuse in broader agent environments.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The schema title and description are written entirely in Chinese, which imposes a specific language/locale in the skill's natural-language interface metadata. There is no indication here that users can choose another language or that the Chinese-only locale is intentionally justified as region-specific.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The description includes a natural-language trigger example ('搜索一下AI视频') that is broad enough to resemble an ordinary user request rather than a clearly scoped invocation. In agent environments, this can cause accidental or ambiguous activation, especially since the skill is designed to perform scraping/search actions against external content sources based on user text.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The package description is entirely in Chinese and the only natural-language invocation example is also Chinese, which suggests a language-specific interaction model. There is no indication that users may choose another language or that the Chinese-only constraint is intentional and justified for a region-specific tool.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README states that outputs are automatically saved to a local logs directory and gives filename patterns for search, post, and comment tasks, but it does not prominently warn users up front that scraped public data and comments will be exported to disk. Because this skill collects potentially sensitive public-content datasets at scale, silent local persistence can lead to unintentional retention, sharing, or exposure of collected comments, account identifiers, and activity data on shared machines or in synced workspaces.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This document operationalizes large-scale Douyin scraping, including search, creator post collection, comments, and hotlist retrieval, with limits up to 10,000 records and no warning about privacy, consent, platform terms, or lawful handling of collected data. In the context of a scraping skill, that omission materially increases the risk of misuse for mass data harvesting, profiling, or collection of user-generated content without appropriate safeguards.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This code performs a network request that transmits user-supplied data (url) along with an authentication token to an external API, which is a safety-relevant operation for code files. In this file there is no visible confirmation prompt, user-facing disclosure, or warning comment indicating that the data will be sent off-box.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This function issues a network request containing the user's video URL and authentication token to retrieve task results. The file includes no visible disclosure, prompt, or warning that this network transmission occurs, so users may not realize their input and credentials are being sent to an external service.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The code reads process.env.GUAIKEI_API_TOKEN and uses it to authenticate external comment-retrieval requests, but this file does not include a user-facing prompt, warning, or explanatory comment about accessing credentials from the environment. For code files, sensitive environment variable access should be disclosed unless the warning is clearly provided elsewhere or is unmistakably inherent to the skill's stated purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The tool persists full comment results to a local JSON file without explicit user consent or clear disclosure, which can expose scraped comment content and interaction data to other local users, backups, or later unintended reuse. In this skill context, the data being collected is user-generated content at scale, so silent retention meaningfully increases privacy and data-handling risk beyond the expected one-shot CLI output.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code uses Chinese-only user-facing text in comments, errors, and success messages, such as at L21, L45, L49, L71, and L82. Because the skill does not provide any user opt-in or configuration for language selection, it violates the language/locale policy criterion for natural-language policy issues.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

User-facing descriptions and runtime messages throughout the file are presented only in Chinese, starting with the help text definitions. Under the policy, forcing a specific language without opt-in is a natural-language policy violation unless the locale restriction is explicitly documented and justified, which is not evident in this file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The CLI persists the full search output to a local JSON file whose filename also embeds the user's keyword. This can expose potentially sensitive search terms and returned content to other local users, backup systems, or later unintended disclosure, especially because the tool does not clearly warn users or offer an opt-in/opt-out control for persistence.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Multiple error messages in this file are hard-coded in Chinese, including messages surfaced from API and network/auth failures. This enforces a specific language for user-visible output without offering a language choice or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The user-facing messages in this file are entirely in Chinese, including warnings and status output, with no indication that language is configurable or user-selected. This creates a natural-language policy concern because it forces a specific language/locale without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JavaScript file contains natural-language comments and user-facing error/output strings exclusively in Chinese, including validation errors and formatted result messages. The file does not indicate that the skill is region-specific or provide any user opt-in for language selection, which can violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

该 markdown 文件适用 SQP-1。L214 使用“帮我搜抖音里…/抖音今天有什么热点/分析这条抖音视频评论区…”作为触发示例,其中“分析…”这类自然语言表达偏口语化,若仅依赖示例而非明确触发词清单,可能与日常请求重叠,增加误调用风险。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

This manifest file contains user-facing natural-language fields entirely in Chinese, including the title and parameter descriptions. The file does not offer any language/locale choice or explain that the skill is intentionally region-specific, which can violate a language/locale policy requiring user opt-in or documented justification.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:16