Back to skill

Security audit

抖音搜索爬虫

Security checks for vulnerabilities and agentic risk

Overview

The skill appears to perform the advertised Douyin data retrieval, but it needs review because it handles an API token insecurely and automatically saves scraped result datasets locally.

Install only if you are comfortable sending Douyin keywords, profile or video URLs, and a GUAIKEI API token to guaikei.com, and with retrieved public Douyin results being saved locally by default. Prefer a low-privilege or revocable token, avoid running it on shared or CI machines, review and delete the logs directory regularly, and use the data only where Douyin terms and privacy law allow it.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
src/utils/request.js:99
Finding

API Token Exposed Through URL Query Strings

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
src/douyin/search-cli.js:258
Finding

Automatic Plaintext Retention of Complete Result Datasets

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (45)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The reported implementation focuses on token handling and promotional guidance for obtaining a private token rather than the advertised analysis features. This is risky because it encourages users to provision secrets to a skill whose real behavior is unclear, increasing exposure of credentials and enabling deceptive or unnecessary secret collection.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The reported implementation focuses on token handling and promotional guidance for obtaining a private token rather than the advertised analysis features. This is risky because it encourages users to provision secrets to a skill whose real behavior is unclear, increasing exposure of credentials and enabling deceptive or unnecessary secret collection.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The reported implementation focuses on token handling and promotional guidance for obtaining a private token rather than the advertised analysis features. This is risky because it encourages users to provision secrets to a skill whose real behavior is unclear, increasing exposure of credentials and enabling deceptive or unnecessary secret collection.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The reported implementation focuses on token handling and promotional guidance for obtaining a private token rather than the advertised analysis features. This is risky because it encourages users to provision secrets to a skill whose real behavior is unclear, increasing exposure of credentials and enabling deceptive or unnecessary secret collection.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/changelog.md (reported line 72)May include surrounding context.

md
- 技能重命名为“douyin-search-keyword”。
- 在SKILL.md中添加了openclaw元数据、使用帮助、许可证、标签和示例,以实现更好的集成与文档化。
- 移除了两个本地文件(.env 和 scripts/last-search.json),以优化代码结构并提升安全性。
- 文档现已更加简洁且以用户为中心,重点在于提供清晰的使用说明和数据字段解释。
- 突出技能特性、合规要点及技术流程。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The trigger list mixes specific phrases like '抖音搜索' with broader phrases such as '短视频选题', '抖音舆情', and '抖音数据分析', which can overlap with common strategy, writing, or analysis requests that do not necessarily require this scraping skill. Although nearby text narrows scope somewhat, the trigger list itself lacks explicit exclusions for these broad phrases and could cause unintended invocation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The README explicitly states that scraped results are automatically saved as JSON logs, including video, author, and comment data, but it does not clearly warn users that potentially sensitive or regulated third-party content will persist on local disk. This creates a real data-handling risk: users may unknowingly retain large volumes of scraped content and comments, increasing exposure to privacy, compliance, and accidental disclosure issues if logs are shared, synced, or left unsecured.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The file documents large-scale scraping of Douyin search results, creator posts, hot topics, and comments, including up to 10,000 records, but provides no guidance on privacy, lawful use, rate limiting, consent boundaries, or platform terms. In an agent skill context, this omission increases the chance the tool will be used for unauthorized collection, profiling, or policy-violating surveillance of user-generated content and creator activity.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The CLI writes fetched comment results to a local JSON file automatically after execution, but this side effect is not clearly disclosed in the skill description or in runtime prompts. Because comment data may contain personal identifiers, opinions, or sensitive content, silent persistence increases the risk of unintended retention, local data exposure, and mishandling on shared systems.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Comment results are saved locally without any explicit warning, confirmation, or consent flow, which creates a transparency and privacy problem. In the context of a scraping/analysis skill that processes social-media comments at scale, undisclosed storage makes accidental collection and retention more dangerous, especially on multi-user machines or automation hosts.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The code reads a credential from GUAIKEI_API_TOKEN and sends the normalized URL and limit to createPostTask/getPostTask, which implies transmission of user-provided data to an external service. This file does not clearly warn the user that invoking the CLI will send that data off-host.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The CLI writes the full scraped results to a local JSON file automatically, without any opt-in, path disclosure, or redaction. Because the data may include comments, author identifiers, URLs, and analysis output, this creates an unintended local data retention channel that can expose collected data to other local users, backups, log collectors, or later compromise of the host.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The CLI persists the full search output to a local JSON file automatically after execution, but there is no explicit user consent, warning, or visibility into where the file is stored. In this skill’s context, search queries and returned Douyin content may contain sensitive research targets, monitoring subjects, or business-intelligence data, so silent local persistence can create unintended data exposure on shared machines, CI runners, or multi-user environments.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code contains hard-coded Chinese user-facing error strings, and the success/error output later in the file continues the same pattern. For a general utility module, forcing a specific language without offering user choice or documenting a locale-specific scope is a natural-language policy violation under the language/locale rule.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code returns user-facing error text in Chinese, such as "请求失败", without offering a language choice or indicating that the skill is intentionally limited to a Chinese-speaking context. That creates a natural-language locale policy issue because the file forces a specific language for runtime messages.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The authentication error message is entirely in Chinese and includes locale-specific support instructions, with no evidence of user opt-in to that language. This forces a specific language/locale in a user-visible path and therefore violates the natural-language policy criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The user-facing strings in this module are entirely in Chinese, including warnings and status messages, with no indication that the skill is China-specific or that users can opt into another language. This creates a natural-language locale policy issue because the skill implicitly forces one language on all users without documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This JavaScript file contains user-facing natural-language strings exclusively in Chinese, including the module description and banner text. Under the policy for natural-language constraints, forcing a specific language without explicit user choice can be a locale-policy violation unless clearly documented as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The file’s docstrings, validation errors, and formatted result messages are all hard-coded in Chinese, including user-visible strings such as error messages and search result labels. This imposes a specific language/locale on users without any apparent opt-in or configuration path, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The manifest description and usage guidance are presented entirely in Chinese and assume Chinese-language interaction, but do not state that the skill is Chinese-only or offer an alternative language option. Under the language/locale policy, forcing a specific language without user opt-in can be a natural-language policy issue unless clearly justified and documented.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language fields in this schema, including the title and descriptions, are entirely in Chinese and do not indicate that the skill is region-specific or that users may choose another language. Under the policy rules, a fixed language without opt-in or justification is a locale-policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The schema title and description are written only in Chinese, which imposes a specific language on users or downstream tooling without indicating that the schema is locale-specific or offering an alternative. Under the policy, language constraints should be justified or optional rather than implicit.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

This JSON schema uses Chinese-only natural-language metadata in the title and field descriptions, which can amount to a language/locale policy constraint if the skill is expected to support users more generally. There is no accompanying indication that the schema is intentionally region-specific or that language choice is configurable.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

This manifest-style JSON file uses Chinese-only natural-language titles and descriptions for the skill inputs, but it does not state that the skill is intended only for Chinese-speaking users or offer any language choice. Under the policy, forcing a specific language without opt-in can be a locale/language policy violation.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:24