Back to skill

Security audit

douyin-download

Security checks for vulnerabilities and agentic risk

Overview

The skill largely matches a Douyin download/transcription purpose, but its documentation understates where transcript data goes and the URL fetching is broader than that purpose requires.

Install only if you are comfortable with Douyin media being downloaded locally, audio being sent to SiliconFlow, and transcripts potentially being sent to MiniMax when segmentation is enabled. Avoid using it on private or sensitive videos, prefer --no-segment unless MiniMax processing is intended, and do not pass arbitrary non-Douyin URLs.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
douyin.js:163
Finding

Unrestricted User-Controlled URL Fetching Enables SSRF

Content
View full analysis
{ const parsedUrl = new URL(url); const client = parsedUrl.protocol === 'https:' ? https : http; const opts = { method: options.method || 'GET', headers: { ...HEADERS, ...options.headers } }; const req = client.request(url, opts, (res) => { if (res.statusCode >= 300 && res.statusCode < 400 && res.headers.location) { httpRequest(res.headers.location, options) .then(resolve) .catch(reject); return; } let data = ''; res.on('data', chunk => data += chunk); res.on('end', () => { resolve({ statusCode: res.statusCode, headers: res.headers, body: data, url: url }); }); }); ``` ```js async function parseShareUrl(shareText) { const modalIdMatch = shareText.match(/(?:modal_id[=:])?(\d{16,})/); if (modalIdMatch) { const modalId = modalIdMatch[1]; return await getVideoInfoByModalId(modalId); } const urlMatch = shareText.match(/https?:\/\/[^\s]+/); if (!urlMatch) { throw new Error('No valid sharing URL was found'); } const shareUrl = urlMatch[0]; const response1 = await httpRequest(shareUrl); const finalUrl = response1.url; ``` ### Technical Analysis The skill extracts an arbitrary HTTP or HTTPS URL from caller-controlled input and sends a request to it. There is no hostname allowlist, destination IP validation, DNS rebinding defense, or rejection of loopback, private, link-local, and cloud metadata addresses. The HTTP helper also follows redirects recursively without validating the new destination or imposing an explicit redirect limit. Therefore, even if the initial URL points to a public server, ...[truncated 1781 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
douyin.js:276
Finding

Bearer API Keys Are Exposed Through Child Process Arguments

Content
View full analysis
Remediation
View remediation

other

Warning
Location
douyin.js:329
Finding

Transcripts Are Sent to an Undocumented External MiniMax Service by Default

Content
View full analysis
{ const { spawn } = require('child_process'); const proc = spawn('curl', [ '-X', 'POST', url, '-H', `Authorization: Bearer ${apiKey}`, '-H', 'Content-Type: application/json', '-d', JSON.stringify(data) ]); ``` ```js let textContent = await transcribeAudio(audioPath, apiKey, showProgress); if (doSegment) { textContent = await semanticSegment(textContent, null, showProgress); } ``` The documented behavior states that semantic segmentation uses an OpenClaw built-in LLM. The implementation instead sends the complete transcript to `https://api.minimaxi.com`. Semantic segmentation is enabled by default, while `MINIMAX_API_KEY` and this additional third-party transmission are not declared in the documented skill metadata requirements. ### Technical Analysis After audio is uploaded to SiliconFlow for transcription, the resulting transcript is passed as a user message to the external MiniMax chat-completion API. The `doSegment` option defaults to `true`, so this disclosure occurs in the normal extraction path whenever a MiniMax key is available. This behavior materially differs from the documentation, which represents segmentation as an operation performed by an OpenClaw built-in LLM. Users therefore may not understand that transcript content is transmitted to a second external service. Transcripts can contain personal conversations, names, contact information, business information, copyrighted ...[truncated 1382 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (6)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill declares shell and environment-dependent behavior but does not define any explicit tool scope such as permissions or allowed-tools. That omission weakens containment and reviewability, making it easier for the skill to invoke broader command execution or access sensitive environment data than users may expect.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill documentation instructs users to download videos and send audio/content to an external ASR service and an LLM, but it does not clearly warn that user-supplied media or extracted content may leave the local environment and that files will be written to disk. This can lead to unintended disclosure of private data, copyrighted media, or sensitive content through normal use.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

The tool uploads extracted audio to a third-party transcription API and optionally sends transcript text to a second external LLM service. This is an actual data-exposure risk because video-derived audio and text may contain sensitive or copyrighted content, and the transmission happens as part of normal operation with limited disclosure, consent controls, or destination restrictions.

Content

Scanner excerpt · douyin.js (reported line 39)May include surrounding context.

js
'Accept-Language': 'zh-CN,zh;q=0.9',
};

const SILI_FLOW_BASE_URL = 'https://api.siliconflow.cn/v1/audio/transcriptions';
const SILI_FLOW_MODEL = 'FunAudioLLM/SenseVoiceSmall';
const MINIMAX_BASE_URL = 'https://api.minimaxi.com';
const DEFAULT_DOWNLOAD_PATH = '/tmp/douyin-download/';

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest description and primary documentation are presented solely in Chinese, and the file does not indicate that the skill is region-specific or provide an alternate language option. This can violate language/locale policy when a skill forces one language without user opt-in or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The request headers hard-code Accept-Language: zh-CN,zh;q=0.9, which forces a specific language preference for outbound requests. This is a natural-language/locale policy concern because the skill does not offer a user choice or document a justified region-specific restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The generated transcript uses toLocaleString('zh-CN'), which forces Chinese locale formatting in output regardless of user preference. This violates the language/locale policy unless the skill offers opt-in or clearly documents that it is region-specific.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
douyin.js:143