T09 · Insecure Skill Coding Practices
- Location
scripts/transcribe.js:104- Finding
Arbitrary Command Execution Through Shell Command Injection
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is a real Douyin/video transcription helper, but it needs Review because it handles secrets unsafely and its command execution path can be abused with crafted URLs or filenames.
Install only after reviewing the command-injection and API-key handling risks. Do not paste API keys into chat; configure secrets out of band and rotate any key already shared. Treat links and filenames as untrusted, and avoid running this skill on crafted input until command execution is changed to non-shell argument arrays with strict URL validation.
scripts/transcribe.js:104Arbitrary Command Execution Through Shell Command Injection
SKILL.md:257API Secret Requested Through Conversational Input
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
cd ~/.openclaw/workspace/skills/douyin-transcribe
cp .env.example .env
# 编辑 .env,填入 GROQ_API_KEY
The manifest presents the skill as a simple transcription assistant, but the instructions include remote media extraction, file persistence, browser scripting, shell execution, and external API submission of content. This mismatch prevents informed consent and can cause users to expose private video/audio data or local environment state without realizing the full behavior.
Referenced artifact was not completely inspected
node scripts/transcribe.js "<原始链接>"
Referenced artifact was not completely inspected
node scripts/transcribe.js "<原始链接>"
Referenced artifact was not completely inspected
node scripts/transcribe.js "<原始链接>"
Instructing the agent to read the skill's .env file exposes a credential-bearing file to the model and toolchain even though only the presence of configuration is needed. Reading .env content can reveal API keys or other secrets and is broader than necessary for normal transcription behavior.
当用户第一次触发时,按以下顺序检查环境。缺什么补什么。
read: <skill目录>/.env
Automatically copying or managing a .env file is not inherently malicious, but in the context of this skill it is part of a credential-handling workflow performed by the agent. That increases the chance of misconfiguration or later exposure of secrets if the agent also reads or writes values in that file.
read: <skill目录>/.env
不存在则从 .env.example 复制。
The skill explicitly asks users to send a Groq API key in chat without any warning that the value is a sensitive credential. Secrets pasted into chat may be stored in logs, appear in model context, or be exposed to other tools, creating a significant credential-compromise risk.
The skill directs the agent to solicit a raw API key from the user and persist it into local configuration. This is dangerous because it combines secret collection, local storage, and agent-mediated handling, increasing the attack surface for leakage, misuse, and later exfiltration.
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
const path = require('path');
const https = require('https');
// 加载 .env 文件
function loadEnvFile() {
const envPath = path.join(__dirname, '..', '.env');
if (fs.existsSync(envPath)) {
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
const path = require('path');
const https = require('https');
// 加载 .env 文件
function loadEnvFile() {
const envPath = path.join(__dirname, '..', '.env');
if (fs.existsSync(envPath)) {
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
const path = require('path');
const https = require('https');
// 加载 .env 文件
function loadEnvFile() {
const envPath = path.join(__dirname, '..', '.env');
if (fs.existsSync(envPath)) {
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
const path = require('path');
const https = require('https');
// 加载 .env 文件
function loadEnvFile() {
const envPath = path.join(__dirname, '..', '.env');
if (fs.existsSync(envPath)) {
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
// 加载 .env 文件
function loadEnvFile() {
const envPath = path.join(__dirname, '..', '.env');
if (fs.existsSync(envPath)) {
const envContent = fs.readFileSync(envPath, 'utf-8');
envContent.split('\n').forEach(line => {
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
**Windows:** 从 [gyan.dev](https://www.gyan.dev/ffmpeg/builds/) 下载 release full 版本,解压后将 `bin` 目录加入 PATH。
**Linux:** `sudo apt install ffmpeg`
### 2. 获取 Groq API Key(免费)
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
1. 打开 [console.groq.com](https://console.groq.com)
2. Google 或 GitHub 登录(不需要信用卡)
3. **API Keys** → **Create API Key** → 复制(`gsk_` 开头)
### 3. 配置
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
1. 打开 [console.groq.com](https://console.groq.com)
2. Google 或 GitHub 登录(不需要信用卡)
3. **API Keys** → **Create API Key** → 复制(`gsk_` 开头)
### 3. 配置
The README explains that audio is processed by Groq, but it does not present this as a clear user-facing privacy warning before use or installation. Because the skill handles user-supplied video/audio that may contain personal or sensitive content, sending transcripts or audio to a third-party API without explicit disclosure can cause unintended data exposure and informed-consent failures.
The skill invokes sensitive capabilities including environment file access and modification, browser automation, and command execution, but does not declare any tool scope or allowed-tools boundary. That makes the effective privilege set implicit and harder to review, increasing the chance the agent will overreach into host or credential-handling actions beyond user expectations.
The skill saves uploaded video files and transcript outputs to local directories and later supports archiving them, but gives no upfront warning about retention, storage location, or possible sensitivity of spoken content. This can lead to unintended persistence of personal, copyrighted, or confidential material.
The skill tells the agent to ask the user for a raw Groq API key in chat and then write it into a local .env file. Collecting and handling credentials directly in the conversation creates a clear secret-exposure risk through chat logs, model context, and accidental reuse or disclosure.
The trigger list includes generic terms such as "转录", "transcribe", and "视频转文本" that are not unique to Douyin and can match many unrelated user requests. This can cause unintended invocation of a skill that has access to powerful tools like browser and exec, increasing the chance of unnecessary exposure to untrusted links or files and accidental execution paths.
The script uploads full audio content to Groq or OpenAI for transcription, but the user-facing flow does not clearly warn that media leaves the local machine and is processed by a third party. Because Douyin videos and local files may contain personal, copyrighted, or sensitive information, silent remote transmission creates a real privacy and compliance risk.
The transcription request hard-codes the language parameter to 'zh', which enforces Chinese-language processing regardless of the user's actual content or preference. The file does not present this as an opt-in locale choice or explain a compliance-bound reason for restricting processing to Chinese only.
After transcription, the script sends transcript text to Groq again for punctuation and formatting without a separate disclosure that another third-party processing step occurs. This expands data exposure beyond the minimum necessary operation and may transmit sensitive speech-derived text even when the user only expected speech-to-text processing.
Detected: suspicious.dangerous_exec, suspicious.exposed_secret_literal