Back to skill

Security audit

TikTok/Douyin 创作流水线

Security checks for vulnerabilities and agentic risk

Overview

The skill's TikHub scraping and transcription purpose is disclosed, but unsafe command execution, detached background jobs, arbitrary URL downloads, and weak secret/dependency handling warrant Review before installation.

Install only after reviewing or fixing the unsafe background transcription and URL-download paths; use an isolated environment, pin dependencies, avoid command-line API keys, restrict scraping/downloading to content you are authorized to process, and confirm paid TikHub calls before running batch jobs.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/tikhub.py:296
Finding

Shell Command Injection in Background Whisper Transcription

Content
View full analysis
{log_file} 2>&1 &" print(f"🚀 Whisper 后台转写启动,日志: {log_file}") subprocess.run(nohup_cmd, shell=True) ``` ### Technical Analysis The function constructs a command string by joining an argument list and then executes that string with `shell=True`. Values such as `audio_path`, `model`, and `language` are not shell-escaped or restricted to safe values. Because the resulting string is interpreted by a command shell, shell metacharacters in any attacker-influenced argument can introduce additional commands. Using a Python list initially does not provide protection because the list is converted back into an unquoted string before execution. The log-file path is also incorporated into the shell command without quoting. Although `os.path.basename()` removes directory components, it does not remove shell metacharacters. ### Attack Path 1. An attacker causes the Agent or an integrating application to call `whisper_transcribe()` with a crafted `audio_path`, `model`, or `language`. 2. The supplied value contains shell syntax, such as a command separator followed by an attacker-selected command. 3. The function joins the arguments into `nohup_cmd` without quoting or validation. 4. `subprocess.run(..., shell=True)` passes the command to the system shell. 5. The shell interprets the injected syntax and executes the additional command with the privileges of the Agent process. ### Impact Assessment Successful exploitation permits arbitrary local command execution with the same op ...[truncated 637 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/tikhub.py:126
Finding

Unrestricted URL Fetching Enables SSRF and Unbounded Downloads

Content
View full analysis
str: """ 下载视频到本地,返回文件路径 aweme_id: 视频 ID video_url: 可选,传入直链可跳过 API 调用 请求头需要 Referer + User-Agent,否则 403 """ os.makedirs(output_dir, exist_ok=True) if not video_url: video_url = get_high_quality_url(aweme_id) if not video_url: print("❌ 无法获取视频地址,请确认 API 余额充足") return None local_path = os.path.join(output_dir, f"{aweme_id}.mp4") headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) " "AppleWebKit/537.36 (KHTML, like Gecko) " "Chrome/120.0.0.0 Safari/537.36", "Referer": "https://www.douyin.com/", } print(f"⬇️ 开始下载: {video_url[:80]}...") resp = requests.get(video_url, headers=headers, stream=True, timeout=120) ``` The response body is subsequently streamed to disk without a maximum-size check: ```python with open(local_path, "wb") as f: for chunk in resp.iter_content(chunk_size=1024 * 1024): if chunk: f.write(chunk) ``` ### Technical Analysis `download_video()` accepts a caller-provided `video_url` and sends a request to it without validating: - The URL scheme. - The destination hostname. - The resolved IP address. - Redirect destinations. - Whether the destination is loopback, private, link-local, or otherwise internal. - The response content type. - The maximum permitted response size. This permits the Agent's network position to be used to contact services that may not be reachable by the attacker directly. The TikHub bearer authorization header is not attached to this particular request, so this issue does not directly disclose that key. Nev ...[truncated 1816 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/batch.py:37
Finding

TikHub API Key Exposed Through Command-Line Arguments

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Third-Party Dependencies Create Supply-Chain Risk

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (22)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This is the strongest issue in the file: user-influenced parameters are embedded into a shell command and executed with shell=True, enabling shell metacharacter injection. Because the skill processes downloaded/local media files and derives names from paths, an attacker who controls or influences audio_path can achieve arbitrary OS command execution in the agent environment.

Content

Scanner excerpt · scripts/tikhub.py (reported line 307)May include surrounding context.

python
log_file = f"/tmp/whisper_{os.path.basename(audio_path)}.log"
        nohup_cmd = f"nohup {' '.join(cmd)} > {log_file} 2>&1 &"
        print(f"🚀 Whisper 后台转写启动,日志: {log_file}")
        subprocess.run(nohup_cmd, shell=True)
        print(f"📝 文字稿将保存到: {output_path}")
        print(f"⏱️  medium 模型 CPU 转写 1 分钟音频约需 1-2 分钟,请耐心等待")
        return output_path

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill demonstrates network access, shell execution, and file-writing behavior but does not declare any tool scope or permission boundaries. In an agent environment, this increases the chance the skill will be invoked with overly broad capabilities, enabling unintended downloads, local file creation, and command execution without explicit user- or platform-level constraint.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The activation description is broad and trigger-based, covering scraping, downloading, user information retrieval, fan lists, and transcription requests across multiple platforms without clear exclusions or consent checks. This can cause the agent to route sensitive or policy-relevant requests into the skill too aggressively, including large-scale collection of third-party data or copyrighted media processing.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
79% confidence
Finding

The skill is explicitly designed to send data to an external third-party API and also facilitates downloading and processing user-supplied platform content. In context, external transmission is expected, but it is still security-relevant because URLs, identifiers, metadata, and potentially paid API usage can be sent off-platform without explicit disclosure, creating privacy, compliance, and cost risks.

Content

Scanner excerpt · SKILL.md (reported line 160)May include surrounding context.

md
- **视频下载计费**:每次调用付费端点都会被计费,注意余额
- **转写速度**:mlx-whisper(Apple GPU)> openai-whisper small(CPU)> openai-whisper medium(CPU)
- **faster-whisper 不支持 Apple MPS**:不要用!
- **API 文档**:https://api.tikhub.io/docs(Swagger UI)

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The inline comments say GPU mode returns text content while CPU mode returns a path, but the code then stores the GPU-mode value in a variable named text_path and records it as a path. This is an active contradiction in the code documentation/annotation around what full_pipeline_douyin_to_text returns, which can mislead operators about outputs and side effects.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/tikhub.py (reported line 52)May include surrounding context.

python
"""通用 POST 请求"""
    for i in range(retries):
        try:
            resp = requests.post(f"{BASE_URL}{endpoint}", headers=HEADERS, json=json_data, timeout=30)
            if resp.status_code == 429:
                print(f"⚠️ 频率限制,等待 5 秒...")
                time.sleep(5)

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest frames the skill primarily as a TikHub-based multi-platform scraping/downloading tool, with link-to-text described as a download→audio→Whisper pipeline. The implementation realizes that feature by invoking local executables via subprocess, including ffmpeg and a shell-based nohup Whisper launch, which is a materially broader capability than ordinary API use and local file handling.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/tikhub.py (reported line 262)May include surrounding context.

python
"""
    if not output_path:
        output_path = video_path.rsplit(".", 1)[0] + ".wav"
    result = subprocess.run([
        "ffmpeg", "-i", video_path, "-vn",
        "-acodec", "pcm_s16le", "-ar", "16000", "-ac", "1",
        output_path, "-y"

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The function whisper_transcribe defaults language to Chinese, and its docstring presents that as the expected setting rather than offering a neutral default or explicit user choice. This is a natural-language locale policy issue because the skill imposes a specific language by default without opt-in or justification that the tool is region-specific only.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

The design explicitly supports running Whisper under nohup in the background, creating a detached process that persists beyond the immediate session. In an agent-skill context this is risky because it bypasses normal lifecycle expectations, complicates monitoring, and can be abused to leave unauthorized long-running jobs on the host.

Content

Scanner excerpt · scripts/tikhub.py (reported line 287)May include surrounding context.

python
model: tiny/base/small/medium/large,medium 精度速度平衡好
        language: 语言代码,Chinese
        output_path: 可选,文字稿输出路径
        background: True 用 nohup 后台跑(推荐,CPU 慢);False 同步等待
    返回: 文字稿文件路径(后台模式立即返回路径,不等待完成)
    """
    if not output_path:

Session Persistence

Medium
Category
Rogue Agent
Confidence
96% confidence
Finding

This code constructs and launches a nohup-backed background process, allowing execution to survive the parent context. In hosted or agent environments, that persistence can be abused for covert long-running computation, resource exhaustion, or evasion of expected execution boundaries.

Content

Scanner excerpt · scripts/tikhub.py (reported line 305)May include surrounding context.

python
if background:
        log_file = f"/tmp/whisper_{os.path.basename(audio_path)}.log"
        nohup_cmd = f"nohup {' '.join(cmd)} > {log_file} 2>&1 &"
        print(f"🚀 Whisper 后台转写启动,日志: {log_file}")
        subprocess.run(nohup_cmd, shell=True)
        print(f"📝 文字稿将保存到: {output_path}")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
99% confidence
Finding

This code builds a shell command from user-influenced values and executes it with shell=True under nohup. Because audio_path is incorporated into the joined command string without shell escaping, an attacker can craft a filename containing shell metacharacters to trigger arbitrary command execution and leave it running in the background.

Content

Scanner excerpt · scripts/tikhub.py (reported line 307)May include surrounding context.

python
log_file = f"/tmp/whisper_{os.path.basename(audio_path)}.log"
        nohup_cmd = f"nohup {' '.join(cmd)} > {log_file} 2>&1 &"
        print(f"🚀 Whisper 后台转写启动,日志: {log_file}")
        subprocess.run(nohup_cmd, shell=True)
        print(f"📝 文字稿将保存到: {output_path}")
        print(f"⏱️  medium 模型 CPU 转写 1 分钟音频约需 1-2 分钟,请耐心等待")
        return output_path

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/tikhub.py (reported line 312)May include surrounding context.

python
print(f"⏱️  medium 模型 CPU 转写 1 分钟音频约需 1-2 分钟,请耐心等待")
        return output_path
    else:
        result = subprocess.run(cmd, capture_output=True, text=True)
        if result.returncode != 0:
            print(f"❌ Whisper 错误: {result.stderr[-300:]}")
        print(f"📝 转写完成: {output_path}")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The mlx_whisper_transcribe function sets language to zh, and the surrounding natural-language description frames Chinese/English as fixed options with Chinese as the default. That constitutes a locale constraint imposed by the skill without explicit user selection or a documented region-specific justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The example sets language="zh" for transcription, which forces a specific language in the documented behavior without mentioning that the user can choose another language. This is a natural-language locale constraint and may violate language/locale policy when presented as a default workflow without opt-in or justification.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency list includes requests without a version pin, which makes builds non-reproducible and can silently pull in a vulnerable or breaking release. In a skill that performs network scraping and API access, dependency drift increases supply-chain risk and makes it impossible to verify whether known fixes are present.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests
openai-whisper
mlx-whisper
ffmpeg

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
93% confidence
Finding

requests has known advisories, and because no version is pinned, there is no way to determine whether the deployed environment includes a fixed or vulnerable release. This is particularly relevant for a scraping/API skill, where attacker-controlled URLs, redirects, or credential-handling edge cases may be reachable during normal operation.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

openai-whisper is unpinned, so installation may resolve to different versions over time with different security or behavior characteristics. Because this skill processes downloaded media and transcription pipelines, uncontrolled dependency changes can introduce supply-chain exposure or unexpected execution paths.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests
openai-whisper
mlx-whisper
ffmpeg

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

mlx-whisper is specified without a version, preventing reliable auditing of the installed package and allowing unreviewed updates to enter the environment. In a media-processing toolchain, this creates avoidable supply-chain risk even if no specific exploit is demonstrated in the manifest itself.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
requests
openai-whisper
mlx-whisper
ffmpeg

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
87% confidence
Finding

ffmpeg is listed without a version, so the environment may install different releases with varying security posture and codec-parsing behavior. Since FFmpeg commonly handles untrusted media input, lack of version control is more concerning in this skill context than in a purely offline utility.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
requests
openai-whisper
mlx-whisper
ffmpeg

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The top-level docstring describes the skill only in Chinese ('抖音/TikTok 数据爬取工具'), which can indicate a fixed language/locale presentation without user opt-in. In this file there is no accompanying English alternative, opt-in mechanism, or justification that the skill is intended only for a Chinese-language context.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The docstring for mlx_whisper_transcribe states '返回: 文字稿内容(str)', and full_pipeline_douyin_to_text treats the return value as transcript text, but elsewhere the surrounding pipeline and naming conventions imply a path-oriented contract. More concretely, full_pipeline_douyin_to_text returns output_path while whisper_transcribe returns a path, creating contradictory intent around what transcription helpers return.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.