T01 · Skill Instruction Hijacking
- Location
src/douyin_extractor.py:413- Finding
Persistent Referral Advertising Injected into Skill Output
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill has a coherent Douyin transcription purpose, but it can download and execute unverified FFmpeg binaries and sends media to external speech services, so users should review it carefully before installing.
Install only if you are comfortable with the skill downloading video content, uploading extracted audio to transcription providers, and running local FFmpeg. Prefer manually installing FFmpeg from a trusted package manager, pinning dependencies, and avoiding private or sensitive videos until the auto-installer, URL validation, referral output, and privacy disclosures are corrected.
src/douyin_extractor.py:413Persistent Referral Advertising Injected into Skill Output
scripts/install_ffmpeg.py:20Mutable and Unverified FFmpeg Payload Is Downloaded and Executed
scripts/install_ffmpeg.py:119Unsafe Archive Extraction Permits Path Traversal and Link-Based File Writes
src/douyin_extractor.py:256User-Controlled URLs Permit Server-Side Request Forgery
package.json:26Dependencies Are Installed Without Version or Integrity Pinning
The documented behavior goes beyond simple Douyin parsing and transcription by describing automatic FFmpeg download, platform detection, archive extraction, permission changes, and local executable installation. Downloading and installing binaries at runtime materially increases supply-chain and local-execution risk, especially when those system-modifying behaviors are not the core expected function of a text-extraction skill.
The documented behavior goes beyond simple Douyin parsing and transcription by describing automatic FFmpeg download, platform detection, archive extraction, permission changes, and local executable installation. Downloading and installing binaries at runtime materially increases supply-chain and local-execution risk, especially when those system-modifying behaviors are not the core expected function of a text-extraction skill.
Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.
# 10. 验证安装
print("\n🧪 验证安装...")
test_env = os.environ.copy()
test_env["PATH"] = str(ffmpeg_bin_dir) + os.pathsep + test_env["PATH"]
try:
The README promotes transcript extraction but does not clearly warn that shared Douyin links, downloaded media, extracted audio, and resulting transcript data may be sent to external speech-recognition providers. In this skill context, users may submit copyrighted, private, or regulated content, so the missing disclosure increases the risk of unintended third-party data sharing.
The README states that the API key 'will not be uploaded to any server', but the tool's documented transcription workflow necessarily sends authenticated requests and media-derived content to external speech-recognition services. This is a misleading security/privacy claim that can cause users to process sensitive content under false assumptions about data exposure and trust boundaries.
The skill documentation exposes capabilities that imply access to environment variables, file read/write, network, and shell execution, but it does not declare any explicit tool scope or permissions boundary. In an agent/MCP context, this makes the effective privilege set opaque to users and reviewers, increasing the chance that the skill can access secrets or modify the local system beyond what users expect.
The README explicitly states that FFmpeg will be automatically downloaded and installed on first use, but does not provide a clear warning that this modifies the local system and may execute downloaded software. In a skill intended for agent-driven use, silent system modification is risky because users may trigger it indirectly and may not realize they are allowing software installation.
The package description is written entirely in Chinese and does not indicate any user language choice or that the skill is limited to a Chinese-only audience for compliance or regional reasons. Under the stated policy, forcing a specific language or locale without opt-in is a natural-language policy concern.
This Python file contains its title, description, prompts, and status messages entirely in Chinese, starting with the module docstring. The policy requires flagging language or locale constraints when the skill forces a specific language without user opt-in, and this file does not offer any alternative language or justify a Chinese-only audience.
The installer downloads archives from external URLs, extracts them, changes permissions, and later executes the resulting binaries, creating a software supply-chain risk. In the context of a text-extraction skill, bundling an auto-installer for third-party executables increases attack surface significantly, especially because mirror URLs use mutable 'latest' release endpoints rather than immutable pinned artifacts.
This code executes a freshly downloaded external binary from the local skill directory after download and extraction, but the script does not verify signatures, checksums, or pinned release artifacts before execution. If the download source, mirror, network path, or extracted contents are compromised, the script will grant execute permission and run attacker-controlled code.
test_env["PATH"] = str(ffmpeg_bin_dir) + os.pathsep + test_env["PATH"]
try:
result = subprocess.run(
[str(ffmpeg_bin_dir / "ffmpeg"), "-version"],
capture_output=True,
text=True,
The description in L04 is entirely in Chinese, and the environment variable descriptions in L42 and L47 also assume Chinese-language use. For a general-purpose skill manifest, this effectively forces a specific language for users without any stated opt-in or justification, which matches the language/locale policy concern.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def check_ffmpeg():
"""检查 FFmpeg 是否可用"""
try:
result = subprocess.run(
["ffmpeg", "-version"],
capture_output=True,
text=True,
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def check_ffmpeg():
"""检查 FFmpeg 是否可用"""
try:
result = subprocess.run(
["ffmpeg", "-version"],
capture_output=True,
text=True,
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def check_ffmpeg():
"""检查 FFmpeg 是否可用"""
try:
result = subprocess.run(
["ffmpeg", "-version"],
capture_output=True,
text=True,
A Douyin text extraction tool is not expected to install system software automatically, and doing so materially increases host-side risk. This capability is more dangerous in this context because users may run the skill for simple extraction and not realize it can launch local installation code.
The skill can automatically execute a local installer script, which expands its behavior from media processing into code execution on the host. If the repository or scripts/install_ffmpeg.py is modified, replaced, or supplied from an untrusted source, the tool will run arbitrary Python code with the user's privileges.
if install_script.exists():
print("\n🚀 启动 FFmpeg 自动安装...")
subprocess.run([sys.executable, str(install_script)])
# 验证安装
installed, version = check_ffmpeg()
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
"""抖音文案提取器"""
# 硅基流动 API 配置
SILICONFLOW_API_URL = "https://api.siliconflow.cn/v1/audio/transcriptions"
SILICONFLOW_MODEL = "FunAudioLLM/SenseVoiceSmall"
# 邀请码信息
During extraction, the workflow can trigger FFmpeg installation logic without prominent advance warning in the primary execution path. Even though there is an interactive prompt, the capability to move from content processing to software installation is security-relevant and easy for users to underestimate.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
"-y", str(audio_file)
]
subprocess.run(cmd, capture_output=True, check=True)
return str(audio_file)
def _get_audio_info(self, audio_file: str) -> Tuple[float, int]:
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
"-y", str(audio_file)
]
subprocess.run(cmd, capture_output=True, check=True)
return str(audio_file)
def _get_audio_info(self, audio_file: str) -> Tuple[float, int]:
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
audio_file
]
result = subprocess.run(cmd, capture_output=True, text=True, check=True)
duration = float(result.stdout.strip())
file_size = os.path.getsize(audio_file)
The tool uploads extracted audio to a third-party transcription API, which transmits potentially sensitive spoken content off-device. This is especially relevant because the skill handles user media and the main workflow does not present a clear consent/privacy warning at the moment of upload.
The module docstring and later user-facing guidance/messages are presented only in Chinese, which effectively forces a specific language experience. The file does not indicate that the skill is China-specific only, nor does it offer any opt-in or alternative locale handling.
The manifest describes a Douyin text extractor and MCP server for extracting watermark-free video and transcribing speech. This code invokes a local executable via subprocess to inspect FFmpeg, which is a host-level execution capability beyond the core stated purpose and not explicitly declared in the manifest.
Detected: suspicious.exposed_secret_literal