Back to skill

Security audit

Douyin Video Transcribe

Security checks for vulnerabilities and agentic risk

Overview

This skill is mostly for Douyin transcription, but it should be reviewed because it can start a persistent Docker transcription service from an unpinned third-party image and expose it on the host network.

Install only if you are comfortable letting the skill download media, process local audio/video, and start a Docker Whisper server. Safer use would pin the Docker image to a reviewed digest, bind the ASR port to 127.0.0.1, require explicit approval before cloud transcription or container creation, and stop/remove the container when finished.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
scripts/whisper_local.py:22
Finding

Automatic Execution of an Unpinned Third-Party Docker Image

Content
View full analysis

Vulnerability Details

File Location: scripts/whisper_local.py:22-23, 97-111
Additional Location: SKILL.md:44-54
Vulnerability Type: Unpinned mutable third-party dependency
Risk Level: Medium

Vulnerable Code

python
CONTAINER_NAME = "whisper-asr"
DOCKER_IMAGE = "onerahmet/openai-whisper-asr-webservice:latest"
python
# 容器不存在,创建新的
print(f"🆕 创建新容器 {self.CONTAINER_NAME} (模型: {self.model})...")
cmd = [
    "docker", "run", "-d",
    "-p", "9000:9000",
    "-e", f"ASR_MODEL={self.model}",
    "-e", "ASR_ENGINE=faster_whisper",
    "--name", self.CONTAINER_NAME,
    self.DOCKER_IMAGE
]

result = subprocess.run(cmd, capture_output=True, text=True, timeout=120)
if result.returncode == 0:
    return True
else:
    print(f"❌ 创建容器失败: {result.stderr}")
    return False

The documentation also recommends executing the same mutable image:

bash
docker run -d -p 9000:9000 --name whisper-asr onerahmet/openai-whisper-asr-webservice:latest

Technical Analysis

The Skill executes onerahmet/openai-whisper-asr-webservice:latest without pinning the image to a reviewed cryptographic digest. The latest tag is mutable, meaning its contents can change after the Skill has been audited without requiring any modification to this repository.

When the expected container does not already exist, start_docker_container() invokes docker run. Docker may retrieve the current image associated with the mutable tag and immediately execute it. Consequently, the effective runtime payload is controlled by the current state of an external container registry and upstream publisher.

The documentation similarly recommends unversioned installation of openai-whisper and execution of the same latest image. No signature verification, digest validation, software bill of materials, or explicit user confirmation is present.

This finding does not establish that the current upstream image is malicious. The vulnerability is the absence of integrity ...[truncated 1705 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin the container image to a reviewed immutable digest:
python
DOCKER_IMAGE = (
    "onerahmet/openai-whisper-asr-webservice"
    "@sha256:<reviewed-image-digest>"
)
  1. Maintain an explicit dependency-update process that:

    • Retrieves a candidate image.
    • Verifies its signature and provenance.
    • Scans it for known vulnerabilities.
    • Reviews its software bill of materials.
    • Updates the pinned digest only after approval.
  2. Require explicit user approval before downloading or launching a previously unavailable image.

  3. Pin Python dependencies to reviewed versions and hashes, for example through a locked requirements file using hash verification.

  4. Harden the container with a non-root user, a read-only root filesystem, dropped Linux capabilities, resource limits, and restricted outbound networking where compatible with the service.

  5. Document the exact reviewed image version and digest in SKILL.md rather than recommending the mutable latest tag.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/whisper_local.py:97
Finding

Unauthenticated ASR Service Exposed on All Host Network Interfaces

Content
View full analysis

Vulnerability Details

File Location: scripts/whisper_local.py:97-111
Additional Location: SKILL.md:51-54
Vulnerability Type: Insecure network service exposure
Risk Level: Medium

Vulnerable Code

python
# 容器不存在,创建新的
print(f"🆕 创建新容器 {self.CONTAINER_NAME} (模型: {self.model})...")
cmd = [
    "docker", "run", "-d",
    "-p", "9000:9000",
    "-e", f"ASR_MODEL={self.model}",
    "-e", "ASR_ENGINE=faster_whisper",
    "--name", self.CONTAINER_NAME,
    self.DOCKER_IMAGE
]

result = subprocess.run(cmd, capture_output=True, text=True, timeout=120)
if result.returncode == 0:
    return True
else:
    print(f"❌ 创建容器失败: {result.stderr}")
    return False

The same exposure is recommended in the documentation:

bash
docker run -d -p 9000:9000 --name whisper-asr onerahmet/openai-whisper-asr-webservice:latest

Technical Analysis

Docker port publication in the form 9000:9000 normally binds the host port to all available interfaces, equivalent to a wildcard address such as 0.0.0.0:9000, unless Docker daemon or host networking configuration overrides that behavior.

The transcription client only requires local access, but the container is exposed more broadly than necessary. The audited configuration does not add authentication, authorization, TLS, request-size controls, or network-level access restrictions to the ASR endpoint.

The container is started in detached mode and the scripts do not stop or remove it after transcription. The exposed service can therefore remain reachable after the Skill operation has completed.

Actual external reachability depends on host firewall rules, network topology, cloud security groups, and Docker configuration. However, the Skill itself does not enforce loopback-only access.

Attack Path

  1. A user invokes the Skill, causing the whisper-asr container to be created.
  2. Docker publishes container port 9000 on all host interfaces.
  3. The container continues running after the transcripti ...[truncated 1299 chars]
Remediation
View remediation

Remediation Suggestions

  1. Bind the service exclusively to the loopback interface:
python
cmd = [
    "docker", "run", "-d",
    "-p", "127.0.0.1:9000:9000",
    "-e", f"ASR_MODEL={self.model}",
    "-e", "ASR_ENGINE=faster_whisper",
    "--name", self.CONTAINER_NAME,
    self.DOCKER_IMAGE
]
  1. Update SKILL.md to use:
bash
docker run -d -p 127.0.0.1:9000:9000 \
  --name whisper-asr \
  <pinned-image-reference>
  1. If remote access is genuinely required:

    • Place the service behind an authenticated reverse proxy.
    • Use TLS.
    • Restrict source networks through firewall rules.
    • Add request-size limits and rate limiting.
    • Use per-user authorization and audit logging.
  2. Add container resource restrictions, such as memory, CPU, process, and storage limits, to reduce denial-of-service impact.

  3. Stop and remove automatically created containers after processing unless persistent operation was explicitly requested by the user.

  4. Verify service readiness through a specific health endpoint and require a successful expected response rather than treating any HTTP response as proof that the intended ASR service is running.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (22)

Tp4

High
Category
MCP Tool Poisoning
Confidence
84% confidence
Finding

The code substantially matches part of the declared purpose: it does extract audio from video files and transcribe audio/video content. However, the declared description presents a broader Douyin transcription suite that supports Douyin links, local files, and image notes, and can analyze content. This code chunk only accepts local file paths through a CLI/API, performs media preprocessing, and transcribes using one of several ASR backends. There is no code here for parsing or downloading Douyin links, no image-note processing/OCR, and no summarization or analysis logic. These are material gaps in represented capabilities, so the description is not fully accurate for this supplied code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a broader Douyin transcription suite that can handle Douyin links, local files, and image notes, then transcribe and analyze content. The supplied code chunk does not implement any Douyin-specific functionality, video/audio extraction, image processing, summarization, or analysis. Instead, it is a narrow local ASR utility: it talks to a local Whisper webservice, checks service health, starts or creates a Docker container, and submits an existing local audio file for transcription. Docker management is an additional operational capability not mentioned in the description. Because the actual behavior is substantially narrower and materially different from the declared end-user purpose, this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill documents use of browser automation, network access, shell commands, and local file output, but it does not declare any tool scope or permissions boundary. This creates an authorization gap where an agent may invoke broader capabilities than users or reviewers expect, increasing the chance of unintended downloads, command execution, and file writes.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill instructs the agent to download remote videos and save them locally without prominently disclosing that behavior in the skill description or permission boundary. Hidden or insufficiently disclosed data transfer and local persistence can surprise users, create storage/privacy issues, and increase the risk of handling untrusted media without informed consent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The workflow sends audio to an HTTP Whisper service on localhost, but the skill does not clearly warn that user content is transmitted to another service for processing. Even on localhost, this is a separate network-exposed component that may log, retain, or mishandle sensitive audio, making the omission a meaningful privacy and trust issue.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The fallback logic can automatically send user-provided local audio to third-party cloud transcription services when the preferred local method is unavailable, without an explicit consent check at the decision point. That creates a privacy and data-governance risk because sensitive local media may leave the host unexpectedly.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The manifest describes a Douyin transcription suite, but this file checks Docker availability and later relies on ffmpeg/ffprobe subprocesses for media handling. While media conversion is expected, spawning external tools and container services is a broader host-execution capability that is not explicitly justified by the manifest text.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/transcriber.py (reported line 165)May include surrounding context.

python
if method == "whisper_local":
            # 检查 Docker 是否可用
            try:
                result = subprocess.run(
                    ["docker", "version"],
                    capture_output=True, text=True, timeout=3
                )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/transcriber.py (reported line 257)May include surrounding context.

python
output_path, "-y"
        ]
        
        result = subprocess.run(cmd, capture_output=True, text=True, timeout=300)
        if result.returncode != 0:
            raise RuntimeError(f"音频提取失败: {result.stderr[:500]}")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/transcriber.py (reported line 288)May include surrounding context.

python
"-show_streams",
                wav_path
            ]
            result = subprocess.run(cmd, capture_output=True, text=True, timeout=10)
            if result.returncode == 0:
                info = json.loads(result.stdout)
                for stream in info.get("streams", []):

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The --language argument only permits zh and en, which imposes a fixed language constraint in the user-facing interface. There is no accompanying justification that this is a region-specific tool or any broader language-choice mechanism beyond those two options.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module title, docstrings, user-facing messages, and CLI usage text are written entirely in Chinese, which imposes a specific language on users. The file does not offer an opt-in language choice or explain that the skill is intentionally limited to a Chinese-speaking context.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code grants the skill the ability to inspect, start, and run Docker containers on the host, which is broader than simple media transcription and can be dangerous on systems where Docker access is privileged. In context, the capability is adjacent to the stated purpose but still overpowered for a transcription client, increasing blast radius if the skill is misused or modified.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/whisper_local.py (reported line 45)May include surrounding context.

python
"""
        try:
            # 检查 Docker 是否可用
            result = subprocess.run(
                ["docker", "version"],
                capture_output=True, text=True, timeout=5
            )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/whisper_local.py (reported line 53)May include surrounding context.

python
return "docker_not_available"
            
            # 检查容器状态
            result = subprocess.run(
                ["docker", "inspect", "--format", "{{.State.Status}}", self.CONTAINER_NAME],
                capture_output=True, text=True, timeout=5
            )

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The helper does more than local transcription: it automatically creates and starts a Docker container, which materially expands the skill's authority and attack surface. In a skill triggered by user content, unexpected environment modification and execution of a third-party container can violate least privilege and expose the host to supply-chain or container breakout risk.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/whisper_local.py (reported line 87)May include surrounding context.

python
if container_status in ("exited", "created"):
            # 容器存在但未运行,启动它
            print(f"🔄 启动已有容器 {self.CONTAINER_NAME}...")
            result = subprocess.run(
                ["docker", "start", self.CONTAINER_NAME],
                capture_output=True, text=True, timeout=30
            )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/whisper_local.py (reported line 108)May include surrounding context.

python
self.DOCKER_IMAGE
        ]
        
        result = subprocess.run(cmd, capture_output=True, text=True, timeout=120)
        if result.returncode == 0:
            return True
        else:

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The example deployment command pulls and runs a third-party image using the mutable 'latest' tag, which can change over time and silently introduce malicious or compromised code. Because this skill also auto-starts and creates the same image elsewhere in the file, the supply-chain risk is real in context, not merely documentary.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The template hardcodes Chinese section titles such as 作者, 来源, 摘要, 正文, 要点, and 标签, which imposes a specific output language. The file does not indicate that this locale is optional, user-selected, or required for a justified region-specific compliance reason.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description says this is a 'Douyin video transcription suite' and repeatedly frames the trigger and scope around Douyin links, local files, and image notes. The SKILL.md edge-case guidance expands behavior to 'Other platforms' such as YouTube and Bilibili using yt-dlp, which is broader than the declared product scope.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.