Back to skill

Security audit

volcengine-video-generate

Security checks for vulnerabilities and agentic risk

Overview

This video-generation skill is mostly coherent, but it can read any user-supplied local file and send it to an external video API as a first-frame image without validating that it is actually an image.

Install only if you trust the environment and will control the arguments passed to the script. Do not pass sensitive local paths as the first-frame image; use a dedicated non-sensitive image file and a safe output directory, and keep Ark or Volcengine credentials in environment variables rather than prompts or source files.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/video_generate.py:31
Finding

Arbitrary Local File Disclosure Through Insufficient Image Validation

Content
View full analysis

Vulnerability Details

File Location: scripts/video_generate.py, lines 31–42 and 70–90
Vulnerability Type: Arbitrary local-file disclosure to an external API
Risk Level: Medium

Vulnerable Code:

python
def get_image_content(image_input: str) -> str:
    """
    Process image input. If it's a local file, convert to base64 data URI.
    Otherwise, assume it's a URL and return as is.
    """
    if os.path.isfile(image_input):
        try:
            mime_type, _ = mimetypes.guess_type(image_input)
            if not mime_type:
                # Fallback or default
                mime_type = "image/png"

            with open(image_input, "rb") as image_file:
                encoded_string = base64.b64encode(image_file.read()).decode("utf-8")
                return f"data:{mime_type};base64,{encoded_string}"
        except Exception as e:
            print(f"Failed to read or encode image file {image_input}: {e}")
            return None
    return image_input
python
if first_frame_image:
    image_url_or_base64 = get_image_content(first_frame_image)
    if image_url_or_base64:
        content.append(
            {"type": "image_url", "image_url": {"url": image_url_or_base64}}
        )

response = client.content_generation.tasks.create(
    model=model_name,
    content=content,
)

Technical Analysis

The first-frame argument is documented as an image path, but the implementation only checks whether the supplied path refers to a regular file. It does not verify that the file is actually an image.

mimetypes.guess_type() infers a type from the filename rather than the file contents. If no type is recognized, the code labels the file as image/png regardless of its actual format. It then reads and Base64-encodes the complete file and places it in a data URI submitted to the external Ark API.

Consequently, any regular file ...[truncated 1800 chars]

Remediation
View remediation

Remediation Suggestions

  1. Resolve local paths to canonical paths and restrict access to an explicitly approved workspace or upload directory.
  2. Reject files whose detected content type is not an approved image format; do not default unknown files to image/png.
  3. Validate file signatures and decode the input with a trusted image library before transmission.
  4. Reject symbolic links where appropriate, and verify the canonical path after resolution to prevent directory-boundary bypasses.
  5. Apply a conservative maximum file-size limit before reading the file into memory.
  6. Require explicit user confirmation before uploading a local file to an external service, especially in Agent-driven execution.
  7. Clearly document that local first-frame image contents are transmitted to the Ark API.
  8. Prefer a constrained file-selection interface over accepting unrestricted filesystem paths.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill invokes code with network access and reads environment-based credentials, but it does not declare any explicit tool scope or permissions boundary. This can cause operators or automated systems to underestimate the skill's access to secrets and external connectivity, increasing the chance of unsafe execution in broader-than-intended environments.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill allows a first-frame image to come from a local path or remote URL and notes it will be converted and sent for processing, but it does not clearly warn about privacy and data transfer implications. This may cause users to unintentionally upload sensitive local images or trigger retrieval of remote content they did not fully assess.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill states that it automatically downloads the generated video to a user-specified file, but it does not clearly warn that this performs a local write operation. Users may supply sensitive or unsafe paths without realizing the script will overwrite or create files, which can lead to accidental data loss or unsafe filesystem effects.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The authentication section references API keys and cloud credentials but does not include explicit handling guidance for sensitive secrets. This can lead users to expose credentials in logs, shells, prompts, or less secure execution contexts while setting up the skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

Most of the document is in English, but one example heading is presented in Chinese. This introduces an implicit language inconsistency without user opt-in or explanation, which can violate language/locale expectations for organizational content.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.