Back to skill

Security audit

Wan Video Generation and Editting

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a coherent Alibaba DashScope video tool, but it can upload arbitrary local files and voice samples without clear user consent or documentation.

Review before installing. Use this skill only with media and prompts you are authorized to send to Alibaba DashScope, avoid private files or internal URLs, and do not use reference voices unless you have explicit consent from the voice owner. Hosts should restrict or patch local-file upload behavior before broad agent use.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/wan27-magic.py:32
Finding

Undocumented Arbitrary Local File Encoding and Upload to DashScope

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (29)

Tainted flow: 'headers' from os.environ.get (line 470, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 97)May include surrounding context.

python
if args.audio_url:
        print(f"  Audio     : {args.audio_url}")

    resp = requests.post(url, headers=headers, json=payload, timeout=60)
    result = resp.json()

    if "code" in result:

Tainted flow: 'headers' from os.environ.get (line 470, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 213)May include surrounding context.

python
if args.audio_url:
        print(f"  Audio     : {args.audio_url}")

    resp = requests.post(url, headers=headers, json=payload, timeout=60)
    result = resp.json()

    if "code" in result:

Tainted flow: 'headers' from os.environ.get (line 470, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 339)May include surrounding context.

python
if args.audio_url:
        print(f"  Audio     : {args.audio_url}")

    resp = requests.post(url, headers=headers, json=payload, timeout=60)
    result = resp.json()

    if "code" in result:

Tainted flow: 'headers' from os.environ.get (line 470, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 446)May include surrounding context.

python
if args.audio_url:
        print(f"  Audio     : {args.audio_url}")

    resp = requests.post(url, headers=headers, json=payload, timeout=60)
    result = resp.json()

    if "code" in result:

Tainted flow: 'headers' from os.environ.get (line 470, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 125)May include surrounding context.

python
print(f"[text2video-get] Checking task: {args.task_id}")

    resp = requests.get(url, headers=headers, timeout=60)
    result = resp.json()

    if "code" in result:

Tainted flow: 'headers' from os.environ.get (line 470, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 241)May include surrounding context.

python
print(f"[text2video-get] Checking task: {args.task_id}")

    resp = requests.get(url, headers=headers, timeout=60)
    result = resp.json()

    if "code" in result:

Tainted flow: 'headers' from os.environ.get (line 470, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 367)May include surrounding context.

python
print(f"[text2video-get] Checking task: {args.task_id}")

    resp = requests.get(url, headers=headers, timeout=60)
    result = resp.json()

    if "code" in result:

Tainted flow: 'headers' from os.environ.get (line 470, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 474)May include surrounding context.

python
print(f"[text2video-get] Checking task: {args.task_id}")

    resp = requests.get(url, headers=headers, timeout=60)
    result = resp.json()

    if "code" in result:

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill advertises executable behavior that uses environment variables and external network access, but it does not declare an explicit tool/permission scope. That creates a transparency and least-privilege problem: users and hosting systems cannot easily assess or constrain what the skill may access, increasing the risk of unintended secret exposure or outbound requests beyond user expectations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation shows that user prompts, image URLs, video URLs, and audio URLs are sent to Alibaba ModelStudio APIs, but it does not clearly warn users that this data leaves the local environment and is processed by a third party. This can lead to accidental disclosure of sensitive prompts or private media, especially because URLs may point to user-controlled or confidential assets.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The documentation directs users to send prompts and externally hosted media URLs to a third-party video-generation endpoint, but it does not warn that this transfers potentially sensitive user data off-platform. In a media-generation skill, users may supply private images, videos, audio, or confidential prompts; absent disclosure increases the risk of unintended data sharing and poor consent practices.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The manifest describes the skill as supporting text2video, image2video, reference2video, and video-editing capabilities, but does not mention audio or voice features. This file's overview explicitly includes 'Voice cloning from reference', introducing a materially different capability beyond the declared video-focused scope.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documentation explicitly supports uploading reference voice media for voice cloning but provides no warning about consent, privacy, biometric sensitivity, or lawful use. Because voiceprints are sensitive personal data and cloning can enable impersonation, fraud, or non-consensual synthesis, omission of these safeguards increases the likelihood of misuse in real deployments.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The documentation explicitly supports supplying an external audio URL to a third-party video-generation API but does not warn that user-provided media references and associated content may be transmitted to Alibaba Cloud/DashScope for processing. In a skill centered on media generation, this omission can lead to accidental disclosure of private or copyrighted audio sources, especially if users assume URLs are only referenced locally or do not appreciate the privacy implications.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation shows users providing externally hosted media URLs to a third-party Alibaba DashScope endpoint but does not explicitly warn that videos and reference images are sent to an external service for processing. In a video-editing skill, users may upload sensitive or proprietary visual content, so omission of this disclosure can lead to unintended data sharing and privacy/compliance issues.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Although the table notes that auto may regenerate audio, the documentation does not frame this behavior as a user-facing warning despite media-editing workflows often expecting original audio preservation. This can cause accidental replacement or synthesis of speech/sounds, creating privacy, integrity, or content-authenticity issues if users assume the original soundtrack will remain untouched.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

encode_local_file() silently reads arbitrary local files, base64-encodes them, and includes their contents in outbound API requests. In a skill context, this is privacy-sensitive because users may provide local media paths without clearly realizing the files are uploaded off-host to a third-party service, increasing the risk of unintended data disclosure.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

This POST transmits user prompts and potentially local audio content to an external cloud API. In this skill, the transfer is core functionality, but it remains security-relevant because users may unintentionally send sensitive prompts or media off-device, and the code provides no explicit privacy warning or consent mechanism beyond implicit execution.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 97)May include surrounding context.

python
if args.audio_url:
        print(f"  Audio     : {args.audio_url}")

    resp = requests.post(url, headers=headers, json=payload, timeout=60)
    result = resp.json()

    if "code" in result:

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

This request sends image/video/audio inputs and prompts to a remote API, including base64-encoded local files when paths are provided. The risk is unintentional external disclosure of potentially sensitive local media, made more relevant in an agent skill where users may assume processing is local unless told otherwise.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 213)May include surrounding context.

python
print(f"  Resolution : {args.resolution}")
    print(f"  Duration   : {args.duration}s")

    resp = requests.post(url, headers=headers, json=payload, timeout=60)
    result = resp.json()

    if "code" in result:

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

The script uploads reference images/videos and optional voice samples to an external vendor service. Because these inputs can contain highly sensitive biometric, identity, or proprietary content, the lack of explicit privacy/consent messaging makes this more dangerous in context than a generic outbound request.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 339)May include surrounding context.

python
print(f"  Ratio      : {args.ratio}")
    print(f"  Duration   : {args.duration}s")

    resp = requests.post(url, headers=headers, json=payload, timeout=60)
    result = resp.json()

    if "code" in result:

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

This POST uploads an input video and optional reference images to a third-party API for editing. That external transmission is intentional, but without explicit disclosure it can cause inadvertent leakage of sensitive video content, which is especially relevant for a general-purpose agent skill handling local user files.

Content

Scanner excerpt · scripts/wan27-magic.py (reported line 446)May include surrounding context.

python
print(f"  Duration   : {args.duration}s")
    print(f"  Audio      : {args.audio_setting}")

    resp = requests.post(url, headers=headers, json=payload, timeout=60)
    result = resp.json()

    if "code" in result:

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The reference-to-video prompt instructions require users to refer to assets using Chinese labels like '视频1' and '图片1'. This imposes a specific language/locale convention in the natural-language interface without opt-in or justification, which can conflict with language-choice policies.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

Several example prompts are written in Chinese, which can imply a language expectation for users of the skill. The document does not state that prompt language is optional or that the API supports multiple languages, nor does it justify a Chinese-only locale requirement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

Several example prompts are written in Chinese, which can imply a language expectation for users of the skill. The document does not state that prompt language is optional or that the API supports multiple languages, nor does it justify a Chinese-only locale requirement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

Several example prompts are written in Chinese, which can imply a language expectation for users of the skill. The document does not state that prompt language is optional or that the API supports multiple languages, nor does it justify a Chinese-only locale requirement.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.