T09 · Insecure Skill Coding Practices
- Location
scripts/wan27-magic.py:32- Finding
Undocumented Arbitrary Local File Encoding and Upload to DashScope
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is mostly a coherent Alibaba DashScope video tool, but it can upload arbitrary local files and voice samples without clear user consent or documentation.
Review before installing. Use this skill only with media and prompts you are authorized to send to Alibaba DashScope, avoid private files or internal URLs, and do not use reference voices unless you have explicit consent from the voice owner. Hosts should restrict or patch local-file upload behavior before broad agent use.
scripts/wan27-magic.py:32Undocumented Arbitrary Local File Encoding and Upload to DashScope
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
if args.audio_url:
print(f" Audio : {args.audio_url}")
resp = requests.post(url, headers=headers, json=payload, timeout=60)
result = resp.json()
if "code" in result:
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
if args.audio_url:
print(f" Audio : {args.audio_url}")
resp = requests.post(url, headers=headers, json=payload, timeout=60)
result = resp.json()
if "code" in result:
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
if args.audio_url:
print(f" Audio : {args.audio_url}")
resp = requests.post(url, headers=headers, json=payload, timeout=60)
result = resp.json()
if "code" in result:
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
if args.audio_url:
print(f" Audio : {args.audio_url}")
resp = requests.post(url, headers=headers, json=payload, timeout=60)
result = resp.json()
if "code" in result:
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
print(f"[text2video-get] Checking task: {args.task_id}")
resp = requests.get(url, headers=headers, timeout=60)
result = resp.json()
if "code" in result:
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
print(f"[text2video-get] Checking task: {args.task_id}")
resp = requests.get(url, headers=headers, timeout=60)
result = resp.json()
if "code" in result:
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
print(f"[text2video-get] Checking task: {args.task_id}")
resp = requests.get(url, headers=headers, timeout=60)
result = resp.json()
if "code" in result:
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
print(f"[text2video-get] Checking task: {args.task_id}")
resp = requests.get(url, headers=headers, timeout=60)
result = resp.json()
if "code" in result:
The skill advertises executable behavior that uses environment variables and external network access, but it does not declare an explicit tool/permission scope. That creates a transparency and least-privilege problem: users and hosting systems cannot easily assess or constrain what the skill may access, increasing the risk of unintended secret exposure or outbound requests beyond user expectations.
The documentation shows that user prompts, image URLs, video URLs, and audio URLs are sent to Alibaba ModelStudio APIs, but it does not clearly warn users that this data leaves the local environment and is processed by a third party. This can lead to accidental disclosure of sensitive prompts or private media, especially because URLs may point to user-controlled or confidential assets.
The documentation directs users to send prompts and externally hosted media URLs to a third-party video-generation endpoint, but it does not warn that this transfers potentially sensitive user data off-platform. In a media-generation skill, users may supply private images, videos, audio, or confidential prompts; absent disclosure increases the risk of unintended data sharing and poor consent practices.
The manifest describes the skill as supporting text2video, image2video, reference2video, and video-editing capabilities, but does not mention audio or voice features. This file's overview explicitly includes 'Voice cloning from reference', introducing a materially different capability beyond the declared video-focused scope.
The documentation explicitly supports uploading reference voice media for voice cloning but provides no warning about consent, privacy, biometric sensitivity, or lawful use. Because voiceprints are sensitive personal data and cloning can enable impersonation, fraud, or non-consensual synthesis, omission of these safeguards increases the likelihood of misuse in real deployments.
The documentation explicitly supports supplying an external audio URL to a third-party video-generation API but does not warn that user-provided media references and associated content may be transmitted to Alibaba Cloud/DashScope for processing. In a skill centered on media generation, this omission can lead to accidental disclosure of private or copyrighted audio sources, especially if users assume URLs are only referenced locally or do not appreciate the privacy implications.
The documentation shows users providing externally hosted media URLs to a third-party Alibaba DashScope endpoint but does not explicitly warn that videos and reference images are sent to an external service for processing. In a video-editing skill, users may upload sensitive or proprietary visual content, so omission of this disclosure can lead to unintended data sharing and privacy/compliance issues.
Although the table notes that auto may regenerate audio, the documentation does not frame this behavior as a user-facing warning despite media-editing workflows often expecting original audio preservation. This can cause accidental replacement or synthesis of speech/sounds, creating privacy, integrity, or content-authenticity issues if users assume the original soundtrack will remain untouched.
encode_local_file() silently reads arbitrary local files, base64-encodes them, and includes their contents in outbound API requests. In a skill context, this is privacy-sensitive because users may provide local media paths without clearly realizing the files are uploaded off-host to a third-party service, increasing the risk of unintended data disclosure.
This POST transmits user prompts and potentially local audio content to an external cloud API. In this skill, the transfer is core functionality, but it remains security-relevant because users may unintentionally send sensitive prompts or media off-device, and the code provides no explicit privacy warning or consent mechanism beyond implicit execution.
if args.audio_url:
print(f" Audio : {args.audio_url}")
resp = requests.post(url, headers=headers, json=payload, timeout=60)
result = resp.json()
if "code" in result:
This request sends image/video/audio inputs and prompts to a remote API, including base64-encoded local files when paths are provided. The risk is unintentional external disclosure of potentially sensitive local media, made more relevant in an agent skill where users may assume processing is local unless told otherwise.
print(f" Resolution : {args.resolution}")
print(f" Duration : {args.duration}s")
resp = requests.post(url, headers=headers, json=payload, timeout=60)
result = resp.json()
if "code" in result:
The script uploads reference images/videos and optional voice samples to an external vendor service. Because these inputs can contain highly sensitive biometric, identity, or proprietary content, the lack of explicit privacy/consent messaging makes this more dangerous in context than a generic outbound request.
print(f" Ratio : {args.ratio}")
print(f" Duration : {args.duration}s")
resp = requests.post(url, headers=headers, json=payload, timeout=60)
result = resp.json()
if "code" in result:
This POST uploads an input video and optional reference images to a third-party API for editing. That external transmission is intentional, but without explicit disclosure it can cause inadvertent leakage of sensitive video content, which is especially relevant for a general-purpose agent skill handling local user files.
print(f" Duration : {args.duration}s")
print(f" Audio : {args.audio_setting}")
resp = requests.post(url, headers=headers, json=payload, timeout=60)
result = resp.json()
if "code" in result:
The reference-to-video prompt instructions require users to refer to assets using Chinese labels like '视频1' and '图片1'. This imposes a specific language/locale convention in the natural-language interface without opt-in or justification, which can conflict with language-choice policies.
Several example prompts are written in Chinese, which can imply a language expectation for users of the skill. The document does not state that prompt language is optional or that the API supports multiple languages, nor does it justify a Chinese-only locale requirement.
Several example prompts are written in Chinese, which can imply a language expectation for users of the skill. The document does not state that prompt language is optional or that the API supports multiple languages, nor does it justify a Chinese-only locale requirement.
Several example prompts are written in Chinese, which can imply a language expectation for users of the skill. The document does not state that prompt language is optional or that the API supports multiple languages, nor does it justify a Chinese-only locale requirement.
No suspicious patterns detected.