T09 · Insecure Skill Coding Practices
- Location
scripts/douyin_downloader.py:85- Finding
Unrestricted URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This Douyin downloader/transcriber is mostly coherent, but it can make unvalidated outbound web requests and upload audio to a remote transcription service with limited user-facing safeguards.
Review before installing. Use this only in an isolated environment and only with trusted Douyin links. Avoid transcribing sensitive, private, or copyrighted audio unless you are comfortable sending it to SiliconFlow. Prefer pinned dependencies and consider adding domain allowlists, request timeouts, size limits, and clearer API-key documentation before regular use.
scripts/douyin_downloader.py:85Unrestricted URL Fetching Enables Server-Side Request Forgery
scripts/douyin_downloader.py:256Configurable Transcription Endpoint Can Receive API Credentials and Private Audio
SKILL.md:22Unpinned Third-Party Dependency Installation
The downloader fetches video_info['url'] and that URL is derived from remote HTML/JSON content, not from a trusted local allowlist. If the upstream page or parsing is manipulated, the script can be induced to download arbitrary URLs, enabling SSRF, unexpected large downloads, or retrieval of malicious files from attacker-chosen infrastructure.
if show_progress:
print(f"正在下载视频: {video_info['title']}")
response = requests.get(video_info['url'], headers=HEADERS, stream=True)
response.raise_for_status()
# 获取文件大小
The script uploads extracted audio content to an external transcription API, which is an intentional data exfiltration path to a third party. In the skill context this may include sensitive spoken content from downloaded media, and users are not forced through an explicit consent step or warned at the point of transmission.
}
try:
response = requests.post(self.api_base_url, files=files, headers=headers)
response.raise_for_status()
result = response.json()
The skill documentation describes capabilities that access environment variables, write files, and perform network requests, but it does not declare any explicit tool scope or permission boundaries. This can cause the host or user to invoke the skill without understanding its operational reach, increasing the risk of unintended data access, file writes, or outbound requests.
The skill notes that an API key is needed for transcript extraction, but it does not clearly warn in the main description or usage flow that video-derived audio/content is transmitted to an external speech-recognition service. Users may assume processing is local and inadvertently send sensitive or copyrighted material to a third party without informed consent.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
}
# 硅基流动 API 配置
DEFAULT_API_BASE_URL = "https://api.siliconflow.cn/v1/audio/transcriptions"
DEFAULT_MODEL = "FunAudioLLM/SenseVoiceSmall"
The script extracts the first URL from arbitrary user-supplied text and immediately fetches it with requests.get, without validating the hostname or scheme beyond a regex match. In an agent/server context this can be abused as SSRF to make the host perform outbound requests to attacker-controlled or internal endpoints before any Douyin-specific normalization occurs.
raise ValueError("未找到有效的分享链接")
share_url = urls[0]
share_response = requests.get(share_url, headers=HEADERS)
video_id = share_response.url.split("?")[0].strip("/").split("/")[-1]
share_url = f'https://www.iesdouyin.com/share/video/{video_id}'
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
share_url = f'https://www.iesdouyin.com/share/video/{video_id}'
# 获取视频页面内容
response = requests.get(share_url, headers=HEADERS)
response.raise_for_status()
pattern = re.compile(
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
# 下载文件
downloaded = 0
with open(filepath, 'wb') as f:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
Audio derived from downloaded videos is transmitted to a third-party transcription service without an explicit user-facing warning at the transmission point. In this skill context the feature is expected, but it still creates privacy and compliance risk because spoken content may contain personal, copyrighted, or confidential information.
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
# 保存文案为 Markdown 格式
transcript_path = video_folder / "transcript.md"
with open(transcript_path, 'w', encoding='utf-8') as f:
f.write(f"# {video_info['title']}\n\n")
f.write(f"| 属性 | 值 |\n")
f.write(f"|------|----|\n")
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
# 保存视频 (可选)
if save_video:
saved_video_path = video_folder / f"{video_info['video_id']}.mp4"
shutil.copy2(video_path, saved_video_path)
if show_progress:
print(f"视频已保存到: {saved_video_path}")
The module docstring says the transcription feature uses the API_KEY environment variable, but later user-facing code and one runtime error message instruct users to set DOUYIN_API_KEY, while the actual lookup in extract_text uses API_KEY. This is an active documentation contradiction about how the skill is configured, not merely an omission.
extract_text_from_audio raises an error telling users to set DOUYIN_API_KEY, but the higher-level extraction flow populates self.api_key from API_KEY or the --api-key argument. This contradiction misrepresents how the feature is configured and may lead to failed use despite following the message.
The code accesses a sensitive credential via os.getenv('API_KEY') to authenticate with an external API. Although the docstring documents the environment variable, it does not clearly disclose to users at the point of use that the credential will be used for remote requests to a third-party service.
The CLI epilog and --api-key help state that users can configure the API key with DOUYIN_API_KEY, but extract_text actually reads os.getenv('API_KEY'). This directly misstates the skill's behavior and can cause users to believe the skill is using one configuration source when it is using another.
No suspicious patterns detected.