Back to skill

Security audit

YM-MediaToolkit(媒体处理工具集)

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed media-processing toolkit, but its unauthenticated local HTTP API can process local files and arbitrary remote URLs in ways that need careful review before use.

Install only in a trusted local environment. Do not expose the HTTP server to a network, avoid --host 0.0.0.0 unless you add authentication and restricted CORS, do not accept media_roots from untrusted callers, and prefer pinned dependencies plus resource limits before production use.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (5)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
utils.py:220
Finding

Redirect and DNS Resolution Gaps Allow Server-Side Request Forgery

Content
View full analysis
str: validate_video_url(url) import requests temp_file = tempfile.NamedTemporaryFile(suffix='.mp4', delete=False) temp_path = temp_file.name temp_file.close() response = requests.get(url, stream=True, timeout=timeout) with open(temp_path, 'wb') as f: for chunk in response.iter_content(chunk_size=8192): f.write(chunk) return temp_path ``` The remote frame extractor performs the same initial-only validation: ```python def __init__(self, video_url: str, timeout: int = 30): validate_video_url(video_url) self.video_url = video_url self.timeout = timeout self.session = requests.Session() ``` ```python def _get_file_size(self) -> int: response = self.session.head(self.video_url, timeout=self.timeout) if 'content-length' in response.headers: return int(response.headers['content-length']) response = self.session.get(self.video_url, stream=True, timeout=self.timeout) return int(response.headers.get('content-length', 0)) def _download_range(self, start: int, end: int) -> bytes: headers = {'Range': f'bytes={start}-{end}'} response = self.session.get(self.video_url, headers=headers, timeout=self.timeout) if response.status_code in [200, 206]: return response.content raise Exception(f"HTTP Range request failed: {response.status_code}") ``` FFmpeg-backed processing similarly passes the validated URL to a separate network client: ```python cmd = ['ffmpeg', '-i', video_url] ``` ### Technical Analysis `validate_video_url()` rejects an initial hostname if ...[truncated 2245 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
utils.py:106
Finding

Request-Controlled Media Roots Permit Arbitrary Local File Access

Content
View full analysis
list: if media_roots is None: env_roots = os.environ.get('YM_MEDIA_ROOTS') if env_roots: media_roots = [root for root in re.split(r'[;,]', env_roots) if root.strip()] else: media_roots = [default_dir] elif isinstance(media_roots, str): media_roots = [root for root in re.split(r'[;,]', media_roots) if root.strip()] elif not media_roots: media_roots = [default_dir] return [Path(root).expanduser().resolve() for root in media_roots] ``` ```python def validate_media_source(source: str, default_dir: str = '.', media_roots=None) -> str: parsed = urlparse(source) if parsed.scheme in ('http', 'https'): validate_video_url(source) return source if parsed.scheme and not _is_windows_drive_path(source): raise ValueError( f"Forbidden protocol: '{parsed.scheme}'; only http/https or local paths are allowed" ) p = Path(source) if any(part == '..' for part in p.parts): raise ValueError(f"Path traversal is forbidden: {source}") resolved = p.resolve() roots = get_media_roots(media_roots=media_roots, default_dir=default_dir) allowed = False for root in roots: try: resolved.relative_to(root) allowed = True break except ValueError: continue if not allowed: roots_text = ', '.join(str(root) for root in roots) raise ValueError( f"Local input is outside allowed media_roots: {source} " f"(allowed: {roots_text})" ) if not resolved.exists(): raise ValueError(f"Local inp ...[truncated 2579 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
run.py:1030
Finding

Unauthenticated Cross-Origin HTTP API Exposes File and Media Operations

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Open-Ended Dependency Constraints Produce Non-Reproducible and Unsafe Installations

Content
View full analysis
=2.28.0 opencv-python>=4.8.0 numpy>=1.24.0 aiohttp>=3.8.0 flask>=2.3.0 flask-cors>=4.0.0 faster-whisper>=1.0.0 paddlepaddle>=2.6.0 paddleocr>=2.7.0 ``` The documented installation procedure is: ```bash pip install -r requirements.txt ``` ### Technical Analysis Every dependency uses an open-ended lower-bound constraint. Consequently, installation may select any future direct dependency release and any compatible transitive dependency release. There is no lockfile, exact version constraint, or package hash to ensure that an audited set of artifacts is installed. Python packages may execute code while being built or imported. An upstream account compromise, malicious future release, dependency takeover, or unexpected transitive dependency change can therefore alter the effective code installed with the Skill without changing the reviewed repository. The issue also prevents reproducible builds and makes it difficult to associate the deployed environment with the dependency versions covered by testing. ### Attack Path 1. An operator follows the documented `pip install -r requirements.txt` instruction. 2. The package resolver queries configured package indexes. 3. It selects current versions satisfying the open-ended `>=` constraints. 4. A newly published, compromised, or behaviorally incompatible direct or transitive dependency is selected. 5. Package code executes during build, installation, import, or normal Skill operation under the installer or service account. 6. The deployed behavior differs from the audited repository and tested environment. ### Impact Assessment The impact depends on the privileges used during installation and execution. A compromised package could access files, credentials, ...[truncated 273 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
job_manager.py:64
Finding

Unbounded Job Submission and Remote Media Processing Enable Resource Exhaustion

Content
View full analysis
dict: if action not in self.action_handlers: return protocol_error('invalid_action', f'Invalid action: {action}') if params is None: params = {} if not isinstance(params, dict): return protocol_error('invalid_params', 'params must be an object') if metadata is None: metadata = {} if not isinstance(metadata, dict): metadata = {} job_id = uuid.uuid4().hex job_path = self._job_path(job_id) job = { 'job_id': job_id, 'action': action, 'params': params, 'created_by': metadata.get('created_by'), 'intent': metadata.get('intent'), 'source': metadata.get('source'), 'metadata': metadata, 'status': 'queued', 'code': 'ok', 'reply': f'Job submitted: {job_id}', 'hint': 'Use poll_url to query job status.', 'created_at': utc_now(), 'started_at': None, 'finished_at': None, 'result': None, 'output_paths': [], 'error': None, } self._write_job(job) self.queue.put(job_id) ``` The range reader accepts a ...[truncated 2247 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (70)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

该代码块的核心功能是音频提取与音频信息探测,而不是自然语言媒体助手的综合能力实现。虽然“音频转换”与实际行为部分相关,但声明还强调视频压缩、MP4/MOV封面提取、字幕识别,这些在本代码中均未出现。相反,代码明确具备从远程URL直接流式读取视频并提取音频、批量处理多个视频、以及使用 ffprobe 分析音频流信息的能力,这些属于更具体且未在描述中明确说明的能力。整体上,描述与该代码块的实际职责存在明显偏差,因此应判定为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

该代码块没有体现视频压缩、MP4/MOV封面提取或音频转换能力,也没有执行语音转字幕的“识别”流程。它处理的是已有字幕数据:安全读取JSON、规范化受保护词、按长度和语言线索切分字幕、重新分配每段起止时间,并输出新的字幕JSON。虽然这可被视为字幕处理相关的辅助能力,但与声明的主要功能集合相比,实际代码的核心用途明显不同且未被描述,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个综合型“自然语言媒体助手”,包含多项媒体处理功能。但提供的代码只聚焦于视频帧/封面提取,且主要特色是对远程MP4视频进行部分下载、解析moov/trak/stbl等结构并解码目标帧,再保存为JPG。代码中没有自然语言交互逻辑,也没有视频压缩、音频转换或字幕识别相关实现。因此描述明显宽于且偏离代码实际功能,属于能力声明与实际行为不一致。虽然“MP4/MOV封面提取”与代码部分吻合,但其余核心宣称能力均未体现。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

声明与代码有较大重合:都围绕自然语言媒体处理,且包含压缩、封面/缩略图提取、音频处理、字幕识别。但代码还实现了若干未在描述中体现的能力:1) 查询媒体信息和音频信息;2) 接收结构化JSON动作与多步骤pipeline;3) 支持更广泛的媒体扩展名,而非仅MP4/MOV。虽然这些能力仍属媒体助手范畴,但已构成未声明的实际功能,因此应判定为存在描述与行为不完全一致的情况。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

该代码块的核心职责是作业调度与生命周期管理,而不是描述中所说的媒体处理功能。虽然 action_handlers 可能在别处挂接媒体操作,但此代码本身没有实现视频压缩、封面提取、音频转换或字幕识别逻辑,也未体现与这些媒体能力直接相关的资源访问或处理流程。相反,它新增了未在描述中体现的重要能力:任务提交、后台执行、结果持久化、轮询、清理和错误恢复。因此,描述与实际代码行为存在明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

代码与声明存在明显描述不一致。声明描述了一个包含视频压缩、封面提取、音频转换和字幕识别的多功能媒体助手,但该代码片段仅实现“字幕识别/提取”相关功能:通过 ASR 和 OCR 获取字幕,进行清洗、融合并输出 JSON。虽然“字幕识别”这一项与声明部分一致,但其余被强调的核心能力在代码中完全没有体现,因此当前代码块的实际行为只覆盖声明中的一部分,且主功能范围明显更窄,属于实质性不匹配。

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 48)May include surrounding context.

md
字幕识别依赖已内置在 `requirements.txt`,包括 `faster-whisper`、`paddlepaddle`、`paddleocr`。

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 138)May include surrounding context.

md
字幕识别依赖已内置在 `requirements.txt`,包括 `faster-whisper`、`paddlepaddle`、`paddleocr`。

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 265)May include surrounding context.

md
字幕识别依赖已内置在 `requirements.txt`,包括 `faster-whisper`、`paddlepaddle`、`paddleocr`。

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The maintenance guide specifies that the reply field must be a Chinese response suitable for chat display. This imposes a fixed language requirement in natural-language behavior without documenting user choice, opt-in, or a justified region-specific constraint, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill accepts remote http/https media URLs but does not prominently warn at the point of feature introduction that providing a URL triggers outbound network access and retrieval of attacker-controlled content. This matters because network fetches expand the trust boundary and can expose the host to malicious media, privacy leakage, or SSRF-like misuse if URL validation is weak in implementation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill states that file-writing actions default to overwrite=true, which can silently replace existing media outputs and job artifacts if callers reuse paths or names. In automation or chat-driven contexts, this increases the risk of unintended data loss because users may not realize destructive writes are the default.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 192)May include surrounding context.

HTTP / Claw 长任务推荐:

bash
curl -X POST http://127.0.0.1:8080/skill/chat \
  -H 'Content-Type: application/json' \
  -d '{"message":"识别 \"sample.mp4\" 的字幕","async":"auto"}'

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 225)May include surrounding context.

提交任务:

bash
curl -X POST http://127.0.0.1:8080/skill/jobs \
  -H 'Content-Type: application/json' \
  -d '{"action":"pipeline","params":{"source":"sample.mp4","steps":[{"id":"metadata","action":"info","enabled":true}]}}'

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code returns a Chinese-only error string to users, which imposes a specific language regardless of user preference. The policy allows locale constraints only when explicitly justified or when users are given a choice, neither of which is present here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The returned error text forces Chinese output for engine load failures, which is a natural-language policy issue when no user opt-in or documented regional scope exists. This can violate language/locale requirements for general-purpose skills.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · asr_engine.py (reported line 51)May include surrounding context.

python
'-y',
                audio_path,
            ]
            ffmpeg_result = subprocess.run(cmd, capture_output=True, text=True, timeout=300)
            if ffmpeg_result.returncode != 0:
                return {
                    'status': 'error',

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This user-visible error string is emitted only in Chinese, with no indication that the skill is limited to Chinese-speaking users or that localization is supported. That makes the skill force a specific language in a general code path.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes a media assistant for video compression, cover extraction, audio conversion, and subtitle recognition, but this module explicitly supports operating on remote videos directly via URL. Code uses HTTP HEAD requests and invokes ffmpeg/ffprobe against remote URLs, adding network-retrieval behavior not stated in the manifest description.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module-level natural-language description is entirely in Chinese and does not indicate that users may choose another language or locale. This can violate language/locale policy when a skill implicitly forces a specific language without documented opt-in or justification.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The code performs outbound requests against user-supplied URLs using requests.head and later passes the same remote source to ffmpeg/ffprobe, which can enable SSRF-style access to internal services or other unintended network destinations if validation is insufficient. In a media-processing skill, remote URL support is plausible, but it is more dangerous here because the component acts as a network-capable proxy and parser for attacker-controlled endpoints.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · audio_extractor.py (reported line 135)May include surrounding context.

python
logger.info(f"命令: {' '.join(cmd[:5])}...")
    
    try:
        result = subprocess.run(
            cmd,
            capture_output=True,
            text=True,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · audio_extractor.py (reported line 277)May include surrounding context.

python
cmd = ['ffprobe', '-v', 'quiet', '-print_format', 'json', '-show_streams', video_url]
    
    try:
        result = subprocess.run(cmd, capture_output=True, text=True, timeout=30)
        if result.returncode != 0:
            return {'status': 'error', 'message': 'ffprobe failed'}

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Several error messages are hard-coded in Chinese, such as the path traversal, out-of-directory, and missing-file exceptions. This imposes a specific language on users without offering a locale choice or documenting that the tool is intentionally Chinese-only.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.