Back to skill

Security audit

Byted Kickart Video Subtitler

Security checks for vulnerabilities and agentic risk

Overview

The skill can support remote video subtitling, but it handles cloud credentials and updates in ways that need careful review before installation.

Review this skill before installing. Do not paste long-lived cloud AK/SK credentials into chat, and do not run any update command returned by the service unless you can independently verify it. Prefer scoped temporary credentials, sanitized logs, pinned dependencies, and a private skill-specific cache directory before using it with sensitive media or accounts.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/core/api/meida/chunks.py:117
Finding

Authentication credentials are exposed through command output and HTTP logging

Content
View full analysis
>> {method.upper()} {url} {headers} {body}") response = requests.request( method=method.upper(), url=url, headers=headers, data=body, timeout=30 ) logging.info(f"<<< {response.headers} {response.text}") ``` ```python headers["Authorization"] = f"Bearer {self.token}" ... logging.info(f">>> {method.upper()} {url} {headers} {body}") response = requests.request( method=method.upper(), url=url, headers=headers, data=body, timeout=30 ) logging.info(f"<<< {response.headers} {response.text}") ``` ### Technical Analysis The skill instructions explicitly require printing all supported authentication secrets, including the bearer API key and the long-term secret access key. This unnecessarily places credentials in command output, which may be retained in agent transcripts, execution telemetry, shell history, or orchestration logs. The HTTP clients also serialize the complete request header mapping into application logs. For bearer authentication, this records the reusable token verbatim. For AK/SK authentication, the log contains the access-key identifier and a valid request signature. Although a signature is less reusable than the secret key itself, recording signed requests still expands the exposure sur ...[truncated 1613 chars]
Remediation
View remediation

T03 · Remote Payload Retrieval and Execution

Error
Location
SKILL.md:147
Finding

Server-provided update command can be executed without authenticity or allowlist validation

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/core/utils/downloader.py:222
Finding

User-controlled media URLs can trigger unrestricted server-side requests and unbounded downloads

Content
View full analysis
" ``` ```python async with session.get(url, timeout=timeout) as response: if response.status != 200: return DownloadResult( url=url, success=False, error=f"HTTP {response.status}" ) filename = self.filename_generator.generate(url) file_path = os.path.join(self.output, filename) with open(file_path, "wb") as f: async for chunk in response.content.iter_chunked(8192): f.write(chunk) ``` ### Technical Analysis The documented workflow accepts a user-provided “public URL” and follows redirects with `curl -L`, but it does not enforce HTTPS, validate the resolved address, reject loopback/private/link-local networks, or revalidate every redirect destination. The Python downloader similarly performs arbitrary `GET` requests without scheme or destination restrictions. Its timeout limits request duration but does not enforce a maximum response size. Content is streamed until the server ends the response, permitting disk exhaustion. The subsequent file-type check does not prevent the initial network request and therefore cannot mitigate SSRF. Downloading a user-provided video is part of the declared functionality, but access to internal addresses and arbitrary response sizes exceeds the minimum network and storage privileges required. ### Attack Path 1. An attacker supplies a URL targeting an internal endpoint, such as a loopback service, RFC1918 address, or cloud metadata address. 2. The skill executes `curl -L` or the downloader issues a request from the agent environment. 3. The internal service receives ...[truncated 893 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/requirements.txt:1
Finding

Mutable dependency ranges are installed automatically during every workflow

Content
View full analysis
=2.31.0 qrcode>=8.2 jsonpath>=0.82.2 Pillow>=10.1.0 urllib3>=2.1.0 pydantic==2.12.5 pandas==2.3.3 python-dotenv>=1.1.1 click>=8.3.2 ``` ### Technical Analysis Seven of the nine direct dependencies use open-ended lower bounds. The mandatory prerequisite workflow installs these packages at runtime. Consequently, a future release satisfying the range can be installed without having been audited with this skill. The requirements file also lacks artifact hashes and does not constrain package indexes to an approved repository. Dependency resolution may additionally select mutable transitive dependency versions. No specific package in the reviewed file was proven malicious. The confirmed security defect is the non-reproducible, mutable installation process, which creates an avoidable supply-chain execution channel. ### Attack Path 1. The user invokes the skill. 2. The mandatory prerequisite process runs `pip install`. 3. The resolver selects the newest versions satisfying the open-ended constraints. 4. A compromised publisher account, malicious future release, dependency-confusion source, or compromised package index supplies an altered artifact. 5. Installation or later import executes attacker-controlled package code in the skill environment. 6. That code can access the same files, credentials, and network resources as the skill process. ### Impact Assessment A compromised dependency can execute arbitrary Python code with the runner's privileges. Accessible assets include Volcengine credentials, user video and subtitle files, task results, logs, and other files visible to the agent environment. ]]>
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/core/__init__.py:48
Finding

Signed media metadata is stored in a shared predictable cross-skill directory

Content
View full analysis
Path: return self.base_dir / f"{media_id}.json" ... with open(path, "w", encoding="utf-8") as f: json.dump(data, f, ensure_ascii=False, indent=2) ``` ```python media_info = { "id": matriel.id, "url": matriel.url, "type": matriel.type, "width": matriel.width, "height": matriel.height, "size": matriel.size, "timestamp": time.time(), } ... self.repository.save(matriel.id, media_info) ``` ### Technical Analysis The subtitling skill stores media metadata under a directory named for a different skill, `byted-kickart-viral-replicator`. The cache includes the remote media URL. The skill documentation states that produced video URLs can contain authentication parameters, so such URLs must be treated as bearer-like sensitive data. Files are created using default process permissions, with no explicit restrictive mode, lifecycle expiration, ownership verification, or cleanup after task completion. The predictable shared path creates an unnecessary cross-skill boundary and increases the chance that another component can read or overwrite cached media metadata. Local caching is used by `subtitler.py` to look up the uploaded media, but indefinite storage in another skill's namespace is not required for that purpose. ### Attack Path 1. A user uploads a video. 2. `SimpleMediaService` writes the media ID and remote URL to a predictable JSON file in the unrelated `byted-kickart-viral-replicator` directory. 3. Another local skill or process with access to `/tmp/openclaw` enumerates or reads the cache files. 4. The process ob ...[truncated 592 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (87)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose is a concrete user-facing capability: adding or embedding subtitles into videos. The actual code chunk does not perform that task or any clear part of it beyond generic infrastructure. It sets constants for media extensions and size limits, defines a Result schema, and initializes logging/directories under /tmp. Those are supporting utilities at best, but in isolation they do not substantiate the declared functionality. Because the code shown has a materially different primary behavior (generic initialization/configuration) and lacks any subtitle-related processing, this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的核心能力是“给视频添加/嵌入字幕”,应包含视频文件处理、字幕生成或写入视频等逻辑。但提供的代码仅是通用远程 API 客户端层:从环境变量读取 ACCESS_KEY_ID、SECRET_ACCESS_KEY、ARK_SKILL_API_KEY、ARK_SKILL_API_BASE,构造带签名或 Bearer Token 的 HTTP 请求,调用 requests.request 访问外部服务并记录日志。该代码片段没有体现任何与视频、音轨、字幕文件、转录、封装或字幕烧录相关的行为。虽然这可能是某更大系统中的底层支撑模块,但就该代码块本身而言,其实际行为与声明用途明显不一致,属于材质性目的不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says this skill adds or embeds subtitles into videos. However, the supplied code does not process videos, subtitles, subtitle files, media tracks, or video rendering. Instead, it is a general service-layer API client for submitting and querying AI template tasks through ICCP endpoints. The submit method even uses a hardcoded image URL in ResourceList, which is inconsistent with a video-subtitle workflow. This is a material purpose mismatch, not just an implementation detail.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明的核心能力是“给视频添加/嵌入字幕”。但代码中没有任何与字幕相关的处理,例如语音识别生成字幕、读取/写入 SRT/VTT/ASS 文件、时间轴字幕对齐、FFmpeg 字幕烧录/封装、字幕轨添加等。相反,这段代码主要是媒体平台通用基础设施:根据不同认证方式构造 HTTP 请求,调用 ListUsers、GetUploadState、StreamUploadData、CreateMaterial、GetMediaInfo 等接口,完成媒体素材上传、创建、状态轮询和结果格式化;并且支持 image/video/audio 多类素材,而非专门针对视频字幕。故该代码行为与技能描述的主要目的存在明显不一致。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的核心能力是“给视频添加/嵌入字幕”,这通常需要字幕文本生成或读取、时间轴处理、字幕轨合成/烧录、视频导出等操作。而实际代码只实现了媒体文件管理相关能力:计算文件哈希、调用远程 Muse/Kickart 服务上传图片或视频、轮询媒资信息、保存媒资元数据到本地文件,并提供查询与删除接口。代码甚至明确支持 image/video 分类,但没有任何与字幕文件(如 SRT/ASS/VTT)、字幕识别、字幕合成、FFmpeg 处理或视频渲染有关的行为。因此其主要目的与声明严重不符,属于明显的描述-行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear description-behavior mismatch. The declared purpose is a media-processing skill for adding or embedding subtitles into videos. The actual code only defines the public interface of an authentication module by importing and exporting auth-related classes. It does not process video files, generate subtitles, embed subtitles, or implement any trigger logic related to subtitle tasks. This is not a supporting implementation detail for subtitle addition; it is an unrelated authentication component.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared purpose is video subtitle generation/embedding for video files, but the supplied code does not process videos, subtitles, media files, transcription, or embedding at all. Instead, it implements auth configuration logic, including loading ACCESS_KEY_ID, SECRET_ACCESS_KEY, ARK_SKILL_API_KEY, and ARK_SKILL_API_BASE from environment variables and choosing an authentication strategy. This is a materially different primary purpose and accesses resources (credentials/env vars) unrelated to the declared subtitle-adding functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明的核心能力是视频字幕处理(添加、生成、嵌入字幕),但提供的代码块完全没有任何视频处理、字幕生成、字幕轨嵌入、媒体转码或字幕文件操作逻辑。相反,代码的主要功能是从URL下载文件并保存到本地,属于通用下载器能力。这是与声明用途 materially different 的主要行为,且包含未声明的网络访问与文件下载能力,因此应判定为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear description-behavior mismatch. The declared purpose is a subtitle-adding skill for videos, but the code chunk only implements media metadata extraction utilities. For images, it gets file size and dimensions; for videos, it gets file size, width, height, FPS, frame count, duration, and nearest standard aspect ratio. No subtitle generation, subtitle parsing, subtitle timing, subtitle rendering, muxing, or video modification is present. This is not a supporting detail of subtitle embedding; it is a distinct utility with a different primary purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared purpose is specifically about adding or embedding subtitles into videos. However, the supplied code chunk contains only low-level hash/checksum helper functions and no logic related to video processing, subtitle generation, subtitle embedding, media handling, or subtitle file manipulation. While utility code can support a larger system, this chunk's behavior is materially unrelated to the declared subtitle-adding functionality, so it is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this skill adds or embeds subtitles into videos. However, the supplied code chunk only contains schema/model definitions (Matriel, ImageMatriel, VideoMatriel) for storing media metadata. There is no logic for reading video files, generating subtitles, synchronizing captions, invoking speech recognition, muxing subtitle tracks, or burning subtitles into video. This is a materially different primary purpose from the declared subtitle-adding functionality, so the description does not accurately represent the code shown.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的核心能力是对视频添加字幕,但提供的代码仅包含一个抽象校验器、图片校验器、视频校验器及工厂分发逻辑。其行为只是检查输入素材的宽高、总像素和视频分辨率是否可获取。这与“添加字幕到视频”的主要目的明显不一致。虽然这类校验可能是某些媒体处理流程的辅助步骤,但当前代码片段本身没有任何字幕识别、字幕文件生成、字幕轨嵌入、视频转码或输出处理逻辑,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says this skill adds or embeds subtitles into videos. The supplied code does not operate on video files, subtitles, or media processing at all. Instead, it creates a command-line entry point that calls an external service endpoint named 'RegisterArkClawCombo', apparently to query/register a free package, and prints the result. This is a materially different primary purpose and an undeclared capability involving external service interaction. Therefore, the description does not accurately represent the code's behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this skill adds or embeds subtitles into videos. However, the supplied code only defines an upgrade.py command that submits a request containing the skill name to a remote ICCP service, then polls for completion and prints the result. Its docstring literally indicates 'get latest skill version.' There is no handling of video inputs, no subtitle generation, no subtitle file parsing, no ffmpeg/media operations, and no embedding of subtitles into a video. Therefore the actual behavior shown is materially different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear description-behavior mismatch. The declared purpose is subtitle addition/embedding for videos, but the supplied code only handles validating and uploading a local video file to a media service. No subtitle generation, subtitle file handling, speech-to-text, subtitle muxing/burning, or video rewriting appears in the code. The primary purpose is materially different: media upload, not subtitle processing.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill explicitly instructs the agent to ask users to paste cloud ACCESS_KEY_ID and SECRET_ACCESS_KEY directly into chat, then use them in-session. Collecting long-lived cloud credentials through chat is highly dangerous because chats may be logged, exposed to operators, reused by other components, or leaked, leading to full account compromise and abuse of the user's cloud resources.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill directs the agent to obtain cloud credentials via chat and immediately operationalize them by exporting them into the session environment. This creates a direct secret-handling vulnerability: compromise of chat logs, prompt traces, environment dumps, or downstream tools could expose credentials with potentially broad cloud-account access.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill collects AK/SK credentials directly in chat and provides no meaningful warning that the user is sharing secrets into a potentially logged conversational channel. In this context, the subtitle use case makes the behavior especially inappropriate because secret collection is unrelated to the user's primary task and can normalize unsafe credential-sharing practices.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 128)May include surrounding context.

md
python3.12 -m pip install -r ./scripts/requirements.txt

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The version-check flow allows execution of an install_command returned from an external source. This is a classic remote code execution and supply-chain risk: a compromised update service or tampered response could cause arbitrary commands to run with the agent's privileges.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The client logs full request headers and body before sending requests, which includes Authorization credentials and potentially sensitive request payloads. Anyone with access to logs could recover bearer tokens, signed authorization data, or user content, enabling credential reuse, lateral movement, or exposure of sensitive input data.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The service hard-codes a remote image URL into the submitted payload instead of operating on user-supplied video/subtitle assets, which is inconsistent with the skill’s stated purpose. In a video subtitling skill, this can cause undisclosed outbound access to a third-party resource, incorrect task execution, and suggests the backend is performing actions unrelated to the user’s requested media, increasing the risk of misuse or deceptive behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The file claims to implement a video subtitling skill, but the actual behavior constructs an ICCP service client and calls RegisterArkClawCombo, which is unrelated to subtitle generation or video processing. This mismatch indicates deceptive functionality and can cause unauthorized account/package registration or external side effects when users believe they are only invoking a local media operation.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill declares no explicit tool scope while its documented behavior requires environment-variable access, file reads/writes, and network operations. This weakens least-privilege controls and makes accidental or unauthorized capability expansion harder to detect or constrain.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description and the rest of the file prescribe the skill's behavior entirely in Chinese, including fixed user prompts and templates, with no indication that the assistant should match the user's preferred language. This creates a locale/language policy issue because the skill effectively enforces a specific language without user opt-in or documented regional justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.