Back to skill

Security audit

小剪刀视频剪辑

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly matches a cloud video-editing workflow, but it handles tokens unsafely and can upload overly broad local files to external services.

Review before installing. Use this only for media you are willing to upload to the listed cloud services, stage files in a dedicated folder, and do not provide arbitrary local paths. Avoid pasting long-lived tokens into chat or command lines; prefer a scoped, revocable token through a secure secret mechanism if available.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
oss_upload.py:20
Finding

Unrestricted Local File Upload to an External Service

Content
View full analysis
dict: """ 上传文件到 OSS (QA环境) """ if not os.path.isfile(file_path): return {"success": False, "error": f"File not found: {file_path}"} target_url, inferred_usage = infer_route_and_usage(file_path) final_usage = usage or inferred_usage filename = os.path.basename(file_path) try: with open(file_path, "rb") as f: files = { "files": (filename, f), } data = { "usage": final_usage } ...[truncated 2104 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:126
Finding

Authentication Token Exposure Through Conversation and Command-Line Arguments

Content
View full analysis
" --token "" --task-id --frames 15 ``` ```markdown # SKILL.md:126-132 ## 🔑 Token 配置 (交互式) 为了方便分享,本 Skill 支持**无需预设环境变量**: 1. **首次运行**:如果你没有设置环境变量,我会提示你“请输入小剪刀 Token”。 2. **输入方式**:你可以直接在对话中发送你的 Token。 3. **有效期**:建议使用环境变量 `XJD_TOKEN` 以实现长期免输,但直接在对话中提供也是完全支持的。 ``` ```python # bridge.py:424,439,454-460 parser.add_argument("--token", help="API Token for MScissors") args = parser.parse_args() token = get_token(args) if not token: print(json.dumps({ "err_code": 401, "err_message": "Token missing. Please provide a valid token in the conversation.", "success": False }, ensure_ascii=False)) sys.exit(0) ``` ```python # subtitle_ocr.py:199-209 if __name__ == "__main__": import argparse parser = argparse.ArgumentParser(description="字幕识别 - 云端优化版") parser.add_argument("--url", required=True, help="视频 URL") parser.add_argument("--token", required=True, help="用户 Token") parser.add_argument("--task-id", type=int, default=0, help="任务 ID") parser.add_argument("--frames", type=int, default=15, help="抽帧数量") args = parser.parse_args() result = recognize_subtitle( video_url=args.url, biyi_token=args.token, task_id=args.task_id, num_frames=args.frames ) ``` ### Technical Analysis The documented workflow explicitly asks users to submit an authentication token in the conversation. It also supports passing the token through the `--token` command-line argument. Conversation content can be retained in chat history, agent traces, support exports, telemetry, or orchestration logs. Command-line arguments may be ex ...[truncated 1417 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Note
Location
client.py:18
Finding

Unnecessary Host Fingerprinting Transmitted with API Requests

Content
View full analysis
Dict[str, Any]: sys_name = platform.system().lower() os_name = "macos" if sys_name == "darwin" else ("windows" if sys_name.startswith("win") else sys_name) os_version = platform.version() arch = platform.machine() or "unknown" return { "app_os": f"{os_name} {os_version}", "app_vn": "1.0.0", "appid": appid, "pkg_name": pkgname, "mac": "00:00:00:00:00:00", "arch": arch, "cpu_model": platform.processor() or "Generic CPU", "api_vn": api_vn, "cl": "official", "device_id": str(uuid.uuid4()), } ``` ```python # client.py:50-63 def _build_headers(self, *, unix_ms: int, device: Optional[Dict[str, Any]] = None) -> Dict[str, str]: aes = self._aes() device_info = device or self.device_info or default_device_info( pkgname=self.pkgname, appid=self.appid, api_vn=self.api_vn ) device_json = json.dumps(device_info, ensure_ascii=False, separators=(",", ":")) encrypted_device = aes.encrypt_to_base64(device_json, unix_ms=unix_ms) return { "Content-Type": "application/x-aes-ms-txt", "pkgname": self.pkgname, "device": encrypted_device, "appsecret": self.appsecret, "X-API-Token": self.appsecret, "x-ms-at": str(unix_ms), } ``` ### Technical Analysis Each request made through `HookHttpClient` automatically gathers the operating-system version, machine architecture, and CPU model. It also generates a device identifier. The resulting profile is encrypted and included in the `device` request header. Encryption does not prevent the intended service from reading these values. The primary Skill documenta ...[truncated 1254 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
crypto.py:31
Finding

Hardcoded Shared AES Key and Unauthenticated Application-Layer Encryption

Content
View full analysis
str: mapping = { "com.biyi.wscissors": "wd.srv.yz.2025!@", "com.biyi.mscissors": "bz.srv.yz.2025!@", } return mapping.get(pkgname, "biyi#def!918.dyj") def get_iv_by_pkg(pkgname: str, unix_ms: int) -> str: if not pkgname: raise ValueError("pkgname is empty") if unix_ms < 50 * 86400 * 365 * 1000: raise ValueError("unix_ms is less than 2020") sign = f"{pkgname}{unix_ms}" md5_hex = hashlib.md5(sign.encode("utf-8")).hexdigest() return md5_hex[8:24] ``` ```python # crypto.py:57-76 def encrypt_to_base64(self, text: str, *, unix_ms: int) -> str: self._ensure_aes() key = get_aes_key_by_pkg(self.pkgname) iv = get_iv_by_pkg(self.pkgname, unix_ms) if len(key) != 16 or len(iv) != 16: raise ValueError(f"key/iv length must be 16 (key={len(key)}, iv={len(iv)})") cipher = AES.new(key.encode("utf-8"), AES.MODE_CBC, iv.encode("utf-8")) padded = _pkcs7_pad(text.encode("utf-8"), 16) encrypted = cipher.encrypt(padded) return base64.b64encode(encrypted).decode("utf-8") def decrypt_base64(self, b64_text: str, *, unix_ms: int) -> str: self._ensure_aes() key = get_aes_key_by_pkg(self.pkgname) iv = get_iv_by_pkg(self.pkgname, unix_ms) if len(key) != 16 or len(iv) != 16: raise ValueError(f"key/iv length must be 16 (key={len(key)}, iv={len(iv)})") encrypted = base64.b64decode(b64_text.encode("utf-8")) cipher = AES.new(key.encode("utf-8"), AES.MODE_CBC, iv.encode("utf-8")) padded = cipher.decrypt(encrypted) plain = _pkcs7_unpad(padded, 16) return plain.decode("utf-8") ``` ### Technical Analysis The application-layer encr ...[truncated 2203 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (38)

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The declared description suggests an AI video editing skill, but this code chunk only provides infrastructure for communicating with a backend service. It gathers platform/device information, encrypts headers and payloads, posts data over the network, handles streamed responses, and polls asynchronous tasks. There is no concrete video editing functionality in the supplied code, so the actual behavior is materially broader/different than the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

声明的核心用途是视频剪辑,但实际代码仅是通用/业务定制的加密解密实现,主要处理字符串的AES加密、解密、填充和IV生成。虽然加密模块可能作为更大系统的辅助组件存在,但就该代码块本身而言,其行为与“AI视频剪辑”没有直接对应关系,属于 materially different primary purpose。未见视频处理、剪辑、转码、媒体访问或AI推理相关实现,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个视频剪辑 skill 的整体用途,但该代码块并未执行视频剪辑、编辑、拼接、转码或生成等核心视频处理行为,而是专注于将本地文件上传到远程 OSS/接口服务。虽然上传素材可能是视频剪辑流程中的辅助步骤,但从该代码块本身看,其主要功能是网络文件上传,且还支持音频、图片、字幕等多类资源上传,属于与“AI视频剪辑”不一致的实际行为范围。声明中也未体现任何网络上传或远程存储访问能力,因此应判定为描述与代码行为不匹配。

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Requesting sensitive access tokens in chat without a strong privacy/security warning exposes users to credential theft through transcript retention, accidental sharing, and downstream logging. Because the token appears sufficient to operate remote editing APIs, compromise could let others access or manipulate user tasks and media.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

Treating direct token submission in chat as normal operation is a clear secret-handling vulnerability. It creates a predictable channel for credential capture and replay, and the surrounding skill context makes this more dangerous because the token is required to upload/process user media with a third-party service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

Telling users to obtain a token 'from packet capture' is a direct instruction to intercept credentials outside normal authorization controls. This strongly suggests bypass of intended access mechanisms and can facilitate account compromise, unauthorized API use, and abuse of the service infrastructure.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This is a real privacy and transport-security issue. The skill explicitly instructs clients to upload extracted video frames containing potentially sensitive visual data to a public OSS/CDN endpoint '明文,不走加密通道', then exposes them via publicly accessible URLs for OCR fetching. That creates clear risks of interception in transit, unauthorized third-party access, leakage of personal or confidential on-screen content, and unintended retention in a public location.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
84% confidence
Finding

The skill explicitly directs use of networked helper scripts and token-based remote APIs, but it declares no tool scope or permissions boundaries. That creates an authorization gap: an agent may invoke network and environment capabilities without an auditable least-privilege declaration, increasing the chance of unintended data access or exfiltration.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Natural-language policy review applies to all file types. The description, headings, and prescribed welcome behavior are all fixed in Chinese, and the skill does not indicate that users may choose another language or locale, which can violate language/locale choice policy.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This is a markdown file, so vague-trigger review applies. The listed triggers include broad everyday phrases like "帮我剪辑视频" and "剪辑视频", and the file does not provide scope limits, negative examples, or a constrained invocation context, which could cause unintended activation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill instructs users to paste authentication tokens directly into chat, normalizing unsafe credential handling and expanding token exposure to chat logs, model context, operators, plugins, and retention systems. For a video editing workflow, collecting secrets in-band is broader than necessary and materially increases the chance of credential leakage or misuse.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The script restricts voice handling to Chinese voice names mapped to zh-CN Azure voices and defaults to a Chinese locale voice when no preference is provided. There is no visible opt-in or alternative language selection path, which can violate language/locale policy when used in a general-purpose skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code uploads a local file via oss_upload() and then submits the resulting media URL to remote services without any built-in user confirmation, consent checkpoint, or data minimization. In a skill context, this can expose sensitive local media or embedded metadata to third-party infrastructure unexpectedly, which is a real privacy and data exfiltration risk.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This request sends a user-supplied video URL and bearer token to an external service for frame extraction. In this skill's context, external transmission is expected for cloud video processing, but it still constitutes real data egress of potentially sensitive media references and credentials to third-party endpoints.

Content

Scanner excerpt · bridge.py (reported line 125)May include surrounding context.

python
frames = []
        try:
            resp = requests.post(url, json=payload, headers=headers, stream=True, timeout=None)
            for line in resp.iter_lines():
                if line:
                    line_str = line.decode("utf-8")

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
96% confidence
Finding

Using timeout=None on a streaming external request allows the connection to hang indefinitely if the remote endpoint stalls or never terminates the stream. In an agent or automation environment, this can cause denial of service, blocked workers, resource exhaustion, and degraded availability.

Content

Scanner excerpt · bridge.py (reported line 125)May include surrounding context.

python
frames = []
        try:
            resp = requests.post(url, json=payload, headers=headers, stream=True, timeout=None)
            for line in resp.iter_lines():
                if line:
                    line_str = line.decode("utf-8")

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This code transmits a video URL and authorization token to a remote audio-extraction endpoint, which can expose sensitive media references and associated account capabilities if intercepted, misused, or logged. Because the feature is cloud-based, the transmission is intentional, but it is still a genuine privacy and data-handling risk.

Content

Scanner excerpt · bridge.py (reported line 155)May include surrounding context.

python
audio_url = None
        try:
            resp = requests.post(url, json=payload, headers=headers, stream=True, timeout=None)
            for line in resp.iter_lines():
                if line:
                    line_str = line.decode("utf-8")

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
96% confidence
Finding

This streaming request for audio extraction also has no timeout, so a malicious or faulty remote service can keep the process blocked indefinitely. In long-running agent workflows, repeated hangs can exhaust worker capacity and prevent other tasks from executing.

Content

Scanner excerpt · bridge.py (reported line 155)May include surrounding context.

python
audio_url = None
        try:
            resp = requests.post(url, json=payload, headers=headers, stream=True, timeout=None)
            for line in resp.iter_lines():
                if line:
                    line_str = line.decode("utf-8")

External Transmission

Medium
Category
Data Exfiltration
Confidence
92% confidence
Finding

The final composition request can transmit a large payload including video paths, subtitle paths, audio ZIP URLs, BGM paths, and task metadata to an external service. In aggregate, this can leak substantial user content and processing metadata, making the impact higher than the earlier single-purpose transmissions if a service is compromised or inputs are sensitive.

Content

Scanner excerpt · bridge.py (reported line 355)May include surrounding context.

python
}
        final_res = {}
        try:
            resp = requests.post(url, json=payload, headers=headers, stream=True, timeout=None)
            for line in resp.iter_lines():
                if line:
                    line_str = line.decode("utf-8")

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
97% confidence
Finding

An unbounded timeout on the final composition stream is more severe because this operation may run longer, carry larger payloads, and sit on critical workflow resources while waiting for completion. A stalled upstream can therefore tie up the skill indefinitely and create a practical denial-of-service condition for users or shared agent infrastructure.

Content

Scanner excerpt · bridge.py (reported line 355)May include surrounding context.

python
}
        final_res = {}
        try:
            resp = requests.post(url, json=payload, headers=headers, stream=True, timeout=None)
            for line in resp.iter_lines():
                if line:
                    line_str = line.decode("utf-8")

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The client gathers host fingerprinting details such as OS version, architecture, CPU model, package name, and a generated device ID, then packages them for transmission. For a video-editing skill, this exceeds what is clearly necessary to perform editing operations and creates privacy and tracking risk if the backend stores or correlates these identifiers.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The request headers include encrypted device metadata plus credential-like values in appsecret and X-API-Token without any visible notice, consent flow, or minimization logic. Even if transport is expected, silently transmitting identifiers and secrets increases privacy and misuse risk, especially in a client library that may run on end-user systems.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The stream_post method accepts an arbitrary URL, headers, and JSON payload, allowing the skill to send data to destinations outside the declared service boundary. In an agent-skill context, this broad network primitive can be abused for unauthorized exfiltration or to interact with attacker-controlled endpoints under the guise of normal streaming functionality.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The streaming method sends arbitrary payloads to arbitrary URLs with no built-in user-facing indication that outbound transmission is occurring. In a skill environment, hidden network egress is security-relevant because it can carry prompts, files, or other sensitive runtime data off-platform without meaningful transparency.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
93% confidence
Finding

Using timeout=None for a streaming HTTP request permits an unbounded wait on network resources. A malicious or malfunctioning server can keep the connection open indefinitely, leading to thread, socket, or worker exhaustion and degraded availability of the host application.

Content

Scanner excerpt · client.py (reported line 120)May include surrounding context.

python
def stream_post(self, url: str, payload: Any, headers: Dict[str, str] = None) -> Iterator[Dict[str, Any]]:
        if headers is None:
            headers = {"Content-Type": "application/json"}
        resp = self._session.post(url, json=payload, headers=headers, timeout=None, stream=True)
        for line in resp.iter_lines():
            if line:
                line_str = line.decode("utf-8")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code contains natural-language strings in Chinese, including a comment and later a user-facing runtime error, with no indication that the skill is region-specific or that users can choose another language. Under the language/locale policy, forcing a specific language without opt-in is a policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.