Back to skill

Security audit

Movie Subtitle Viewer

Security checks for vulnerabilities and agentic risk

Overview

This subtitle skill mostly matches its stated purpose, but it needs review because downloaded subtitle filenames can control where files are written and may overwrite local files.

Review this skill before installing. Use it only in a restricted workspace, provide explicit safe save paths, avoid running downloads where overwriting local files would matter, and be aware that OpenSubtitles credentials, search terms, and possibly subtitle text may be sent to external services. Prefer a version that confines downloads to a dedicated directory, validates returned URLs, checks response status and size, and pins dependencies.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
src/subtitle_client.py:99
Finding

Untrusted Subtitle Filename Allows Arbitrary File Overwrite

Content
View full analysis

Vulnerability Details

File Location: src/subtitle_client.py, lines 99–103
Vulnerability Type: Path traversal and unrestricted file write
Risk Level: Medium

Vulnerable Code:

python
# 保存文件
if save_path is None:
    save_path = subtitle.get('file_name', 'subtitle.srt')
    
with open(save_path, 'wb') as f:
    f.write(r2.content)

Technical Analysis

When the caller does not provide save_path, the application takes file_name directly from the subtitle dictionary and uses it as the destination passed to open().

The filename is not reduced to a basename, normalized against an approved download directory, or checked for absolute paths and parent-directory components. Consequently, values such as ../../configuration.json or /home/user/.config/application.conf can cause downloaded content to be written outside the intended subtitle workspace.

The subtitle dictionary is normally derived from an OpenSubtitles API response, but download() is a public method that also accepts caller-constructed dictionaries. Exploitation therefore does not strictly require control of the remote API: any party able to influence the dictionary supplied to this method and provide a valid file_id can control the default output path.

The use of write-binary mode ('wb') truncates an existing destination file before writing the downloaded response, making this an overwrite primitive rather than merely an unauthorized file creation issue.

Attack Path

  1. The attacker obtains or identifies a valid OpenSubtitles file_id.

  2. The attacker causes download() to receive a dictionary such as:

    python
    {
        "file_id": VALID_FILE_ID,
        "file_name": "../../target-file"
    }
    
  3. The caller invokes download() without an explicit trusted save_path.

  4. The client requests a download link and retrieves the remote subtitle content.

  5. The relative path escapes the c ...[truncated 1058 chars]

Remediation
View remediation

Remediation Suggestions

Store downloads exclusively under a dedicated, trusted directory and never treat a remote filename as a path.

  1. Reduce the supplied filename to its basename using Path(filename).name.
  2. Reject empty names, absolute paths, path separators, and parent-directory components.
  3. Restrict accepted extensions to supported subtitle formats such as .srt and .ass.
  4. Resolve the final path and verify that it remains below the approved download directory.
  5. Avoid silently overwriting existing files; use exclusive creation or generate a unique filename.
  6. Treat an explicitly supplied save_path as untrusted unless it is similarly constrained by the surrounding application.
  7. Validate the download response with raise_for_status() and apply a maximum response-size limit before writing.

Example hardened approach:

python
from pathlib import Path
import uuid

download_dir = Path("workspace/subtitles").resolve()
download_dir.mkdir(parents=True, exist_ok=True)

raw_name = subtitle.get("file_name") or f"{uuid.uuid4()}.srt"
safe_name = Path(raw_name).name

if Path(safe_name).suffix.lower() not in {".srt", ".ass"}:
    raise ValueError("Unsupported subtitle extension")

destination = (download_dir / safe_name).resolve()
if download_dir not in destination.parents:
    raise ValueError("Invalid subtitle path")

r2.raise_for_status()
with destination.open("xb") as output:
    output.write(r2.content)
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (18)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 22)May include surrounding context.

1. 设置环境变量

bash
# 创建 .env 文件(不要提交到 Git!)
OPENSUBTITLES_API_KEY=your_api_key
OPENSUBTITLES_USERNAME=your_username
OPENSUBTITLES_PASSWORD=your_password

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README explicitly instructs users to send subtitle lines to an AI for plot summarization, but it does not warn that subtitle content may be transmitted to a third-party model provider or logged by external services. Even if subtitles are not traditional secrets, they can contain copyrighted or sensitive content, and users need clear notice before sharing data outside their environment.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrase "看电影" is very broad and maps to ordinary conversational intent rather than a narrowly scoped tool action. This can cause accidental skill activation during unrelated movie discussions, leading the agent to perform network searches/downloads or process files when the user did not explicitly intend to invoke the skill.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger "subtitle" is too generic, especially in mixed-language conversations, and may appear in ordinary requests, filenames, or technical discussion without intent to activate this skill. That ambiguity raises the risk of unintended invocation and downstream actions such as querying OpenSubtitles or saving files to the workspace.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · src/subtitle_client.py (reported line 8)May include surrounding context.

python
class SubtitleClient:
    """OpenSubtitles API 客户端"""
    
    BASE_URL = "https://api.opensubtitles.com/api/v1"
    
    def __init__(self):
        self.token = None

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · src/subtitle_client.py (reported line 28)May include surrounding context.

python
"password": self.password
        }
        
        r = requests.post(url, json=data, headers=headers, timeout=30)
        r.raise_for_status()
        
        self.token = r.json()['token']

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · src/subtitle_client.py (reported line 86)May include surrounding context.

python
"password": self.password
        }
        
        r = requests.post(url, json=data, headers=headers, timeout=30)
        r.raise_for_status()
        
        self.token = r.json()['token']

Tainted flow: 'download_link' from requests.post (line 90, network input) → requests.get (network output)

Medium
Category
Data Flow
Confidence
93% confidence
Finding

The code fetches a second URL taken directly from the first API response without validating the destination host, scheme, or expected path. If the upstream API is compromised, misconfigured, or intercepted, this creates a server-side request forgery style primitive and can also download unexpected content from an attacker-controlled location.

Content

Scanner excerpt · src/subtitle_client.py (reported line 95)May include surrounding context.

python
raise ValueError("No download link in response")
            
        # 下载实际文件
        r2 = requests.get(download_link, timeout=60)
        
        # 保存文件
        if save_path is None:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The code saves downloaded content directly to a path on the local filesystem using open(save_path, 'wb'), but there is no confirmation prompt, warning message, or other user disclosure in the download method. Because file creation/modification affects user data and system state, this operation should be disclosed unless it is clearly communicated elsewhere in the skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill description is written in Chinese and does not indicate that the language is optional, user-selectable, or required for a region-specific use case. This can violate a language/locale policy when users are not given a choice or justification for the constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

Natural-language policy review applies to all file types. The skill content, trigger examples, and usage are presented only in Chinese, with no opt-in or statement that the skill is intended exclusively for a Chinese-language audience, which can amount to a forced language choice.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
94% confidence
Finding

The dependency is specified with a lower bound only, which allows future unreviewed versions to be installed and makes builds non-reproducible. This increases supply-chain risk because a later vulnerable or malicious release could be pulled without any code change in the skill.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
pysubs2>=1.8.0
requests>=2.28.0
python-dotenv>=0.21.0

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
98% confidence
Finding

requests is not pinned to an exact version, so installations may resolve to different releases over time, including versions with known security issues. Because requests is a network-facing library, version drift can directly expose the skill to transport, credential, or request-handling vulnerabilities.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
pysubs2>=1.8.0
requests>=2.28.0
python-dotenv>=0.21.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The manifest leaves the installed requests version unverifiable, and the package has multiple known advisories across its release history. Without pinning, it is impossible to determine whether deployments may resolve to a vulnerable version, which is especially concerning for a commonly used HTTP client that may handle credentials, redirects, and remote content.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

python-dotenv is also unpinned, allowing arbitrary newer releases to be installed without review. While often used only for local configuration loading, unpinned versions still create supply-chain and reproducibility risks and may introduce newly disclosed flaws.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
pysubs2>=1.8.0
requests>=2.28.0
python-dotenv>=0.21.0

Unverifiable Dependency: python-dotenv has 2 known advisory(ies) (CVE-2026-28684 (python-dotenv: Symlink following in set_key allows arbitrary file overwrite via ); CVE-2026-28684 (python-dotenv reads key-value pairs from a .env file and can set them as environ)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
89% confidence
Finding

The python-dotenv dependency is not pinned, so the actual installed version cannot be checked against known advisories with confidence. If the skill uses .env file manipulation or loading in unsafe filesystem contexts, a vulnerable release could contribute to file overwrite or configuration handling issues.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The module docstring includes a Chinese description alongside an English title, which implies language-specific behavior or presentation without stating that users can choose their preferred language. Under the policy for natural-language violations, forcing or assuming a language without opt-in can be a locale/language policy issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code includes Chinese docstrings and English CLI output strings, but provides no user opt-in or documented locale policy. That can violate a language/locale policy when a skill implicitly forces particular languages rather than allowing user choice.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.