Back to skill

Security audit

派蒙.skill - 原神特化 AI 游戏伴侣

Security checks for vulnerabilities and agentic risk

Overview

This game-assistant skill is not overtly malicious, but it needs review because it can capture screens, upload images to a cloud vision provider, store an API key, and click or type on the desktop with weak window-scope safeguards.

Install only if you are comfortable with a local game tool that can see your screen, save screenshots, send screenshots to Alibaba Cloud/DashScope for click_text, and control mouse/keyboard input. Prefer environment variables over saved API keys, keep sensitive windows and notifications hidden, use dry-run first, and avoid using this until the full-desktop screenshot fallback, global click fallback, eval parser, plaintext key storage, and dependency pinning are fixed.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/screenshot.py:204
Finding

Silent Full-Desktop Capture Can Be Uploaded to a Remote Vision Service and Used for Global Clicking

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/recognition.py:45
Finding

Arbitrary Python Execution Through Asset-Controlled ROI Expressions

Content
View full analysis
int: """ 解析 ROI 表达式 Args: expr: 表达式 (如 "width/4", "height*3/4", 或整数) base_value: 基础值 (width 或 height) Returns: 计算结果 """ if isinstance(expr, int): return expr if isinstance(expr, str): expr = expr.replace('width', 'base').replace('height', 'base') try: return int(eval(expr, {'base': base_value})) except: return 0 return 0 ``` The expressions are loaded from the game asset configuration: ```python config_path = os.path.join(ASSETS_DIR, game_name, 'assets', 'buttons.json') if os.path.exists(config_path): with open(config_path, 'r', encoding='utf-8') as f: _button_configs[game_name] = json.load(f) ``` ### Technical Analysis `eval()` executes a string as Python code. Supplying `{'base': base_value}` as the globals dictionary does not disable builtins; Python can insert or expose `__builtins__`. Consequently, a malicious expression can invoke builtins such as `__import__` and execute arbitrary operating-system commands. The bundled `games/genshin-impact/assets/buttons.json` contains only ordinary arithmetic expressions such as `width/4` and `height-70`. No embedded malicious expression was identified in the audited file. The vulnerability becomes exploitable if the asset file is tampered with, replaced, or supplied through an untrusted game package. The broad `except` block suppresses evidence of malicious or malformed expressions and returns zero, making exploitation attempts or configuration corruption more difficult to diagnose. ### Attack Path 1. An attacker gains the ability to modify `games/genshin-impact/assets/buttons.json`, or convinces the user to install an ...[truncated 921 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/config.py:36
Finding

Aliyun API Key Is Persisted in Plaintext in the Project Directory

Content
View full analysis
None: global _config_cache config_path = get_config_dir() / CONFIG_FILE_NAME with open(config_path, 'w', encoding='utf-8') as f: json.dump(config, f, indent=2, ensure_ascii=False) _config_cache = config ``` ```python def set_api_key(api_key: str, provider: str = "aliyun") -> None: config = load_config() if "gui_agent" not in config: config["gui_agent"] = {} config["gui_agent"]["api_key"] = api_key config["gui_agent"]["provider"] = provider save_config(config) ``` The destination is project-local: ```python def get_config_dir() -> Path: return Path(__file__).parent.parent CONFIG_FILE_NAME = "config.json" ``` ### Technical Analysis The documented `config --set-api-key` workflow writes the API key directly into `config.json` under the project root. The value is not protected with an operating-system credential store, encryption, or explicit restrictive ACLs. Masking the key in the `config --show` output does not protect the underlying file. Project directories are commonly copied into backups, support archives, synchronization services, or source-control commits. Other processes or users with read access to the directory can recover the complete credential. ### Attack Path 1. The user runs `python main.py config --set-api-key KEY`. 2. `set_api_key()` inserts the complete key into the configuration dictionary. 3. `save_config()` serializes the key as plaintext into project-local `config.json`. 4. Another local process, user, backup tool, synchronization service, archive, or accidental repository commit obtains the file. 5. The exposed key is reused to make unauthorized requests against the associated Aliyun a ...[truncated 474 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Open-Ended Dependency Constraints Permit Installation of Unreviewed Future Releases

Content
View full analysis
=306 Pillow>=10.0.0 opencv-python>=4.8.0 numpy>=1.24.0 openai>=1.0.0 ``` ### Technical Analysis Every dependency uses an open-ended minimum-version constraint. As a result, `pip install -r requirements.txt` may install arbitrary future versions that were not reviewed or tested with this Skill. No lock file or package hashes are supplied to ensure artifact integrity and reproducibility. The audited package names do not show evidence of obvious typosquatting or dependency confusion. The risk arises from unconstrained resolution and the absence of integrity pinning rather than from a confirmed malicious dependency currently listed. This exposure is particularly important because these packages run inside a process with screenshot access, remote API credentials, and mouse and keyboard control. ### Attack Path 1. A user follows the documented installation command using `requirements.txt`. 2. Pip resolves the newest versions satisfying the lower bounds. 3. A future compromised, malicious, or unexpectedly incompatible release is selected. 4. Package installation hooks or imported runtime code execute on the user's machine. 5. The affected dependency obtains the same process-level access as the Skill, including accessible files, the API key, screenshots, network communication, and desktop automation functions. ### Impact Assessment A compromised dependency could execute arbitrary code under the current user account. It could access sensitive screenshots and credentials, tamper with recognition or model results, redirect network traffic, or misuse desktop automation. Even without malicious compromise, incompatible future versions can alter security-sensitive behavior or break assumptions about API requests, image handling, and native ...[truncated 24 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (65)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The documented behavior goes beyond the high-level description by including direct keyboard input, window enumeration, and API key/config management without clearly declaring those powers in the skill metadata. This mismatch is dangerous because users and orchestrators may trust the description while the skill actually supports more invasive local automation and access than advertised.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The documented behavior goes beyond the high-level description by including direct keyboard input, window enumeration, and API key/config management without clearly declaring those powers in the skill metadata. This mismatch is dangerous because users and orchestrators may trust the description while the skill actually supports more invasive local automation and access than advertised.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The API-key-based multimodal click feature implies that screenshots or on-screen content may be sent to a third-party provider, but the documentation does not clearly disclose that transmission or its privacy implications. Because game windows and desktops can contain personal or account-related information, undisclosed third-party exfiltration risk is significant.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language click feature captures a screenshot and sends screenshot-derived UI content to an external GUI agent provider, but the flow shown here does not provide an explicit warning or consent checkpoint at time of transfer. Screenshots may contain account names, chats, overlays, or other sensitive on-screen information, making silent third-party disclosure a meaningful privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

When the game window cannot be resolved, the code falls back to pyautogui.click(x, y), which performs a global desktop click at arbitrary coordinates. This breaks the intended application boundary and can trigger actions in other windows, including accidental confirmation dialogs, chat apps, browsers, or security-sensitive prompts.

Content

No source excerpt is available for this finding.

eval() call detected

High
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

The code evaluates ROI expressions from a JSON button configuration using Python's eval(), which can execute arbitrary Python expressions rather than just arithmetic. Because game_name influences which buttons.json file is loaded and skill assets should be treated as untrusted/adversarial, a malicious or tampered config could achieve arbitrary code execution in the agent process.

Content

Scanner excerpt · scripts/recognition.py (reported line 54)May include surrounding context.

python
if isinstance(expr, str):
        expr = expr.replace('width', 'base').replace('height', 'base')
        try:
            return int(eval(expr, {'base': base_value}))
        except:
            return 0

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The document is written entirely in Chinese and presents the skill's operational rules and interaction patterns only in that language. There is no indication that users may choose another language or that the Chinese-only constraint is a documented region-specific requirement, which creates a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The file claims the skill is only an assistive companion, but the same document specifies a screenshot-analysis-action loop, button clicking, key presses, hold actions, and background clicking. This mismatch can weaken safety boundaries in the host agent by normalizing iterative autonomous control over a local application, which is especially sensitive because the skill runs on the local machine with access to windows and input APIs.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The document is framed entirely in Chinese from the title onward and repeatedly instructs Claude on how it should behave, but it does not state that language choice is optional or user-selected. Under the policy rule, forcing a specific language or locale without opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README presents the skill as a Paimon persona entirely in Chinese and frames interaction in that language, without any indication that users can choose another language. This creates a natural-language locale constraint that is neither optional nor explicitly justified as a region-specific requirement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README states that screenshots are analyzed by a multimodal cloud model via Alibaba Cloud's OpenAI-compatible endpoint, but it does not clearly disclose that captured game images may be transmitted to a third-party service. Because screenshots can contain account names, chat messages, friend lists, payment-related UI, or other personal data, users may unknowingly expose sensitive information to an external provider.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill documents capabilities that involve screenshots, configuration management, and interaction with local files/environment, but it does not declare any explicit tool scope or permission boundaries. In an automation skill that can inspect windows and operate on local state, missing least-privilege declarations increases the chance that an agent will invoke broader capabilities than the user expects.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill emphasizes screenshot capture and automated clicking but does not clearly warn that screenshots may contain sensitive desktop content or that input automation can affect unintended windows if targeting is wrong. For a local automation skill, missing privacy and safety warnings materially increases the chance of accidental data exposure or unintended system interaction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The file states that after loading the SOUL, the AI '会化身为可爱的派蒙,用派蒙的语气与你交流', and the entire skill description is written as a Chinese persona experience with no indication that users may choose another language. This is a natural-language locale/language constraint that appears mandatory rather than opt-in.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The document says the skill is not for full automation, yet it explicitly describes a screenshot-analyze-act loop that can be repeated, which is functionally close to autonomous operation. That contradiction can normalize unattended control behavior and increase the risk of misuse, policy violations, or unintended actions inside the game or on the desktop.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The SOUL file defines a broad general-purpose assistant with knowledge, automation, file operations, web search, writing, and coding capabilities, which materially exceeds the manifest's stated role as a Genshin-specific companion. This scope expansion is dangerous because downstream agents may treat the file as authorization to perform unrelated high-risk actions under a game-assistant identity, increasing the chance of privilege misuse and user confusion.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation explicitly grants authority to automate tasks, operate software and files, and perform network searches, none of which are necessary for a narrowly scoped in-game companion. In an agentic environment, these statements can be interpreted as permission to touch user systems and external resources, creating a real risk of destructive actions, data exposure, or overbroad autonomy.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The example says the assistant will organize files on the desktop by creating folders and moving files, but it provides no warning, preview, scope limit, or confirmation step. In a system with execution capability, this normalizes direct modification of user data and can lead to accidental data loss, misplaced files, or unauthorized changes.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The manifest hard-codes Chinese-language values for the game name, window title, character, and keywords (e.g. "原神", "派蒙", "提瓦特") without any indication that users can opt into another language or locale. This can violate language/locale policy because the skill appears tailored to a fixed locale while lacking documented user choice or justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly describes continuous screenshot-based observation of the user's game window, including reading UI state and inventory-like information, but it does not provide a clear privacy notice, consent boundary, or data-handling limitation. Even if intended for gameplay assistance, screen capture can expose incidental sensitive information in overlays, notifications, account details, or other visible desktop content if the wrong window is captured or the game is not the only visible content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill advertises automated mouse and keyboard control and includes concrete command examples for clicking, keypresses, and holding movement keys, but it lacks a clear warning about system interaction risks, misclicks, focus errors, and unintended actions outside the game. In context, an agent with input control can affect not just gameplay but any foreground window if focus changes, making accidental or unsafe execution materially risky.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill explicitly documents background clicking and keystroke automation against a live game window, but it does not warn about unintended actions, focus ambiguity, or the risk of interacting with the wrong UI state. In this context, silent background input can cause accidental purchases, destructive game actions, or persistent automation behavior without immediate user awareness.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The entire skill is defined in Chinese and presents a fixed Chinese-language persona, with no indication that the user may choose another language or that the skill is intentionally restricted to a Chinese-speaking context. Under the stated policy, a fixed language/locale without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill explicitly describes continuous screenshot-based observation and the ability to control mouse and keyboard in a live game window, but it does not present clear user-facing consent, privacy, or safety boundaries. That creates a real risk of unexpected screen capture of sensitive content and unintended system-impacting actions, especially because the assistant is framed as being able to proactively observe and act.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill's top-level description and user-facing interface text are presented exclusively in Chinese, which effectively forces a specific language on users. There is no indication of locale selection, multilingual support, or justification for restricting the interface to one language.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/recognition.py:54