T09 · Insecure Skill Coding Practices
- Location
scripts/screenshot.py:204- Finding
Silent Full-Desktop Capture Can Be Uploaded to a Remote Vision Service and Used for Global Clicking
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This game-assistant skill is not overtly malicious, but it needs review because it can capture screens, upload images to a cloud vision provider, store an API key, and click or type on the desktop with weak window-scope safeguards.
Install only if you are comfortable with a local game tool that can see your screen, save screenshots, send screenshots to Alibaba Cloud/DashScope for click_text, and control mouse/keyboard input. Prefer environment variables over saved API keys, keep sensitive windows and notifications hidden, use dry-run first, and avoid using this until the full-desktop screenshot fallback, global click fallback, eval parser, plaintext key storage, and dependency pinning are fixed.
scripts/screenshot.py:204Silent Full-Desktop Capture Can Be Uploaded to a Remote Vision Service and Used for Global Clicking
scripts/recognition.py:45Arbitrary Python Execution Through Asset-Controlled ROI Expressions
scripts/config.py:36Aliyun API Key Is Persisted in Plaintext in the Project Directory
requirements.txt:1Open-Ended Dependency Constraints Permit Installation of Unreviewed Future Releases
The documented behavior goes beyond the high-level description by including direct keyboard input, window enumeration, and API key/config management without clearly declaring those powers in the skill metadata. This mismatch is dangerous because users and orchestrators may trust the description while the skill actually supports more invasive local automation and access than advertised.
The documented behavior goes beyond the high-level description by including direct keyboard input, window enumeration, and API key/config management without clearly declaring those powers in the skill metadata. This mismatch is dangerous because users and orchestrators may trust the description while the skill actually supports more invasive local automation and access than advertised.
The API-key-based multimodal click feature implies that screenshots or on-screen content may be sent to a third-party provider, but the documentation does not clearly disclose that transmission or its privacy implications. Because game windows and desktops can contain personal or account-related information, undisclosed third-party exfiltration risk is significant.
The natural-language click feature captures a screenshot and sends screenshot-derived UI content to an external GUI agent provider, but the flow shown here does not provide an explicit warning or consent checkpoint at time of transfer. Screenshots may contain account names, chats, overlays, or other sensitive on-screen information, making silent third-party disclosure a meaningful privacy and data-handling risk.
When the game window cannot be resolved, the code falls back to pyautogui.click(x, y), which performs a global desktop click at arbitrary coordinates. This breaks the intended application boundary and can trigger actions in other windows, including accidental confirmation dialogs, chat apps, browsers, or security-sensitive prompts.
The code evaluates ROI expressions from a JSON button configuration using Python's eval(), which can execute arbitrary Python expressions rather than just arithmetic. Because game_name influences which buttons.json file is loaded and skill assets should be treated as untrusted/adversarial, a malicious or tampered config could achieve arbitrary code execution in the agent process.
if isinstance(expr, str):
expr = expr.replace('width', 'base').replace('height', 'base')
try:
return int(eval(expr, {'base': base_value}))
except:
return 0
The document is written entirely in Chinese and presents the skill's operational rules and interaction patterns only in that language. There is no indication that users may choose another language or that the Chinese-only constraint is a documented region-specific requirement, which creates a natural-language locale policy concern.
The file claims the skill is only an assistive companion, but the same document specifies a screenshot-analysis-action loop, button clicking, key presses, hold actions, and background clicking. This mismatch can weaken safety boundaries in the host agent by normalizing iterative autonomous control over a local application, which is especially sensitive because the skill runs on the local machine with access to windows and input APIs.
The document is framed entirely in Chinese from the title onward and repeatedly instructs Claude on how it should behave, but it does not state that language choice is optional or user-selected. Under the policy rule, forcing a specific language or locale without opt-in is a natural-language policy concern.
The README presents the skill as a Paimon persona entirely in Chinese and frames interaction in that language, without any indication that users can choose another language. This creates a natural-language locale constraint that is neither optional nor explicitly justified as a region-specific requirement.
The README states that screenshots are analyzed by a multimodal cloud model via Alibaba Cloud's OpenAI-compatible endpoint, but it does not clearly disclose that captured game images may be transmitted to a third-party service. Because screenshots can contain account names, chat messages, friend lists, payment-related UI, or other personal data, users may unknowingly expose sensitive information to an external provider.
The skill documents capabilities that involve screenshots, configuration management, and interaction with local files/environment, but it does not declare any explicit tool scope or permission boundaries. In an automation skill that can inspect windows and operate on local state, missing least-privilege declarations increases the chance that an agent will invoke broader capabilities than the user expects.
The skill emphasizes screenshot capture and automated clicking but does not clearly warn that screenshots may contain sensitive desktop content or that input automation can affect unintended windows if targeting is wrong. For a local automation skill, missing privacy and safety warnings materially increases the chance of accidental data exposure or unintended system interaction.
The file states that after loading the SOUL, the AI '会化身为可爱的派蒙,用派蒙的语气与你交流', and the entire skill description is written as a Chinese persona experience with no indication that users may choose another language. This is a natural-language locale/language constraint that appears mandatory rather than opt-in.
The document says the skill is not for full automation, yet it explicitly describes a screenshot-analyze-act loop that can be repeated, which is functionally close to autonomous operation. That contradiction can normalize unattended control behavior and increase the risk of misuse, policy violations, or unintended actions inside the game or on the desktop.
The SOUL file defines a broad general-purpose assistant with knowledge, automation, file operations, web search, writing, and coding capabilities, which materially exceeds the manifest's stated role as a Genshin-specific companion. This scope expansion is dangerous because downstream agents may treat the file as authorization to perform unrelated high-risk actions under a game-assistant identity, increasing the chance of privilege misuse and user confusion.
The documentation explicitly grants authority to automate tasks, operate software and files, and perform network searches, none of which are necessary for a narrowly scoped in-game companion. In an agentic environment, these statements can be interpreted as permission to touch user systems and external resources, creating a real risk of destructive actions, data exposure, or overbroad autonomy.
The example says the assistant will organize files on the desktop by creating folders and moving files, but it provides no warning, preview, scope limit, or confirmation step. In a system with execution capability, this normalizes direct modification of user data and can lead to accidental data loss, misplaced files, or unauthorized changes.
The manifest hard-codes Chinese-language values for the game name, window title, character, and keywords (e.g. "原神", "派蒙", "提瓦特") without any indication that users can opt into another language or locale. This can violate language/locale policy because the skill appears tailored to a fixed locale while lacking documented user choice or justification.
The skill explicitly describes continuous screenshot-based observation of the user's game window, including reading UI state and inventory-like information, but it does not provide a clear privacy notice, consent boundary, or data-handling limitation. Even if intended for gameplay assistance, screen capture can expose incidental sensitive information in overlays, notifications, account details, or other visible desktop content if the wrong window is captured or the game is not the only visible content.
The skill advertises automated mouse and keyboard control and includes concrete command examples for clicking, keypresses, and holding movement keys, but it lacks a clear warning about system interaction risks, misclicks, focus errors, and unintended actions outside the game. In context, an agent with input control can affect not just gameplay but any foreground window if focus changes, making accidental or unsafe execution materially risky.
The skill explicitly documents background clicking and keystroke automation against a live game window, but it does not warn about unintended actions, focus ambiguity, or the risk of interacting with the wrong UI state. In this context, silent background input can cause accidental purchases, destructive game actions, or persistent automation behavior without immediate user awareness.
The entire skill is defined in Chinese and presents a fixed Chinese-language persona, with no indication that the user may choose another language or that the skill is intentionally restricted to a Chinese-speaking context. Under the stated policy, a fixed language/locale without user opt-in can be a natural-language policy violation.
The skill explicitly describes continuous screenshot-based observation and the ability to control mouse and keyboard in a live game window, but it does not present clear user-facing consent, privacy, or safety boundaries. That creates a real risk of unexpected screen capture of sensitive content and unintended system-impacting actions, especially because the assistant is framed as being able to proactively observe and act.
The skill's top-level description and user-facing interface text are presented exclusively in Chinese, which effectively forces a specific language on users. There is no indication of locale selection, multilingual support, or justification for restricting the interface to one language.
Detected: suspicious.dynamic_code_execution