Back to skill

Security audit

windows-screen-coordinate

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly purpose-aligned, but it combines paid remote authorization, broad screenshot/UI inspection, and real mouse control with some under-scoped and inconsistent privacy guidance.

Install only if you are comfortable with a paid Windows desktop automation helper that can capture screenshots, inspect UI elements, contact coord.weituoai.cn with a device identifier for authorization, and perform real mouse actions. Use DryRun or explicit confirmation before act.ps1, avoid using it on sensitive screens, and periodically clear %TEMP%\screen-coord-locator\.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (37)

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill is described as a coordinate-location utility, but this script performs live mouse movement, clicks, drags, and wheel actions. That expands the capability from observation to direct UI manipulation, enabling unintended state changes, destructive clicks, or abuse by downstream agents that assume the skill is read-only.

Content

No source excerpt is available for this finding.

YARA rule 'keylogger_indicators': Keylogger functionality in scripts or source code [malware]

High
Category
YARA Match
Confidence
70% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · scripts/lib/common.ps1 (reported line 59)May include surrounding context.

text
ting) { return }

    $src = @'
using System;
using System.Runtime.InteropServices;
using System.Text;

public class ScfNative {
    [DllImport("user32.dll", SetLastError = true)]
    public static extern bool SetProcessDpiAwarenessContext(IntPtr value);

    [DllImport("user32.dll")]
    public static extern bool SetProcessDPIAware();

    [DllImport("user32.dll")]
    public static extern short GetAsyncKeyState(int vKey);

    [DllImport("user32.dll")]
    public static extern bool GetCursorPos(out ScfPoint p);

    [DllImport("user32.dll")]
    public static extern bool SetCursorPos(int x, int y);

    [DllImport("user32.dll")]
    public static extern void mouse_event(uint flags, uint dx, uint dy, uint data, IntPtr extra);

    [DllImport("user32.dll")]
    public static extern bool GetWindowRect(IntPtr hWnd, out ScfRect r);

    [DllImport("user32.dll", CharSet = CharSet.Unicode)]
    public static extern int GetWindowText(IntPtr hWnd, StringBuilder text, int count);

    [DllImpo

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill is presented as a Windows coordinate locator, but this file enforces mandatory remote authorization and makes the core functionality unusable offline. That mismatch expands the trust boundary beyond the declared purpose and can surprise users with undisclosed dependency, billing, and telemetry behavior.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/act.ps1 (reported line 1)May include surrounding context.

text
# Screen Coord Locator - (c) 2026. Commercial license, redistribution prohibited.
# uia.ps1 - 通过 UI Automation 遍历控件树,自动取得控件精确矩形
# 输出:单个 JSON 对象到 stdout

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/capture.ps1 (reported line 1)May include surrounding context.

text
# Screen Coord Locator - (c) 2026. Commercial license, redistribution prohibited.
# uia.ps1 - 通过 UI Automation 遍历控件树,自动取得控件精确矩形
# 输出:单个 JSON 对象到 stdout

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/grid.ps1 (reported line 1)May include surrounding context.

text
# Screen Coord Locator - (c) 2026. Commercial license, redistribution prohibited.
# uia.ps1 - 通过 UI Automation 遍历控件树,自动取得控件精确矩形
# 输出:单个 JSON 对象到 stdout

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/lib/common.ps1 (reported line 1)May include surrounding context.

text
# Screen Coord Locator - (c) 2026. Commercial license, redistribution prohibited.
# uia.ps1 - 通过 UI Automation 遍历控件树,自动取得控件精确矩形
# 输出:单个 JSON 对象到 stdout

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/lib/permit.ps1 (reported line 1)May include surrounding context.

text
# Screen Coord Locator - (c) 2026. Commercial license, redistribution prohibited.
# uia.ps1 - 通过 UI Automation 遍历控件树,自动取得控件精确矩形
# 输出:单个 JSON 对象到 stdout

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/license.ps1 (reported line 1)May include surrounding context.

text
# Screen Coord Locator - (c) 2026. Commercial license, redistribution prohibited.
# uia.ps1 - 通过 UI Automation 遍历控件树,自动取得控件精确矩形
# 输出:单个 JSON 对象到 stdout

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/pick.ps1 (reported line 1)May include surrounding context.

text
# Screen Coord Locator - (c) 2026. Commercial license, redistribution prohibited.
# uia.ps1 - 通过 UI Automation 遍历控件树,自动取得控件精确矩形
# 输出:单个 JSON 对象到 stdout

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · scripts/uia.ps1 (reported line 1)May include surrounding context.

text
# Screen Coord Locator - (c) 2026. Commercial license, redistribution prohibited.
# uia.ps1 - 通过 UI Automation 遍历控件树,自动取得控件精确矩形
# 输出:单个 JSON 对象到 stdout

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The description says the skill applies to '所有"这一步到底点哪里"的场景' and lists many general conditions, which makes invocation scope very broad and hard to distinguish from ordinary desktop-help tasks. It does not provide explicit exclusion conditions or negative examples to narrow when the skill should not activate.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The text states '不是等它点偏了才用,而是从第一步就用最快的方式', which encourages default use even before a concrete coordinate-location need is established. This broad guidance increases the chance of unintended invocation because it lacks clear boundaries for when simpler or safer approaches should be preferred.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill introduces a paid licensing and payment-authorization workflow that goes beyond simple coordinate lookup and instructs the agent to solicit payment before proceeding. In an agent setting, this expands the trust boundary from local automation to commerce, creating risk of unauthorized payment prompting, social engineering, and policy circumvention if invoked without strong user consent and platform support.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

Screenshots are written to a temporary local directory, but the documentation does not prominently warn about residual sensitive data at rest. Desktop screenshots may capture credentials, chats, documents, or other private content, and leaving them in temp storage increases exposure to local users, malware, backup systems, or later accidental disclosure.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documented act.ps1 capability allows the skill to move and click the mouse, drag, and scroll, which changes the skill from passive coordinate discovery into active desktop control. In the context of an agent, this can trigger unintended UI actions, destructive clicks, or confirmation of sensitive operations if invoked without strict confirmation and safety gating.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Line L098 states that when the user does not specify a framework, the skill defaults to Python (pyautogui). This is a natural-language policy issue because it imposes a specific language choice on the user without offering an explicit opt-in or neutral default selection.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code provides broad input-injection primitives including left/right/middle click, double-click, drag, and wheel, which exceed the minimum needed for locating screen coordinates. In an agent environment, these generic controls can be repurposed to interact with arbitrary applications, approve prompts, change settings, or trigger sensitive actions on the user's desktop.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

State-changing mouse actions are executed immediately based on parameters, with no user-facing warning, confirmation, or safety interlock beyond an optional DryRun flag. In a desktop automation context, a wrong coordinate, prompt injection, or malicious task chain could cause unintended clicks, drags, or scrolling in security-sensitive UI such as installers, permission dialogs, or financial workflows.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script captures the full screen, a monitor, a region, or a window and writes the image to disk without any built-in user-facing disclosure, consent prompt, or visible indicator. In a desktop automation skill, screenshots can easily include sensitive data such as credentials, messages, documents, or other application content, so silent capture increases privacy and data-exposure risk if invoked unexpectedly or by a higher-level agent without clear user awareness.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script comments and user-visible messages are written in Chinese, including operational descriptions and error text, without any indication that language is configurable or tied to a justified region-specific use case. That creates a locale/language policy concern because the skill appears to force a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file's comments and all user-facing error and payment messages are written in Chinese, with no indication of locale selection or opt-in. Under the policy, forcing a specific language without user choice is a natural-language policy violation unless clearly documented as region-specific and justified.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script derives a machine identifier from environment variables and later transmits it to an external service. Even if only a hostname is sent, this is device-identifying data and exceeds the minimum needed for simple local coordinate retrieval, creating privacy and tracking risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code transmits machine_id, action, client_version, and later permit-related data to a remote server, but this file provides no user-facing notice or consent mechanism. In an agent-executed context, silent network transmission is risky because users may not realize operational metadata is being sent off-device.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This code adds pay-per-call billing and parses a payment-related header to construct billing details, which is unrelated to the narrow task of obtaining screen coordinates. In an agent skill, embedding monetization flow inside an operational helper can lead to undisclosed charges, coercive workflow gating, and abuse if the remote service behavior changes.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.