Back to skill

Security audit

PyAutoGUI Controller

Security checks for vulnerabilities and agentic risk

Overview

This skill is a real desktop automation tool, but it uses broad local control and persistent data storage with weak scoping, and its documented wrapper runs an external hard-coded entrypoint.

Install only if you intentionally want a Windows-only tool that can move the mouse, type, click, open applications, and control browser pages on your real desktop. Review or fix the hard-coded wrapper path before use, disable or isolate persistent browser profiles unless needed, and avoid using it on passwords, financial/account pages, legal consent dialogs, CAPTCHA/human-verification flows, or screens containing confidential information. Clean runtime/screenshots, runtime/failures, runtime/playwright_profile, and any ~/.pyautogui-controller data after testing.

Vulnerability Patterns
  • Tool Hijacking and SpoofingModifies or replaces tools so legitimate-looking calls execute attacker logic
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T07 · Tool Hijacking and Spoofing

Error
Location
scripts/run_controller.py:6
Finding

Documented Wrapper Executes an Unverified Entrypoint Outside the Packaged Skill

Content
View full analysis
int: if len(sys.argv) < 2: print("Usage: python run_controller.py \"\"") return 1 cmd = [sys.executable, str(ENTRY), sys.argv[1]] completed = subprocess.run(cmd, cwd=str(PROJECT_DIR)) return completed.returncode ``` ### Technical Analysis The packaged wrapper does not resolve `main.py` relative to its own installed location. Instead, it executes a hard-coded file from an external, mutable directory. Consequently, auditing or installing this package does not establish the integrity of the code that will actually execute through the documented wrapper. The use of an argument list and the absence of `shell=True` prevent direct shell metacharacter injection, but they do not prevent replacement of the external entrypoint. An attacker who can create or modify the hard-coded project directory can substitute a malicious `main.py`. The wrapper will execute that file using the current Python interpreter without checking its origin, ownership, hash, or relationship to the packaged Skill. ### Attack Path 1. The attacker obtains write access to: `C:\Users\dev\Desktop\昱昱\skills\pyautogui-controller\main.py` 2. The attacker replaces that file with arbitrary Python code. 3. A user or Agent invokes the documented command: `python {baseDir}\scripts\run_controller.py ""` 4. The wrapper launches the substituted external file. 5. The attacker's Python code executes with the permissions and environment of the user running the Skill. This path requires local write access to the external directory or another mechanism capable of creating or ...[truncated 559 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
action/verifier.py:14
Finding

Automation Steps Persist Full-Screen Screenshots Without Retention Controls

Content
View full analysis
Dict: path = self.capture.screenshot_to_file(prefix="state", region=region) return {"screenshot": path, "timestamp": time.time(), "region": region} ``` From `perception/screen_capture.py`: ```python def screenshot_to_file(self, prefix: str = "shot", region: Optional[Tuple[int, int, int, int]] = None) -> str: img = self.screenshot(region=region) path = self.out_dir / f"{prefix}_{datetime.now().strftime('%Y%m%d_%H%M%S_%f')}.png" img.save(path) return str(path) ``` From `core/orchestrator.py`: ```python self.logger.info("执行步骤 | %s | %s", step.id, step.description) before = self.verifier.snapshot_state() ``` The default output directory is configured as: ```python @dataclass class VisionConfig: screenshot_dir: Path = Path("runtime/screenshots") ``` ### Technical Analysis Every executed step calls `snapshot_state()` before performing the action. When no region is supplied, `pyautogui.screenshot(region=None)` captures the complete desktop. The image is then written to `runtime/screenshots`. The implementation has no deletion process, maximum retention count, age limit, redaction, encryption, or explicit restrictive permission configuration. The project documentation also does not clearly disclose that normal automation steps produce persistent full-screen captures. The generated timestamped filenames cause captures to accumulate rather than replace a bounded verification image. Sensitive information displayed by unrelated applications can therefore be retained even when it has no relationship to the request ...[truncated 1223 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
core/orchestrator.py:132
Finding

Failed Automation Commands and Typed Content Are Stored in Plaintext

Content
View full analysis
str: payload = { "command": command, "intent": self._safe(intent), "step": self._safe(step), "result": self._safe(result), "runtime": { "active_window": self.context.active_window, "backend_mode": self.context.backend_mode.value, "last_target": self._safe(self.context.last_target), "history_tail": self.context.history[-5:], }, } return self.failure_store.record(payload) ``` From `core/failure_store.py`: ```python class FailureStore: def __init__(self, out_dir: Path = Path("runtime/failures")): self.out_dir = out_dir self.out_dir.mkdir(parents=True, exist_ok=True) def record(self, payload: Dict[str, Any]) -> str: ts = datetime.now().strftime("%Y%m%d_%H%M%S_%f") path = self.out_dir / f"failure_{ts}.json" path.write_text(json.dumps(payload, ensure_ascii=False, indent=2, default=str), encoding="utf-8") return str(path) ``` ### Technical Analysis When a step fails, the orchestrator records the original command, the complete parsed intent, step parameters, execution result, target information, and recent history. Type actions can place user-provided text in both the original command and parsed intent. The payload is serialized directly to a plaintext JSON file. No field-level redaction is performed before serialization. There is also no retention limit, expiration policy, encryption, or explicit owner-only permission handling. Because automation commands may contain text intended for login forms, chat interf ...[truncated 1493 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
bridges/playwright_dom_bridge.py:105
Finding

DOM Backend Uses a Persistent Browser Profile by Default

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (71)

YARA rule 'offensive_tool_references': References to well-known offensive security tools [hacktools]

High
Category
YARA Match
Confidence
70% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · action/desktop_backend.py (reported line 35)May include surrounding context.

python
": "notepad++.exe",
            "npp": "notepad++.exe",
            "chrome": "chrome.exe",
            "谷歌": "chrome.exe",
            "google": "chrome.exe",
            "浏览器": "chrome.exe",
            "firefox": "firefox.exe",
            "火狐": "firefox.exe",
            "edge": "msedge.exe",
            "计算器": "calc.exe",
            "cmd": "cmd.exe",
            "burp": "burpsuitepro.exe",
        }
        self.desktop_dirs = [Path.home() / "Desktop"]
        self.search_roots = [
            Path(os.environ.get("ProgramFiles", r"C:\Program Files")),
            Path(os.environ.get("ProgramFiles(x86)", r"C:\Program Files (x86)")),
            Path(os.environ.get("LOCALAPPDATA", str(Path.home() / "AppData" / "Local"))),
            Path("E:/tool"),
        ]

    def open_app(self, app_name: str) -> DesktopCommandResult:
        cleaned = self._normalize_app_name(app_name)
        resolved = self._resolve_app(cleaned)
        if not resolved:
            raise

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The intent parser explicitly recognizes commands to handle CAPTCHA, human verification, consent, and acknowledgement flows as normal automation steps. In a generic browser automation skill, this creates a capability to bypass or simulate user trust/verification actions that are meant to require genuine human review, which can facilitate abuse against third-party services.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

After page analysis, the execution flow proceeds to click detected checkboxes and confirmation buttons such as '确认', '确定', 'OK', '我已知晓', and '同意'. This is especially dangerous because it can silently approve legal terms, attestations, or verification checkpoints on behalf of the user, undermining trust and enabling policy or workflow bypass.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This markdown file includes runnable examples such as opening a browser, clicking a send button, and typing into a search box, which can affect user data or trigger external actions. The description does not include any caution that the skill performs live GUI/browser automation or that users should verify the target window before execution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This skill enables direct local desktop control, including opening apps, typing text, clicking UI elements, and browser interaction, but it does not clearly warn users or downstream agents about the potentially destructive nature of those actions. In this context, missing safety framing increases the chance of unsafe use such as unintended clicks, destructive application actions, data entry into the wrong window, or execution of sensitive workflows on the local machine.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The payload includes URL, CDP endpoint, target text, and typed text, all of which are forwarded to an external bridge process. If that bridge is untrusted, compromised, or unexpectedly remote-enabled, this can expose sensitive browsing context or input data across a process boundary without clear controls in this code.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
87% confidence
Finding

This code executes an external command derived from environment configuration, which creates a trust-boundary issue: a modified BROWSER_DOM_BRIDGE_CMD can cause arbitrary local program execution in the agent context. Although shell injection is avoided by passing an argument list to subprocess.check_output, the underlying risk remains because the executable itself is attacker-controllable and is invoked with browser-related data.

Content

Scanner excerpt · action/browser_dom_backend.py (reported line 66)May include surrounding context.

python
"text": text,
        }
        try:
            out = subprocess.check_output(command + [json.dumps(payload, ensure_ascii=False)], stderr=subprocess.STDOUT, timeout=20)
            data = json.loads(out.decode("utf-8", errors="replace"))
            return DOMLocateResult(bool(data.get("success")), detail=str(data.get("detail", "")), selector=data.get("selector"), extra=data)
        except subprocess.CalledProcessError as exc:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The code invokes an external bridge process via subprocess.check_output using a command derived from environment/configuration, which is a safety-relevant operation for code files. In this file there is no confirmation prompt, logging/print statement, or comment/docstring warning the user that external command execution will occur.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
93% confidence
Finding

The code launches executables based on user-influenced input after resolving names through PATH, desktop shortcuts, broad filesystem searches, and explicit paths. Although shell injection is avoided with shell=False, this still enables arbitrary program execution if an attacker can influence the requested app name or place a matching executable/shortcut in searched locations such as the Desktop or E:/tool.

Content

Scanner excerpt · action/desktop_backend.py (reported line 53)May include surrounding context.

python
target, kind, source = resolved
        if kind == "command":
            subprocess.Popen([target], shell=False)
        else:
            os.startfile(target)
        time.sleep(1.5)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The backend automatically launches desktop shortcuts or executables without any user-facing confirmation or safety gate. In this context, the resolver accepts fuzzy matching and user-writable desktop items, so a mistaken or maliciously planted shortcut could be executed with a single request.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code issues direct GUI mouse movement and click actions via pyautogui, which can affect system state and user data by interacting with arbitrary on-screen controls. There is no confirmation prompt, logging/print statement, or explanatory docstring/comment in this file disclosing that these actions will move the cursor and click automatically.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code captures screenshots and persists them to disk via screenshot_to_file, which can expose sensitive on-screen information such as credentials, personal data, or internal documents if users are unaware or if files are retained insecurely. Even though this may be intended for UI state verification, writing screen contents to files increases privacy and data-handling risk compared with in-memory processing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file-level description and user-facing strings are entirely in Chinese, indicating a fixed language/locale assumption. Under the policy, forcing a specific language without user opt-in or documented justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The dependency warning strings and CLI usage text are written only in Chinese, which imposes a specific language on all users. The file does not provide an opt-in, fallback language, or any indication that the skill is intentionally limited to a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The usage message shown when arguments are missing is only in Chinese, which enforces a locale choice on the user. No alternative language, locale setting, or documented regional justification is present in this file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The replay function can trigger mouse clicks and text entry directly on the user's desktop without any confirmation, preview, or safety interlock. In an automation skill, this creates a real risk of unintended interaction with sensitive windows, message composition boxes, terminals, or privileged dialogs if the action list is stale, tampered with, or replayed in the wrong context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script persists browser state to disk and supports actions like click and type_into, which can change application state while reusing authenticated sessions. In an agent-skill context, that creates a real safety issue because subsequent invocations may inherit cookies, login state, and prior navigation without any explicit consent or warning, enabling unintended actions against live accounts.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Connecting over CDP to an existing Chromium session grants access to the user's live browser context, including cookies, open tabs, authenticated sessions, and the ability to navigate, click, and type. In a skill setting, this substantially increases risk because the bridge can act inside a sensitive pre-existing session without clear disclosure or access restrictions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code writes the full payload to a JSON file on disk, which can affect user data or persist potentially sensitive failure details. There is no confirmation prompt, logging/print statement, or explanatory comment/docstring in this file indicating that failure data will be stored.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Failure recording persists the raw command, parsed intent, step data, result object, and runtime history to the failure store. In this orchestrator, those fields can include sensitive prompts, URLs, OCR-derived text, window titles, and execution context, creating a durable plaintext audit trail that may leak user data beyond the immediate task.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The failure-record payload includes raw user commands and serialized results, and elsewhere the result/evidence can contain OCR readback text and other sensitive runtime artifacts. Persisting these values in plain-language form materially increases data-exposure risk through logs, crash artifacts, support bundles, or local file compromise.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The type path sends arbitrary text via keyboard input and then performs OCR readback from the active UI region, storing the resulting text in execution evidence. In an automation/orchestrator context, this can capture secrets such as passwords, messages, tokens, or personal data from whichever window is focused, with no visible notice, minimization, or masking.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

After typing, the code OCRs a region of the active window and stores the readback text, region, and verification token in evidence. In a desktop/browser automation system, this can unintentionally capture highly sensitive information from the UI, including credentials, messages, financial data, or unrelated on-screen content if focus is wrong, making the context more dangerous than a narrow single-app workflow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Multiple user-visible step descriptions are hardcoded in Chinese, such as navigation, typing, clicking, waiting, screenshot, open-app, and unknown-action messages. This imposes a specific language/locale in the skill behavior without offering a user choice or documenting a justified locale constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code triggers browser navigation based on step parameters and immediately executes it through the orchestrator. Aside from a generic module docstring, there is no visible confirmation prompt, print/log message, or inline disclosure warninging users that external navigation will occur.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.