Back to skill

Security audit

Mac Desktop Automation

Security checks for vulnerabilities and agentic risk

Overview

This desktop automation skill is not deceptive, but it gives broad control over the user's desktop and sensitive screen or clipboard data without consistently enforced approval boundaries.

Install only if you are comfortable granting the skill broad control over your active desktop session. Use failsafe, prefer explicit approval, avoid sensitive windows, credentials, private messages, and financial or admin workflows, and treat saved screenshots and logs as potentially sensitive.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
__init__.py:132
Finding

Approval Mode Does Not Protect Several Sensitive Desktop Operations

Content
View full analysis

Vulnerability Details

File Location: __init__.py:132-140, __init__.py:198-228, and __init__.py:348-380
Vulnerability Type: Inconsistent authorization enforcement
Risk Level: High

Vulnerable Code

python
def scroll(self, clicks: int, direction: str = 'vertical',
           x: Optional[int] = None, y: Optional[int] = None) -> None:
    if x is not None and y is not None:
        pyautogui.moveTo(x, y)

    if direction == 'vertical':
        pyautogui.scroll(clicks)
    else:
        pyautogui.hscroll(clicks)
    logger.debug(f"Scrolled {direction} {clicks} clicks")
python
def key_down(self, key: str) -> None:
    """Press and hold a key without releasing."""
    pyautogui.keyDown(key)
    logger.debug(f"Key down: '{key}'")

def key_up(self, key: str) -> None:
    """Release a held key."""
    pyautogui.keyUp(key)
    logger.debug(f"Key up: '{key}'")

def screenshot(self, region: Optional[Tuple[int, int, int, int]] = None,
               filename: Optional[str] = None):
    img = pyautogui.screenshot(region=region)

    if filename:
        img.save(filename)
        logger.info(f"Screenshot saved to: {filename}")
    else:
        logger.debug(f"Screenshot captured (region={region})")
        return img
python
def get_from_clipboard(self) -> Optional[str]:
    try:
        import pyperclip
        text = pyperclip.paste()
        logger.debug(f"Got from clipboard: '{text[:50]}...'")
        return text
    except ImportError:
        logger.error("pyperclip not installed. Run: pip install pyperclip")
        return None
    except Exception as e:
        logger.error(f"Error getting clipboard: {e}")
        return None

Technical Analysis

The controller advertises require_approval=True as a mechanism that requires user confirmation before desktop actions. Mouse movement, clicking, text entry, and ordinary key presses invoke _check_approval(), but several security-sensitive methods do not.

The un ...[truncated 1816 chars]

Remediation
View remediation

Remediation Suggestions

  1. Invoke _check_approval() before every operation that reads from or modifies the desktop environment, including screenshots, clipboard access, scrolling, window operations, key holds, and dialogs.
  2. Separate authorization checks by capability, such as screen_read, clipboard_read, clipboard_write, keyboard_input, and window_control.
  3. Make approval mode fail closed. If approval cannot be obtained, the operation should raise an authorization exception rather than silently continuing.
  4. Require explicit approval for full-screen captures and clipboard reads even when ordinary mouse movements have been preapproved.
  5. Ensure convenience functions preserve the caller's security configuration instead of creating a default global controller with approval disabled.
  6. Add unit tests that instantiate the controller with require_approval=True, mock the approval response, and verify that no underlying PyAutoGUI or clipboard function is reached after denial.
  7. Consider enabling approval mode by default in autonomous workflows and allowing narrowly scoped, time-limited approvals rather than unrestricted access.

T09 · Insecure Skill Coding Practices

Warning
Location
__init__.py:168
Finding

Typed and Clipboard Content Is Exposed Through Application Logs

Content
View full analysis

Vulnerability Details

File Location: __init__.py:168-170, __init__.py:356-358, and __init__.py:372-375
Vulnerability Type: Sensitive information exposure through logging
Risk Level: Medium

Vulnerable Code

python
if self._check_approval(f"type text: '{text[:50]}...'"):
    pyautogui.write(text, interval=interval)
    logger.info(
        f"Typed text: '{text[:50]}{'...' if len(text) > 50 else ''}' "
        f"(interval={interval:.3f}s)"
    )
python
try:
    import pyperclip
    pyperclip.copy(text)
    logger.info(f"Copied to clipboard: '{text[:50]}...'")
except ImportError:
    logger.error("pyperclip not installed. Run: pip install pyperclip")
except Exception as e:
    logger.error(f"Error copying to clipboard: {e}")
python
try:
    import pyperclip
    text = pyperclip.paste()
    logger.debug(f"Got from clipboard: '{text[:50]}...'")
    return text
except ImportError:
    logger.error("pyperclip not installed. Run: pip install pyperclip")
    return None
except Exception as e:
    logger.error(f"Error getting clipboard: {e}")
    return None

Technical Analysis

The controller records up to the first 50 characters of text entered through keyboard automation and data copied to or read from the clipboard. These values are attacker-relevant because desktop automation is commonly used to enter credentials, authentication tokens, private messages, form data, and personal information.

The documentation explicitly demonstrates password entry, making sensitive values a foreseeable input rather than an exceptional case. INFO-level messages are especially exposed because logging.basicConfig(level=logging.INFO) is configured at module import. Debug-level clipboard-read messages can also become visible when an application enables verbose logging.

Truncation does not constitute effective redaction. Passwords, API keys, one-time codes, and many session tokens can be completely disclosed within the first ...[truncated 1084 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove text and clipboard values from all log messages.
  2. Record only non-sensitive metadata, such as operation type, character count, result status, and timing.
  3. Replace content-bearing messages with forms such as Typed text (length=24) and Read clipboard text (length=120).
  4. Do not depend on truncation, pattern-based masking, or debug-level suppression as the primary protection.
  5. Avoid configuring the host application's root logger through logging.basicConfig() in a reusable library.
  6. Add automated tests that submit representative passwords, tokens, and personal data and assert that none of those values appear in captured logs.
  7. Document that clipboard and keyboard operations may process secrets and should not be instrumented with content-level telemetry.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:56
Finding

Installation Instructions Use Unpinned Privileged Desktop-Control Dependencies

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:56, SKILL.md:618, and QUICK_REFERENCE.md:250
Vulnerability Type: Unpinned third-party dependency installation
Risk Level: Medium

Vulnerable Code

bash
pip install pyautogui pillow opencv-python pygetwindow

The quick-reference installation command additionally includes clipboard access:

bash
pip install pyautogui pillow opencv-python pygetwindow pyperclip

Technical Analysis

The installation instructions request packages by name without exact versions, integrity hashes, a lock file, or a documented trusted package index. Consequently, the effective dependency set can change after the Skill has been audited.

These dependencies operate in a high-impact context: they can capture the screen, inject keyboard and mouse input, inspect windows, process images, and access the clipboard. A compromised package release, compromised package-index account, dependency confusion condition, or unexpected incompatible update would execute with the privileges of the Python environment during installation or import.

No evidence was found that the named packages are themselves malicious. The vulnerability is the mutable and unverifiable supply-chain resolution process prescribed by the project.

Attack Path

  1. A user follows the documented pip install command.
  2. Pip resolves the latest package releases and transitive dependencies available from its configured index.
  3. An upstream package or dependency has been compromised, replaced, or altered after this project was reviewed.
  4. Malicious installation hooks or imported package code execute under the user's account.
  5. Because the installed components support desktop-wide automation, compromised code can potentially access screen contents, clipboard information, files available to the user, and active applications.

Impact Assessment

Successful dependency compromise could provide arbitrary code execution with the ...[truncated 464 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin every direct dependency to a reviewed exact version.
  2. Generate and distribute a lock file that captures transitive dependencies.
  3. Require package hashes using a hash-locked requirements file and pip install --require-hashes.
  4. Document the expected package index and discourage installation from untrusted mirrors.
  5. Add automated dependency scanning and update pins only after security review and compatibility testing.
  6. Install the Skill in an isolated virtual environment under a non-administrative account.
  7. Consider separating optional features so clipboard, OpenCV, and window-control packages are installed only when their capabilities are required.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (18)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The guide states the agent 'takes a screenshot of the result' as part of autonomous operation but does not warn that screenshot capture and later screen analysis may collect sensitive on-screen data such as passwords, messages, financial details, or personal documents. Because screenshots are also presented as saved artifacts, this creates a clear confidentiality risk if users are unaware of collection, retention, or sharing behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The guide shows how to disable the failsafe and labels it 'Fast mode' without explaining the resulting loss of emergency-stop protection. For a system that performs autonomous mouse and keyboard control, removing failsafe protections can turn mis-planned or runaway actions into sustained harmful behavior, including destructive clicks, typing, or navigation that a user cannot quickly interrupt.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description emphasizes mouse, keyboard, and screen control, but the documented behavior also includes window enumeration/activation, clipboard read/write, and approval-style interactive prompts. That mismatch can cause downstream systems or users to under-scope the skill's privileges and privacy impact, increasing the risk of unintended data access or misuse.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module-level convenience functions instantiate a global controller with default settings, which means require_approval remains False unless a caller explicitly constructs a controller differently. This effectively bypasses the class's optional safety configuration and enables immediate mouse, keyboard, and screenshot actions without warning, making accidental or unauthorized desktop manipulation much easier.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The guide explicitly promotes broad natural-language control where the agent 'understands what you want' and 'figures out how to do it autonomously' without describing boundaries, confirmation gates, or permitted action scopes. In a desktop-control skill, this ambiguity can cause unintended app launches, clicks, typing, or file interactions based on underspecified prompts, increasing the chance of unsafe or privacy-impacting actions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Examples like 'Type Hello World in Notepad,' 'Write an email saying thank you,' and similar everyday phrases normalize underspecified commands that may act on whichever application is active or whichever account/session is open. In an autonomous desktop agent, such examples encourage unsafe usage patterns that can send messages, alter documents, or enter sensitive text into the wrong destination.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The autonomous text-entry section says the agent 'figures out where to type' and may use the active text editor, but it does not warn that this can inject text into the wrong window, sensitive fields, chats, shells, or admin prompts. In a desktop automation context, misdirected keystrokes can cause data leakage, accidental commands, or unintended communication.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The quick reference presents high-impact desktop automation capabilities and examples such as clipboard exfiltration, window switching, file selection, search-and-replace, launching applications, and disabling failsafes, but it does so as frictionless copy-paste guidance without prominent risk boundaries or misuse warnings. In a desktop-control skill, these examples materially lower the barrier to unsafe automation that can alter user data, interact with sensitive apps, or capture on-screen information if invoked in the wrong context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Direct control of mouse, keyboard, window focus, scrolling, hotkeys, and drag-and-drop can trigger destructive or irreversible user actions, including deleting files, submitting forms, changing settings, or interacting with privileged applications. Although the skill mentions failsafe and approval options later, the capability overview does not clearly warn that these actions can materially alter the user's system if invoked incorrectly or maliciously.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Screenshot capture and clipboard access can expose highly sensitive information such as passwords, tokens, personal messages, documents, or other on-screen secrets. The skill documents these features without a prominent privacy warning, consent boundary, or guidance on limiting capture scope, which makes accidental or abusive collection more likely.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The screenshot method can capture the full screen or arbitrary regions and optionally save the image to disk without any user-facing warning by default. This is dangerous because screenshots may include secrets, personal data, active conversations, or credentials visible on screen, and the saved file can persist sensitive data beyond the current session.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

Window enumeration and activation broaden the skill's capabilities beyond basic mouse/keyboard/screen control into discovery of user activity and targeting of specific applications. This can facilitate focused interaction with sensitive windows such as password managers, email, terminals, or admin tools, especially because activation happens without approval checks.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The clipboard helpers extend the skill from direct desktop input/output automation into access to user data that may contain passwords, tokens, copied documents, or other sensitive content. In this implementation, clipboard reads and writes occur without approval gating, scope restriction, or user-facing disclosure, so another component using this skill could silently inspect or overwrite sensitive clipboard contents.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Reading from the clipboard can silently expose highly sensitive transient user data, including passwords, API keys, private messages, or copied files/URLs. Because get_from_clipboard performs the read without approval or disclosure, any consumer of this skill can access that data with little friction.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The agent autonomously launches applications, types text, draws, and can trigger file-affecting behaviors such as opening editors and invoking save-related workflows without presenting a clear warning, confirmation, or scoped permission model to the user. In a desktop automation skill, this is dangerous because task text can cause unintended modification of user data or interaction with the wrong window if focus changes, leading to accidental data loss or unauthorized edits.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code captures screenshots of the live desktop before and after each step and stores them in the returned result structure without an explicit privacy notice or user consent. Desktop screenshots can contain sensitive information such as emails, documents, credentials, chats, or other applications unrelated to the requested task, so automatic collection increases privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The screenshot step writes an image file to disk using a default filename without explicit user confirmation. Persisting screenshots to disk creates a durable copy of potentially sensitive desktop content, which increases the chance of later exposure through local access, backups, syncing, or accidental sharing.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest describes advanced desktop automation with mouse, keyboard, and screen control, but this demo also exercises clipboard access and modification. Clipboard reading and overwriting are distinct data-access capabilities that are not implied by mouse/keyboard/screen control alone.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.