T09 · Insecure Skill Coding Practices
- Location
scripts/desktop_agent.py:277- Finding
Arbitrary Command Execution Through Shell-Based Application Launching
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill can control a Windows desktop, capture remote-access details, and send WeChat messages or files without strong confirmation or containment.
Review before installing. Only run this in a Windows account or VM where desktop control, screenshots, clipboard use, WeChat sends, file transfers, and ToDesk credential exposure are acceptable. Require fixes for strict app allowlisting, removal of shell=True fallback, blocking safety checks, explicit confirmations for remote credentials and WeChat sends, and secure screenshot handling before using it on a sensitive machine.
scripts/desktop_agent.py:277Arbitrary Command Execution Through Shell-Based Application Launching
scripts/safety.py:34Dangerous Operations Are Detected but Explicitly Allowed
scripts/presets.py:14Sensitive Desktop and Remote-Access Credentials Are Stored in Predictable Plaintext Screenshots
SKILL.md:16Third-Party Dependencies Are Installed Without Version or Integrity Pinning
The code generally matches the declared desktop automation purpose: it can capture screenshots, click, type, press keys, open apps, inspect windows, and do image matching for GUI automation. However, the description explicitly claims WeChat message/file sending and ToDesk remote connection capabilities, which are not implemented in this code chunk. While WeChat can be launched via open_app, there is no dedicated logic for composing/sending messages or files, nor any ToDesk-specific connection workflow. Additionally, the description implies a higher-level agent controlled by natural language or preset commands, but this file only contains low-level action methods and a simple CLI dispatcher. Therefore the description overstates capabilities relative to the supplied code.
The description presents a broad Windows desktop automation agent capable of general GUI control and multiple desktop actions. The supplied code chunk, however, is a limited preset-task script focused specifically on ToDesk and WeChat. While those two declared scenarios are partially represented, the ToDesk flow is materially different from the description: it launches ToDesk, takes a screenshot, and returns a fixed device code, leaving password recognition to an external caller rather than actually obtaining full credentials. Also, the code does not expose the broader general-purpose desktop automation capabilities described. This is best classified as a mismatch because the actual behavior is substantially narrower and partially different in its primary claimed workflow.
The description presents the skill as an active Windows desktop automation agent. However, the supplied code chunk is only a safety restriction module. Its functions inspect text and execution plans for dangerous keywords, warn on risky content, and reject overly long plans. It does not interact with the desktop, launch or control applications, send messages, or create remote connections. Because the actual behavior is materially different from the declared primary purpose, this is a clear description-behavior mismatch.
The skill advertises broad activation for arbitrary desktop control, which can cause the agent to invoke this high-risk skill for vague user requests. In the context of GUI automation with messaging, screenshots, and remote-access actions, ambiguous routing materially increases the chance of unauthorized or unsafe actions being triggered.
The documentation describes screenshot capture, password recognition, and transmission of remote-access details without a clear warning that sensitive data may be collected and disclosed. In a desktop-automation skill, this omission undermines informed consent and can lead to exposure of credentials, personal messages, or other on-screen secrets.
The workflow explicitly instructs the agent to extract a dynamically generated remote-access password from a screenshot and relay it in plain language. This handles authentication secrets in an unsafe way and could expose live credentials to unauthorized parties, logs, chat history, or unintended recipients.
This procedure operationalizes collection of a device code and temporary password from the UI and instructs the agent to send them onward, effectively turning the skill into a credential-harvesting/transmission mechanism. Given the remote-control context, the impact includes unauthorized access, account compromise, and covert persistence through remote sessions.
This duplicate finding identifies the same dangerous sink: Popen(cmd, shell=True) with controllable input. The desktop-agent context increases severity because the tool is likely exposed to natural-language requests, which are especially susceptible to prompt injection and parameter manipulation.
"code": "code",
}
cmd = app_map.get(app_name.lower(), app_name)
subprocess.Popen(cmd, shell=True)
logger.info(f"启动应用: {app_name} -> {cmd}")
return True
except Exception as e:
This duplicate finding identifies the same dangerous sink: Popen(cmd, shell=True) with controllable input. The desktop-agent context increases severity because the tool is likely exposed to natural-language requests, which are especially susceptible to prompt injection and parameter manipulation.
"code": "code",
}
cmd = app_map.get(app_name.lower(), app_name)
subprocess.Popen(cmd, shell=True)
logger.info(f"启动应用: {app_name} -> {cmd}")
return True
except Exception as e:
This function intentionally captures and returns remote-access connection material, including a device code and a screenshot intended for password extraction, without any visible authorization, warning, or confirmation step. In the context of a desktop-control skill, exposing ToDesk credentials can enable unauthorized remote access to the host and is therefore especially dangerous.
The task is expressly designed to collect remote-access credentials for later AI recognition/use, which creates a direct secret-exposure pathway. Given the skill's purpose of controlling a Windows desktop and establishing ToDesk connections, this materially lowers the barrier to unauthorized takeover of the machine if invoked by an untrusted workflow or prompt.
Automated file transfer over WeChat without confirmation or validation can exfiltrate arbitrary local files to external recipients under the user's identity. In a desktop-automation skill, this is more dangerous than text messaging because sensitive documents, credentials, or proprietary data may be transmitted with a single action.
The skill documents capabilities that imply shell execution and network-adjacent remote-access workflows, but it does not declare any explicit tool scope or permission boundaries. For a desktop-control skill, missing scope restrictions increases the chance of unintended command execution, data access, or exfiltration through overbroad agent/tool access.
Trigger phrases such as '连接你' or generic remote-connection wording are ambiguous and prone to accidental invocation. Because the associated workflow can expose remote-access credentials and start a remote-control sequence, collision-prone triggers are dangerous in this skill context.
The screenshot function captures the entire desktop and persists it to C:\temp\desktop_agent automatically. Because desktop screens frequently contain credentials, chats, documents, and other sensitive material, saving full-screen captures to disk increases exposure and leaves recoverable artifacts.
The fallback text-entry path copies arbitrary text to the system clipboard before pasting it. Clipboard contents are globally accessible to other applications and may persist after use, so sensitive tokens, messages, or secrets can be exposed unintentionally.
The open_app API is documented as opening applications, but its implementation allows arbitrary command execution because unknown input is used directly as a shell command. In the context of a desktop automation agent, this materially expands capability from GUI automation to full OS command execution.
This finding reflects that application launch is actually shell-backed arbitrary command execution without meaningful safety controls. In a skill designed to control Windows and apps like WeChat or remote tools, this makes misuse more dangerous because it can pivot from intended automation into system-level command execution.
open_app passes cmd to subprocess.Popen(..., shell=True) after falling back to raw app_name when the input is not in the allowlist. In a desktop-control skill, that means an attacker or prompt-injected agent action can execute arbitrary shell commands rather than merely launching approved applications.
"code": "code",
}
cmd = app_map.get(app_name.lower(), app_name)
subprocess.Popen(cmd, shell=True)
logger.info(f"启动应用: {app_name} -> {cmd}")
return True
except Exception as e:
The docstring says the function returns a dict containing both "device_code" and "temp_password". In the implementation, it instead returns a screenshot path, a fixed device code, and a "password_needs_recognition" flag, with no temporary password included. This is an active mismatch between documented intent and actual behavior.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
agent = DesktopAgent()
# Step 1: 启动ToDesk
subprocess.Popen([TODESK_PATH])
logger.info("ToDesk已启动")
time.sleep(3)
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def wechat_ensure_running():
"""确保微信在运行"""
import subprocess as sp
r = sp.run(['tasklist'], capture_output=True, text=True)
if 'Weixin' not in r.stdout:
sp.Popen([WECHAT_PATH])
import time; time.sleep(5)
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
import subprocess as sp
r = sp.run(['tasklist'], capture_output=True, text=True)
if 'Weixin' not in r.stdout:
sp.Popen([WECHAT_PATH])
import time; time.sleep(5)
logger.info('微信已启动')
The skill can send arbitrary WeChat messages automatically with no confirmation at send time, no recipient verification, and no content review. In a desktop-agent context this increases the risk of social engineering, accidental disclosure, or abuse to message unintended contacts using the user's identity.
The task description and steps claim the preset gets both the device code and temporary password and returns the credentials. However, the handler returns only a fixed device code plus a screenshot path and a flag indicating the password still needs recognition. The registered intent therefore overstates what the code actually does.
No suspicious patterns detected.