Screen Vision

AI screen vision and desktop computer control skill for OpenClaw. Let your AI agent see the screen, understand UI elements, and autonomously perform mouse and keyboard operations (click, type, scroll, drag) via a screenshot-analyze-action loop. Cross-platform: Linux (headless server with XFCE4+noVNC, or desktop), macOS (cliclick), Windows (pyautogui). Supports any OpenAI-compatible vision API (SiliconFlow, OpenAI, DashScope, Zhipu, Ollama, etc.). Smart diff detection saves tokens. Safety mechanisms block dangerous operations. Trigger: when user asks to operate/control computer, view/interact with screen, open applications, browse websites, fill forms, perform desktop GUI tasks, take screenshots, or any task requiring visual screen understanding and desktop automation. Keywords: computer use, screen control, desktop automation, GUI agent, visual agent, 屏幕操控, 桌面控制, 视觉代理, 截屏, 自动化操作, 远程桌面, AI操控电脑, screen vision, desktop agent, mouse keyboard automation, UI automation, computer use agent, CUA, computer control, screen interaction. Examples: "打开Chrome搜索天气", "看看屏幕上有什么", "帮我操作电脑", "截个屏", "open browser and search weather", "click the submit button", "take a screenshot", "fill this form", "帮我打开微信", "操作电脑下载文件", "远程操作桌面".

Install

openclaw skills install @guitu917/ai-screen-vision