Install
openclaw skills install @drumrobot/web-browserEnvironment-aware browser operations. Detects wmux/cmux/tmux and routes to the right backend (wmux/cmux panel → user-visible, plain → Playwright MCP, chrome-devtools → reuse the user's real logged-in session). Topics: ui-test - snapshots, click/fill/verify, closed shadow DOM cascade diagnosis (cdp-trace) [ui-test.md, cdp-trace.md]. credential-issue - open service login via detected backend → wait for user sign-in → issue OR refresh an access key / token / secret / OAuth scope → hand off to follow-up automation (aws-cli, gh secret set, gh auth refresh, etc.) [credential-issue.md]. Covers both new issuance and existing-token scope expansion (PAT scope add, OAuth re-authorize, device-code). Use for: "UI check", "browser test", "screen verify", "Playwright test", "shadow DOM cascade", "::part not working", "CDP trace", "issue token", "service credential", "open login screen", "PAT refresh", "scope expansion", "device-code auth", "browser device-code".
openclaw skills install @drumrobot/web-browserEnvironment-aware browser operations skill. Detects the runtime environment and routes to the
appropriate browser backend, then runs one of two workflows: UI testing/verification (ui-test) or
browser-login-assisted credential issuance (credential-issue).
| Topic | Description | Guide |
|---|---|---|
| ui-test | Snapshot analysis, click/fill/verify, page-state diagnosis | ui-test.md |
| cdp-trace | CDP-based closed shadow DOM cascade diagnosis (DOM.getDocument pierce:true + CSS.getMatchedStylesForNode) | cdp-trace.md |
| credential-issue | service+command param → open login screen → wait for user login → issue access key/token/secret → hand off to automation | credential-issue.md |
web-browser (Step 0: environment detection — shared by all topics)
├─→ ui-test (UI verification)
│ └─→ cdp-trace (extends ui-test for closed shadow DOM)
└─→ credential-issue (browser-login-assisted token/key issuance)
└─→ chrome-devtools backend preferred (reuses the user's real logged-in session)
ui-test, cdp-trace are the UI-testing family.credential-issue reuses the same backend routing + the user-visibility rule, generalized into a
service+command parameterized auth flow.sso-verify) is not included in this skill — it remains in a
separate local-only sso-verify skill (user-environment specific, untracked).The primary purpose of browser diagnosis/verification is "the user sees it on their own screen". Screenshot capture is supporting evidence, not a substitute for visibility.
| # | Don't | Do |
|---|---|---|
| 1 | Launch with chromium.launch({ headless: true }) and only attach a screenshot in chat | chromium.launch({ headless: false, slowMo: 500 }) — let the user follow in real time |
| 2 | "I showed the user a screenshot, so it's fine" | screenshot ≠ visible to the user. If the user says "show me", open a visible browser + slowMo |
| 3 | wmux/cmux/Playwright MCP disconnected → fall back to headless CLI | Even on CLI fallback, force headless: false. On a Windows desktop OS, a chromium GUI is available |
| 4 | "headless is faster and more stable by default" mindset | Speed costs user visibility. If the user says "show me", visibility wins |
| 5 | Playwright MCP disconnected → CLI fallback auto-selects headless | CLI fallback is also headless: false. headless is only for explicit non-interactive cases (e.g., CI assertion) |
| 6 | SaaS/API task lacks credentials → fallback to manual user UI operation | Do NOT recommend manual user UI clicking when API access is available; fallback to credential-issue topic to issue token/key first |
When a task can be performed via API (e.g., Google Forms API, GitHub API, AWS API), but required API tokens or access keys are missing in the environment, do NOT recommend manual user UI clicking or surrender to direct manual UI operation. You MUST recommend credential-issue topic to issue the access key/token via browser login first, then proceed with backend API automation.
| # | Don't | Do |
|---|---|---|
| 1 | API token missing → "Please edit/click manually on the website" | Recommend credential-issue topic to issue API token/key via browser login |
| 2 | Direct UI automation fails → fallback to manual user operation | Check if API automation is available → issue credential via credential-issue → execute API |
headless: falseheadless: falseDuring a closed shadow DOM ak-library cascade investigation, used a npx playwright Bash invocation + chromium.launch({ headless: true }) and only attached a screenshot in chat. The user requested "show it via web-ui-test" and no visible browser was provided. The user reacted angrily that the Chromium UI never appeared.
When a capture/documentation task (report evidence, purchase/registration flow guide, etc.) hits a screen that requires login, and completing that login would reveal materially different information than what's already captured (e.g., the real final price vs. a promotional pre-login price, actual post-login UI state vs. an assumption), do NOT silently stop and paper over the gap with a deferral disclaimer. Ask the user via AskUserQuestion whether to continue (via interactive login in a visible backend) or whether the pre-login capture is sufficient for the purpose at hand.
| # | Don't | Do |
|---|---|---|
| 1 | Hit a login wall → write "please have finance/ops enter payment details themselves for security" and stop, without asking | Decompose the remaining flow: payment/credential entry should be deferred to the user/business owner, but login + viewing the resulting screen is often just informational — ask which is actually needed before deciding to stop |
| 2 | Treat "login" and "entering payment info" as one bundled decision to skip together | They are different risk levels. Login-then-observe (e.g., see the real cart/checkout price) does not require entering card/account credentials — only the latter needs deferral |
| 3 | Report a pre-login/promotional price or state as if it were final, without flagging the gap | If the login-gated final screen wasn't verified, explicitly flag it ("actual payment screen not verified — may differ from the listed price") instead of presenting the pre-login figure as authoritative |
| 4 | Assume the backend can't support interactive login without checking | Check chrome-devtools connection + visibility (per credential-issue.md "Fresh-login flow") first; if visible, open the page there and have the user sign in in that same window, then continue capturing |
| 5 | Decide unilaterally that "this is good enough" when the report's factual accuracy depends on the gated screen | If the gap could make a delivered report/guide factually wrong (e.g., a payment-request report citing a price that turns out incorrect), the stop-vs-continue decision belongs to the user, not the assistant |
While building a domain-registration payment-request report, captured the domain-search-result page (showing a promotional price) and the login screen, then stopped at the login wall with a disclaimer ("have finance/ops enter payment details"), never asking whether to continue via login to verify the real checkout price. The report's stated price differed from the actual payment-screen price. User feedback (paraphrased): "don't arbitrarily skip capturing screens that require login — ask first."
Some SaaS portals allow full browser login automation but selectively trigger CAPTCHA challenges on creation/mutation actions (not just on login). Document confirmed cases here so agents do not repeat failed automation attempts.
| Service | Automatable | CAPTCHA-blocked | Fallback |
|---|---|---|---|
| Discord Developer Portal | Login (via persistent profile with saved credentials) | New application creation, bot token reset | Keep browser visible (headless: false); user handles hCaptcha manually; script polls page.url() for /bot URL and auto-captures token once user navigates there |
| Discord Developer Portal | Reading existing app info, navigating between tabs | (same) | (same) |
launchPersistentContext) with a saved user data directory
retains Discord session cookies. Navigation to discord.com/developers/applications succeeds
without re-authentication.force: true checkbox click and JS dispatchEvent
workarounds successfully activate the Create button, but Discord's backend detects the automated
browser and intercepts submission with CAPTCHA.headless: false + launchPersistentContext (reuses login session).page.url() for the /bot path (every 2s, max ~4min timeout)./bot URL is detected, script resumes: click "Reset Token" → capture input[readonly] value → enable [role="switch"] intents → save changes.Check environment variables AND CLI presence to determine the browser backend:
# wmux
echo "WMUX=$WMUX"
# cmux — detect via ANY of these (cmux app does NOT set CMUX_SESSION; use multi-var OR)
echo "CMUX_BUNDLE_ID=$CMUX_BUNDLE_ID"
echo "CMUX_PANEL_ID=$CMUX_PANEL_ID"
echo "CMUX_BUNDLED_CLI_PATH=$CMUX_BUNDLED_CLI_PATH"
# CLI fallback (env may be unset in nested shells but CLI still works)
command -v cmux && echo "cmux CLI present"
command -v wmux && echo "wmux CLI present"
| Environment | Detect (ANY true → environment matches) | Do (use this) | Don't (forbidden) |
|---|---|---|---|
| wmux | $WMUX set OR command -v wmux succeeds | wmux browser open/snapshot/click/type commands via Bash | Playwright MCP — user cannot see the invisible Playwright window |
| cmux | $CMUX_BUNDLE_ID set OR $CMUX_PANEL_ID set OR $CMUX_BUNDLED_CLI_PATH set OR command -v cmux succeeds (e.g. /Applications/cmux.app/Contents/Resources/bin/cmux) | cmux browser panel commands | Playwright MCP — same reason |
| Plain / tmux | None of wmux/cmux signals present | Playwright MCP (Step 1 below) | — |
cmux app sets several env vars when launching a shell, but CMUX_SESSION is NOT one of them (a legacy guess by analogy with WMUX). Real vars observed in a cmux-launched shell:
CMUX_BUNDLE_ID (e.g. com.cmuxterm.app)CMUX_PANEL_ID (UUID per panel)CMUX_BUNDLED_CLI_PATH (CLI absolute path)CMUX_SHELL_INTEGRATION_DIRCMUX_AGENT_LAUNCH_*GHOSTTY_RESOURCES_DIR (cmux uses Ghostty-based terminal)CMUX_SOCKET is set but often empty — do not use it as the sole signal. Use the OR matrix above.
| # | Don't (single-var assumption) | Do (multi-var OR) |
|---|---|---|
| 1 | [ -n "$CMUX_SESSION" ] only check → false negative on cmux app | OR across CMUX_BUNDLE_ID / CMUX_PANEL_ID / CMUX_BUNDLED_CLI_PATH |
| 2 | Use CMUX_SOCKET as detection (empty in many cases) | Treat empty CMUX_SOCKET as no-signal; rely on the 3 vars above + CLI presence |
| 3 | Assume cmux env var name mirrors wmux (*_SESSION) | Verify against actual cmux app shell environment — vars differ per terminal multiplexer |
When $WMUX is set, use these instead of Playwright MCP.
Invocation form: the rest of this document uses the bare wmux browser … form, which is what runs when wmux is on PATH (the common case). If wmux is not on PATH in the current environment, substitute node "$WMUX_CLI" for wmux in every command below — $WMUX_CLI points to the same entry point. The two forms are interchangeable; pick whichever resolves on the current shell and use it consistently.
wmux browser open <url> # navigate (= playwright navigate)
wmux browser snapshot # get accessibility tree with @eN refs
wmux browser click @eN # click element
wmux browser type @eN <text> # type into element
wmux browser fill @eN <value> # set input value
wmux browser get-text # get page text
wmux browser screenshot # capture screenshot
wmux browser eval <js> # run JavaScript
wmux browser back # go back
wmux browser forward # go forward
wmux browser reload # reload page
Workflow: browser open <url> → browser snapshot → read tree → browser click/type @eN → browser snapshot again.
Refs (@e1, @e2...) expire after page changes — always re-snapshot.
| Action | wmux (Do) | Playwright MCP (Don't in wmux) |
|---|---|---|
| Navigate | Bash("wmux browser open <url>") | mcp__playwright__browser_navigate |
| Snapshot | Bash("wmux browser snapshot") | mcp__playwright__browser_snapshot |
| Click | Bash("wmux browser click @eN") | mcp__playwright__browser_click |
| Type | Bash("wmux browser type @eN text") | mcp__playwright__browser_type |
| Screenshot | Bash("wmux browser screenshot") | mcp__playwright__browser_take_screenshot |
| Evaluate JS | Bash("wmux browser eval <js>") | mcp__playwright__browser_evaluate |
| Wait for text | Re-snapshot + check | mcp__playwright__browser_wait_for |
Key difference: wmux browser is visible to the user in real-time on the right panel. Playwright opens an invisible window the user cannot see.
After Step 0 backend detection, route to the topic:
| Goal | Topic | Entry |
|---|---|---|
| Verify a UI change, snapshot, click/fill | ui-test | ui-test.md |
Diagnose ::part not applying / closed shadow DOM cascade | cdp-trace | cdp-trace.md |
| Open a service login → wait for user login → issue access key/token | credential-issue | credential-issue.md |
Step execution order: Step 0 (this file — detect backend + user-visibility rule) → read the
target topic .md → follow its procedure. The topic .md files hold the actual procedures; this
file is the shared backend-detection + index.