browsing
Use when you need direct browser control - teaches Chrome DevTools Protocol for controlling existing browser sessions, multi-tab management, form automation, and content extraction via use_browser MCP tool
Browsing with Chrome Direct
Overview
Control Chrome via DevTools Protocol using the use_browser MCP tool. Single unified interface with auto-starting Chrome.
Announce: "I'm using the browsing skill to control Chrome."
When to Use
Use this when:
- Controlling authenticated sessions
- Managing multiple tabs in running browser
- Playwright MCP unavailable or excessive
Use Playwright MCP when:
- Need fresh browser instances
- Generating screenshots/PDFs
- Prefer higher-level abstractions
Auto-Capture
Every DOM action (navigate, click, type, select, eval, keyboard_press, hover, drag_drop, double_click, right_click, file_upload) automatically saves:
{prefix}.png— viewport screenshot{prefix}.md— page content as structured markdown{prefix}.html— full rendered DOM{prefix}-console.txt— browser console messages
Files are saved to the session directory with sequential prefixes (001-navigate, 002-click, etc.). You must check these before using extract or screenshot actions.
Credential-shaped pages: when a page shows a token or secret (Slack xoxb-/xapp-, GitHub ghp_/github_pat_, 1Password ops_/A3- keys, otpauth:// seeds, or any element marked data-sen-secret), no files are written for that action and the response says ⚠️ Page shows credential-shaped content; auto-capture and DOM output suppressed. with only metadata. Token-shaped values in extract/eval output are replaced by [REDACTED credential-shaped], and screenshot refuses. Capture secrets with a credential broker. On a token-shaped page, use eval only for value-blind queries (e.g. "is the token field present?"). On a page with any data-sen-secret element, eval refuses outright, and extract/attr refuse for the marked element or strip it from HTML/markdown output. The marker scan covers the top document, open shadow roots, and same-origin iframe/frame/object/embed; it does not see closed shadow roots or cross-origin frames. Markup inside an <iframe srcdoc="…"> attribute is not stripped: extract html and attr srcdoc return it verbatim.
When to mark a secret: mark its element with set_attr (data-sen-secret) as soon as the action that revealed it returns, before any other action on that page, then capture the value with the credential broker. Marking only affects later actions. The capture files the revealing action already wrote (e.g. 003-click.html/.md) still hold the value and are not deleted, so don't read them back. This is an accident guard for a cooperating agent, not a security boundary. SUPERPOWERS_CHROME_ALLOW_CREDENTIAL_CAPTURE=1 turns all of this off.
Separately (and unconditionally): a plain password or one-time code has no shape the check above recognizes, so a page's own handler mirroring one into an attribute, a hidden input, or visible text would otherwise land in the .html/.md/-diff.txt files. Whatever gets written is scrubbed on an inert clone (document.implementation.createHTMLDocument + importNode, never cloneNode on the live document): value and every data-*/aria-* attribute are stripped from input[type="password"] (case-insensitive; remembered on that same element after a "show password" toggle flips it to type="text"), any field whose autocomplete contains current-password, new-password, one-time-code, cc-number, or cc-csc (case-insensitive substring), and data-sen-secret-marked elements. A field with none of those signals is still caught if its own live .value exactly matches one of its OTHER attributes (a self-mirror) -- the case for Google's 2-step verification totpPin field, which has no recognized type or autocomplete but copies the typed code into data-initial-value itself. Self-mirror detection skips inputs nobody types into (submit, button, reset, checkbox, radio, hidden, image) and ignores attributes that name or label a field (type, name, id, aria-label, title, placeholder, for, class). The live .value of every field found any of these ways, plus its HTML-entity-escaped forms (including ), is also replaced with [REDACTED] wherever it occurs in the output -- any attribute, on any element, not just the field's own value -- if it is at least 4 characters long (3 for cc-csc). Dialog captures are redacted with the values from the most recent page capture; a value typed in the same action that opened the dialog is not. Redaction is a plain string replace, so a secret that also appears as ordinary page text (a password of password) replaces that text too; the artifact is degraded but nothing leaks. Self-mirror detection can also flag a non-secret typeable field whose value happens to equal one of its other attributes (a search box synced into data-query), with the same effect. The live page is never touched.
Known gaps (the secret is written to disk in clear): split one-digit OTP boxes feeding an aggregate hidden input (each digit is below the length floor, so the aggregate leaks); a field cleared on submit whose mirror remains; a "show password" toggle that swaps in a new element instead of changing type; a self-mirrored value shorter than the length floor (the field is still found, but its value is too short to substring-redact). Screenshots are pixels, not scrubbed text: a visible (type="text") typed value can appear in a .png.
The use_browser Tool
Single MCP tool with action-based interface. Chrome auto-starts on first use.
Parameters:
action(required): Operation to performselector(optional): CSS or XPath selector for element operationspayload(optional): Action-specific data (string or object)timeout(optional): Timeout in ms for await operations (default: 5000)
Active tab: Every action operates on the current activeTab. Use switch_tab to change it.
Actions Reference
Navigation
-
navigate: Navigate to URL
payload: URL string- Example:
{action: "navigate", payload: "https://example.com"}
-
await_element: Wait for element to appear
selector: CSS selectortimeout: Max wait time in ms- Example:
{action: "await_element", selector: ".loaded", timeout: 10000}
-
await_text: Wait for text to appear
payload: Text to wait for- Example:
{action: "await_text", payload: "Welcome"}
Interaction
-
click: Click element
selector: CSS selector- Example:
{action: "click", selector: "button.submit"}
-
type: Text input
selector: Optional — clicks to focus firstpayload: Text to type (\t=Tab,\n=Enter)- Example:
{action: "type", selector: "#email", payload: "user@example.com"}
-
double_click: Double-click element (fires dblclick event)
selector: CSS selector- Example:
{action: "double_click", selector: ".item"}
-
right_click: Right-click element (fires contextmenu event)
selector: CSS selector- Example:
{action: "right_click", selector: ".row"}
-
select: Select dropdown option
selector: CSS selectorpayload: Option value(s)- Example:
{action: "select", selector: "select[name=state]", payload: "CA"}
-
keyboard_press: Press special keys (Tab, Enter, Escape, Arrow keys, F1-F12)
payload: Key name (string) or{"key": "Tab", "modifiers": {"shift": true, "ctrl": false, "alt": false, "meta": false}}- Example:
{action: "keyboard_press", payload: "Tab"} - Example with modifiers:
{action: "keyboard_press", payload: {"key": "Tab", "modifiers": {"shift": true}}}
Mouse Actions (CDP-Level)
These use CDP Input.dispatchMouseEvent, bypassing synthetic event restrictions.
-
hover: Move mouse over element (CSS :hover, tooltips, menus)
selector: CSS selector- Example:
{action: "hover", selector: ".menu-trigger"}
-
drag_drop: Drag element to target (native drag-and-drop via CDP)
selector: Source elementpayload: Target selector or JSON coordinates{"x":N,"y":N}- Example:
{action: "drag_drop", selector: ".card", payload: ".column-2"}
-
mouse_move: Move mouse to coordinates
payload: JSON{"x":N,"y":N}(optional:steps,fromX,fromYfor smooth movement)- Example:
{action: "mouse_move", payload: "{\"x\":100,\"y\":200}"}
-
scroll: Scroll via mouse wheel events
payload: Direction (up/down/left/right) or JSON{"deltaX":N,"deltaY":N}selector: Optional — scroll within element- Example:
{action: "scroll", payload: "down"}
File Upload
- file_upload: Set files on input[type=file] elements (can't be done via JavaScript)
selector: File input elementpayload: File path or JSON{"files":["/path/a.pdf","/path/b.jpg"]}- Example:
{action: "file_upload", selector: "#upload", payload: "/tmp/doc.pdf"}
Extraction
-
extract: Get page content
payload: Format ('markdown'|'text'|'html')selector: Optional - limit to element- Example:
{action: "extract", payload: "markdown"} - Example:
{action: "extract", payload: "text", selector: "h1"}
-
attr: Get element attribute
selector: CSS selectorpayload: Attribute name- Example:
{action: "attr", selector: "a.download", payload: "href"}
-
set_attr: Write-only attribute setter, restricted to EXACTLY two attribute names —
data-sen-nonceanddata-sen-secretitself (nothing else, not any otherdata-*/aria-*name; seeskills/browsing/lib/set-attribute.js'sALLOWED_ATTRIBUTE_NAMESconstant)selector: CSS or XPath selectorpayload:{"name": "data-sen-nonce"|"data-sen-secret", "value": "..."}(no bare-string form — needs both fields)data-sen-nonceresolves to the single first-VISIBLE match, the same wayextract/click/typeresolve a selector, and refuses if that element already carriesdata-sen-secret.data-sen-secretinstead marks EVERY element the selector matches, hidden duplicates included, and can never remove or weaken an existing mark (re-marking an already-marked element is a no-op, not a refusal) — this is the write path for the marker itself.- Unlike every read action,
set_attris NOT blocked by adata-sen-secretelement existing elsewhere on the page — it takes no caller JavaScript, and its only output isok/no element matched/refused, which acts as a limited prefix oracle (see below). Use it instead ofevalto write onto a page that already has a captured secret (e.g. stamping a credential-broker nonce onto an unmarked digit-input box next to a just-captured TOTP seed, or marking the seed's element in the first place). - Why so narrow: page JS and frameworks routinely read arbitrary
data-*/aria-*attributes and wire them to behavior (data-action,data-href,aria-controls, and more a hostile page could invent), so a prefix allowlist is not guaranteed inert. Widening past these two names is a deliberate, separate change. Itsok/no element matched/refused: target element is markedresponses differ by outcome, which is itself a limited prefix oracle over page content for a caller who varies the selector and watches which result comes back — a known, accepted limitation, not something this guards against. - Example:
{action: "set_attr", selector: "#code-input-0", payload: {"name": "data-sen-nonce", "value": "opaque-nonce"}}
-
eval: Execute JavaScript
payload: JavaScript code- Example:
{action: "eval", payload: "document.title"} - Refuses outright (no value-blind exception) while any element on the page carries
data-sen-secret, checked live at the moment of the call — seeset_attrabove for the write-only escape hatch. This is an accident guard, not a security boundary: eval runs in the same JS realm as the marked element, so it can already read the value directly, exfiltrate it viafetch()/window.name/storage, or erase the marker withremoveAttributeas its own last step — none of which this check can catch, by design. Don't mark an element and then eval on that page expecting the value to stay contained; after marking, use the credential broker for the value andset_attrfor writes.
Export
- screenshot: Capture screenshot of a specific element
payload: Filenameselector: Optional - screenshot specific element- Viewport screenshots are auto-captured after every DOM action. Use this only when you need a specific element.
- Example:
{action: "screenshot", payload: "/tmp/chart.png", selector: ".chart"}
Tab Management
-
list_tabs: List all open tabs
- Example:
{action: "list_tabs"}
- Example:
-
new_tab: Create new tab
- Example:
{action: "new_tab"}
- Example:
-
close_tab: Close the active tab
- Example:
{action: "close_tab"}
- Example:
-
switch_tab: Switch the active tab (sticky — stays until changed)
payload: Tab index (number), URL substring, or title substring- Example:
{action: "switch_tab", payload: 1}(by index) - Example:
{action: "switch_tab", payload: "example.com"}(by URL substring) - Example:
{action: "switch_tab", payload: "GitHub"}(by title substring)
Browser Mode Control
-
show_browser: Make browser window visible (headed mode)
- Example:
{action: "show_browser"} - ⚠️ WARNING: Restarts Chrome, reloads pages via GET, loses POST state
- Example:
-
hide_browser: Switch to headless mode (invisible browser)
- Example:
{action: "hide_browser"} - ⚠️ WARNING: Restarts Chrome, reloads pages via GET, loses POST state
- Example:
-
browser_mode: Check current browser mode, port, and profile
- Example:
{action: "browser_mode"} - Returns:
{"headless": true|false, "mode": "headless"|"headed", "running": true|false, "port": 9222, "profile": "name", "profileDir": "/path"}
- Example:
Profile Management
-
set_profile: Change Chrome profile (must kill Chrome first)
- Example:
{action: "set_profile", "payload": "browser-user"} - ⚠️ WARNING: Chrome must be stopped first
- Side effect: marks the profile as explicit, opting out of auto-disambiguation (see below)
- Example:
-
get_profile: Get current profile name and directory
- Example:
{action: "get_profile"} - Returns:
{"profile": "name", "profileDir": "/path"}
- Example:
Default behavior: Chrome starts in headless mode with "superpowers-chrome" profile on a dynamically allocated port (range 9222-12111). Override the port with CHROME_WS_PORT; override the profile with CHROME_WS_PROFILE.
Auto-disambiguation across parallel MCPs:
When two MCP servers start on the same host with the default profile, the first claims superpowers-chrome (port 9222) and later ones silently fall through to superpowers-chrome-2 (port 9223), superpowers-chrome-3, etc. Each MCP drives its own Chrome with its own profile dir; they don't fight over activeTab. The bridge tracks ownership via a lock file at ~/.cache/superpowers/browser-profiles/<profile>.mcp.lock; stale locks (dead PIDs) are reclaimed automatically.
To opt out of disambiguation — e.g., to intentionally share Chrome between a long-lived chrome-ws CLI session and your MCP — set the profile name explicitly:
- Env var:
CHROME_WS_PROFILE=my-profile - Or:
{action: "set_profile", payload: "my-profile"}at runtime
An explicit profile name still acquires the lock, but on conflict the bridge shares rather than disambiguates — the second process reconnects to the first's Chrome (the original reconnect-on-restart behavior).
Chrome Lifecycle (Recovery)
-
kill_chrome: Kill the Chrome process this MCP is driving
- Example:
{action: "kill_chrome"} - Releases the meta.json; next page action auto-restarts Chrome
- Example:
-
restart_chrome: kill_chrome + immediate spawn
- Example:
{action: "restart_chrome"}
- Example:
Auto-restart banner: when the bridge has to spawn a fresh Chrome (because the previous one died or was killed externally — e.g., kill -9 <pid> from the shell), the first response after the restart prepends:
[Chrome auto-restarted; URL reset to about:blank. Re-navigate to continue.]
Treat this as a signal that your prior URL / tab state is gone — re-navigate before assuming anything about the current page.
Console Logging
Capture browser console output for the active tab. Buffer is keyed by the page session's `sessionId
Files of this skill
skills/browsing/COMMANDLINE-USAGE.mdskills/browsing/EXAMPLES.mdskills/browsing/README.mdskills/browsing/chrome-wsskills/browsing/chrome-ws-lib.jsskills/browsing/host-override.jsskills/browsing/lib/browser-bridge.jsskills/browsing/lib/browser-session.jsskills/browsing/lib/capture.jsskills/browsing/lib/cdp-router.jsskills/browsing/lib/cdp-utils.jsskills/browsing/lib/chrome-launcher-helpers.jsskills/browsing/lib/chrome-process.jsskills/browsing/lib/console-logging.jsskills/browsing/lib/cookies.jsskills/browsing/lib/credential-guard.jsskills/browsing/lib/dialogs-render.jsskills/browsing/lib/dialogs-router.jsskills/browsing/lib/dialogs.jsskills/browsing/lib/element-selector.jsskills/browsing/lib/evaluation.jsskills/browsing/lib/extraction.jsskills/browsing/lib/file-upload.jsskills/browsing/lib/html-diff.jsskills/browsing/lib/key-definitions.jsskills/browsing/lib/keyboard-input.jsskills/browsing/lib/mouse.jsskills/browsing/lib/navigation.jsskills/browsing/lib/page-scripts/dom-summary.jsskills/browsing/lib/page-scripts/html-with-scrub.jsskills/browsing/lib/page-scripts/markdown.jsskills/browsing/lib/page-scripts/permission-shim.jsskills/browsing/lib/page-scripts/rendered-text.jsskills/browsing/lib/page-session.jsskills/browsing/lib/profile-lock.jsskills/browsing/lib/screenshot.jsskills/browsing/lib/secret-marker.jsskills/browsing/lib/select-option.jsskills/browsing/lib/session-state.jsskills/browsing/lib/set-attribute.jsskills/browsing/lib/tabs.jsskills/browsing/lib/viewport.jsskills/browsing/lib/websocket-client.jsskills/browsing/package.jsonskills/browsing/test-chrome-args.jsskills/browsing/test-cookies.jsskills/browsing/test-e2e.shskills/browsing/test-extract.shskills/browsing/test-interact.shskills/browsing/test-navigate.sh
Only the first files are listed.