Skip to content
Agent

Surfaces / browser automation

Pull the agent's work artifacts out of the chat stream into a right-pane canvas, and let it drive a real browser.

What is a surface

Surfaces pull the agent's output and work artifacts out of the chat stream and onto a stable canvas: the conversation stays lean, and artifacts get a dedicated place.

Examples:

  • The agent launches a browser for automation — you see the real browser view
  • The agent writes a Markdown report — rendered in the right pane, editable with a click
  • The agent pulls tabular data — rendered in a Sheet surface the agent can keep reading and writing

Surface types

open_surface supports these types:

TypePurposeSource
docMarkdown documentInput or file
diagramMermaid / DOT diagramInput
imageSingle imageFile or URL
htmlRich HTMLInput
svgVector graphicsInput
webBrowser (interactive)URL
pdfPDF viewerFile
codeCode file previewFile
diffFile changesInput
planTask planLive
sheetSpreadsheetInput or file
docxWord documentInput or file
pptxSlidesInput or file
todoTodo listLive
servicesRunning servicesLive

Sources

source comes in three kinds:

  • File — linked to a project file; the surface updates live as the file changes
  • Input — content the agent generates on the spot
  • URL — a remote resource

Working with surfaces

ToolWhat it does
open_surfaceOpen a new surface (in a new tab)
update_surfaceIncrementally update surface content
edit_plan / update_todosUpdate the corresponding content types
browser_list_surfacesList all currently open surfaces

There is no limit on right-pane tabs and none are closed automatically — they stay until you close them. Draw charts with diagram (Mermaid); command output lives in the services panel.

Browser automation

The agent can drive a real browser directly — navigate, click, fill forms, take screenshots, run complete flows. On desktop it uses a web page in the right pane; in the CLI, a separate Chrome window.

Common tools include:

FunctionExample tools
Navigationbrowser_navigate / browser_back / browser_forward / browser_reload
Screenshotsbrowser_screenshot
Page content queriesbrowser_query / browser_get_text / browser_get_aria_tree, etc.
Interactionbrowser_click / browser_type / browser_press_key / browser_hover / browser_fill_form, etc.
Waiting for loadsbrowser_wait_for / browser_wait_for_navigation
Network controlRequest interception / response mocking
Assertionsbrowser_eval / browser_expect
Cookies & storageGet / set / clear
File uploadsbrowser_set_input_files
Multiple tabsbrowser_new_tab / browser_close_tab

The agent sees the actual page render and uses it to decide where to click next.

Tip

To run several automation flows at once, have the agent open separate tabs — actions on a single page execute one step at a time.

Availability

PlatformBrowser automationRight-pane canvas
Desktop✓ Built-in browser in the right pane, or an external Chrome✓
CLI✓ Launches a separate Chrome window you can watch✗

The CLI needs Chrome installed on your machine.