Surfaces / browser automation
Pull the agent's work artifacts out of the chat stream into a right-pane canvas, and let it drive a real browser.
What is a surface
Surfaces pull the agent's output and work artifacts out of the chat stream and onto a stable canvas: the conversation stays lean, and artifacts get a dedicated place.
Examples:
- The agent launches a browser for automation — you see the real browser view
- The agent writes a Markdown report — rendered in the right pane, editable with a click
- The agent pulls tabular data — rendered in a Sheet surface the agent can keep reading and writing
Surface types
open_surface supports these types:
| Type | Purpose | Source |
|---|---|---|
| doc | Markdown document | Input or file |
| diagram | Mermaid / DOT diagram | Input |
| image | Single image | File or URL |
| html | Rich HTML | Input |
| svg | Vector graphics | Input |
| web | Browser (interactive) | URL |
| PDF viewer | File | |
| code | Code file preview | File |
| diff | File changes | Input |
| plan | Task plan | Live |
| sheet | Spreadsheet | Input or file |
| docx | Word document | Input or file |
| pptx | Slides | Input or file |
| todo | Todo list | Live |
| services | Running services | Live |
Sources
source comes in three kinds:
- File — linked to a project file; the surface updates live as the file changes
- Input — content the agent generates on the spot
- URL — a remote resource
Working with surfaces
| Tool | What it does |
|---|---|
open_surface | Open a new surface (in a new tab) |
update_surface | Incrementally update surface content |
edit_plan / update_todos | Update the corresponding content types |
browser_list_surfaces | List all currently open surfaces |
There is no limit on right-pane tabs and none are closed automatically — they stay until you close them. Draw charts with diagram (Mermaid); command output lives in the services panel.
Browser automation
The agent can drive a real browser directly — navigate, click, fill forms, take screenshots, run complete flows. On desktop it uses a web page in the right pane; in the CLI, a separate Chrome window.
Common tools include:
| Function | Example tools |
|---|---|
| Navigation | browser_navigate / browser_back / browser_forward / browser_reload |
| Screenshots | browser_screenshot |
| Page content queries | browser_query / browser_get_text / browser_get_aria_tree, etc. |
| Interaction | browser_click / browser_type / browser_press_key / browser_hover / browser_fill_form, etc. |
| Waiting for loads | browser_wait_for / browser_wait_for_navigation |
| Network control | Request interception / response mocking |
| Assertions | browser_eval / browser_expect |
| Cookies & storage | Get / set / clear |
| File uploads | browser_set_input_files |
| Multiple tabs | browser_new_tab / browser_close_tab |
The agent sees the actual page render and uses it to decide where to click next.
Tip
To run several automation flows at once, have the agent open separate tabs — actions on a single page execute one step at a time.
Availability
| Platform | Browser automation | Right-pane canvas |
|---|---|---|
| Desktop | ✓ Built-in browser in the right pane, or an external Chrome | ✓ |
| CLI | ✓ Launches a separate Chrome window you can watch | ✗ |
The CLI needs Chrome installed on your machine.

