> ## Documentation Index
> Fetch the complete documentation index at: https://docs.boat.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Computer use

> Every harness on a sandbox can see and use the Linux desktop of the sandbox through the built-in computer tools.

Every sandbox runs a Linux desktop. Every harness on it can see and use that desktop with ordinary tools. You do not install or register anything:

```bash theme={null}
boat prompt "Open example.com in the browser and tell me the page title"
boat prompt --provider codex "Log into the admin app in Chrome and screenshot the dashboard"
```

The tools come from the [Cua Driver](https://github.com/trycua/cua). It is preinstalled on every sandbox. It is registered as an MCP server named `computer` for all seven harnesses:

```mermaid theme={null}
flowchart LR
  P["boat prompt"] --> H["Harness process<br/>claude / codex / pi / opencode / prime / kimi / vibe"]
  H -- "stdio MCP<br/>server name: computer" --> M["cua-driver mcp"]
  M -- "unix socket" --> D["cua-driver daemon<br/>systemd user service"]
  D --> X["The Sandbox desktop<br/>Xorg :0, 1920x1080, Chrome"]
  X -.-> S["boat desktop<br/>watch it live"]
```

The daemon runs as the desktop user. So the tools act on the same screen that `boat desktop` streams. Open a stream in one window and watch the agent work in it.

## What the agent gets

| Group | Tools |
| - | - |
| Look | `get_window_state` (screenshot plus the accessibility tree in one call), `get_desktop_state`, `get_screen_size`, `list_windows`, `list_apps` |
| Point and type | `click`, `double_click`, `right_click`, `drag`, `scroll`, `move_cursor`, `type_text`, `press_key`, `hotkey`, `set_value` |
| Apps and windows | `launch_app`, `kill_app`, `bring_to_front`, `set_window_frame`, `invoke_menu` |
| Browser | `browser_navigate`, `browser_click`, `browser_type`, `get_browser_state`, `browser_download` |
| Clipboard | `clipboard_read`, `clipboard_write` |

By default, the tools send clicks and keystrokes in the background. They go to the target window, but they do not raise it or move the real pointer. So the agent can use several windows without a fight for focus.

Some widgets accept input only when they have focus. In that case the tool says so. The agent can then try again with `delivery_mode: "foreground"`. This activates the window, does the action, and gives focus back.

In `boat events`, each harness shows the tool names with its own prefix:

| Harness | Example tool name |
| - | - |
| Claude Code, Kimi Code | `mcp__computer__click` |
| Codex | `mcp:computer/click` |

## Where it is registered

The registration is an ordinary entry in the config file of each harness. Boat writes it at sandbox start, and merges it with what you put there yourself:

| Harness | File | Entry |
| - | - | - |
| Claude Code | `~/.claude.json` | `mcpServers.computer` |
| Codex | `~/.codex/config.toml` | `[mcp_servers.computer]`, in a marked block |
| pi | `~/.pi/agent/mcp.json` | `mcpServers.computer` |
| OpenCode | `~/.config/opencode/opencode.json` | `mcp.computer` |
| Prime Agent | `~/.prime/agent/settings.json` | `mcpServers.computer` |
| Kimi Code | `~/.kimi-code/mcp.json` | `mcpServers.computer` |
| Mistral Vibe | none: the agent server passes it when it opens each session | `computer` |

* Boat does not touch your own MCP servers in those files.
* If you remove `computer` from a file, that harness loses the tools.
* Boat writes the entry again at the next sandbox start. To remove the tools for good, delete the daemon from the sandbox.

OpenCode has one tool, `browser_prepare`, turned off in `opencode.json` `tools`. The reason: OpenCode sends MCP tool schemas to the model provider without change, and the Anthropic API rejects the schema of that tool. Without this, every OpenCode turn on an Anthropic model fails.

Prime Agent calls the server from inside its Python tool, as `await mcp.call_tool("computer", "click", {...})`.

## Practical notes

* **One screen, shared.** All conversations on a sandbox use the same desktop. Two agents that click at the same time interfere with each other. Give each parallel GUI task its own sandbox, or run the tasks one after the other.
* **It survives stop, resume and fork.** The daemon is a systemd user service on the sandbox. It starts again automatically. The socket is on `/run`, so a resumed sandbox never gets an old socket.
* **Logins persist.** A Chrome profile that the agent signed into is on the disk of the sandbox. It comes back on resume and goes with a fork.
* **Screenshots are large.** A full-desktop screenshot is a real image in the context of the model. When a task runs in a loop, ask for one window (`get_window_state` with a window id).
* **Watch it, and record it.** Run `boat desktop` to open the stream and watch the agent work. Run `ascii-record-desktop` to record the run to an MP4. Read [Desktop Streaming](/desktop-streaming).

## Check the daemon

If a prompt says that it has no computer tools, check the daemon:

```bash theme={null}
boat exec 'systemctl --user status cua-driver.service --no-pager'
boat exec '/opt/ascii/cua-driver/cua-driver status --socket /run/ascii-cua/driver.sock'
boat exec 'systemctl --user restart cua-driver.service'
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.