computer for all seven harnesses:
The daemon runs as the desktop user. So the tools act on the same screen that boat desktop streams. Open a stream in one window and watch the agent work in it.
What the agent gets
By default, the tools send clicks and keystrokes in the background. They go to the target window, but they do not raise it or move the real pointer. So the agent can use several windows without a fight for focus.
Some widgets accept input only when they have focus. In that case the tool says so. The agent can then try again with
delivery_mode: "foreground". This activates the window, does the action, and gives focus back.
In boat events, each harness shows the tool names with its own prefix:
Where it is registered
The registration is an ordinary entry in the config file of each harness. Boat writes it at sandbox start, and merges it with what you put there yourself:- Boat does not touch your own MCP servers in those files.
- If you remove
computerfrom a file, that harness loses the tools. - Boat writes the entry again at the next sandbox start. To remove the tools for good, delete the daemon from the sandbox.
browser_prepare, turned off in opencode.json tools. The reason: OpenCode sends MCP tool schemas to the model provider without change, and the Anthropic API rejects the schema of that tool. Without this, every OpenCode turn on an Anthropic model fails.
Prime Agent calls the server from inside its Python tool, as await mcp.call_tool("computer", "click", {...}).
Practical notes
- One screen, shared. All conversations on a sandbox use the same desktop. Two agents that click at the same time interfere with each other. Give each parallel GUI task its own sandbox, or run the tasks one after the other.
- It survives stop, resume and fork. The daemon is a systemd user service on the sandbox. It starts again automatically. The socket is on
/run, so a resumed sandbox never gets an old socket. - Logins persist. A Chrome profile that the agent signed into is on the disk of the sandbox. It comes back on resume and goes with a fork.
- Screenshots are large. A full-desktop screenshot is a real image in the context of the model. When a task runs in a loop, ask for one window (
get_window_statewith a window id). - Watch it, and record it. Run
boat desktopto open the stream and watch the agent work. Runascii-record-desktopto record the run to an MP4. Read Desktop Streaming.