feat(browser): replace the preview proxy with a real browser panel and an agent web tool (#2883)

The preview panel worked by proxying a dev server through OpenChamber's own
origin and rewriting the HTML that came back. Anything the rewriter did not
anticipate broke, and pages that refuse to be embedded never loaded at all.
This deletes the proxy (-1604 lines and its tests) and merges the preview and
browser panels into one surface backed by a real Chromium view.

What the panel is now

- A `<webview>` in its own session partition: logins and cookies persist, hot
  reload works because nothing is rewritten, DevTools are one click away.
- Annotation: pick one element, drag a region, or draw freehand, write a note,
  and it reaches chat with a screenshot of the visible page with the marks on it.
- Toolbar: hard reload, page zoom, device sizes, a light/dark switch that
  applies to the page rather than the app, and cookie/cache clearing scoped to
  the panel alone.
- Several pages at once, each tab showing the page's own favicon, and an address
  bar that suggests pages already visited in this project.
- Dev servers are listed from what is actually listening on the machine, checked
  against what a project announced, so a server is offered no matter how it was
  started. One that is still starting is waited for instead of failing.

Remote dev servers

The desktop app binds a local port and pipes raw bytes to the OpenChamber host
over the existing authenticated connection, so the page keeps its own origin at
the root of its own host. The reachable set is exactly what discovery reports
and is re-checked per connection, so an authenticated client cannot dial
arbitrary local services on the host. Links and redirects to another loopback
port stay on the machine that served the page. A tunnel that cannot be opened is
reported; it is never replaced by the plain loopback URL, which would answer
from the user's own machine under a remote address.

Agent control

Browser actions are a separate `openchamber_web` tool: open, snapshot, click,
type, scroll, inspect computed styles, resize between mobile/tablet/desktop, and
capture a screenshot into `.openchamber/screenshots/` in the project. The
existing `openchamber` tool keeps sessions, worktrees and scheduled tasks. Each
has its own setting in the new Settings -> General -> OpenChamber Tools section,
and the plugin is not injected at all when both are off.

Capability belongs to the connected client, not to configuration: a client
declares on its event stream that it can drive a page, which only a Chromium
host does. Exactly one client performs each request — it claims the request
before acting, and the first claim wins — because deciding by whose result
arrives first would be too late for a click that already happened. No client
listening is answered immediately with an explanation rather than a timeout.

Runtime boundaries

Web tabs get a plain iframe that can display a page but not inspect one. The
VS Code extension no longer offers the surface at all, since nothing that makes
the panel worth having works there. Mobile is unaffected.

Native boundary

Camera, microphone, location and device-picker requests from panel pages are
denied — Electron grants them by default when no handler is set, and the panel
loads whatever address the user types. Page capture, appearance emulation and
storage clearing verify that their target belongs to the panel's own session
instead of trusting a web-contents id from the renderer.

Persisted state

Stored `preview` tabs migrate to `browser` (v13 -> v14). Context panel tab
limits are now per surface, so filling one surface no longer evicts another's
tabs. Address history is stored per project and per runtime.

Documentation

`preview.mdx` and `desktop-browser.mdx` rewritten across all locales, the agent
tool settings path corrected, new `DOCUMENTATION.md` for the browser-control
broker and the dev tunnel, and the `ui-api-decoupling` skill updated where it
still described the deleted proxy.
This commit is contained in:
Bohdan Triapitsyn
2026-08-13 22:44:13 +03:00
committed by GitHub
parent 50613bb170
commit a5aa32446d
151 changed files with 10431 additions and 5587 deletions
@@ -53,3 +53,12 @@ other.
directory and does not erase other session results.
- Destructive session/worktree deletion and project-path registration are not
part of the action contract.
- `browser.capture` writes its image on the server, into
`.openchamber/screenshots/` under the scoped project directory, and returns
the project-relative path rather than the image bytes. The client that took
the picture may be on a different machine than the repository, and a path is
what an answer, a commit, or a review can use; base64 in a tool result cannot
be any of those. The agent's label is reduced to a filename fragment, never
used as a path. The result also states how to present the image, because chat
renders the image paths written in a finished answer below that message —
saving the file is not what shows it to anyone.
@@ -1,3 +1,12 @@
/**
* Two capabilities, two tools.
*
* Controlling sessions and driving a page are different intents, and a single
* tool description covering both is vaguer than either — which is how a model
* ends up calling the wrong one. Separate tools also mean turning one off
* removes it entirely, parameters included, rather than leaving its inputs
* visible in a shared schema.
*/
export const OPENCHAMBER_CONTROL_ACTION_DEFINITIONS = Object.freeze([
{ action: 'projects.list', title: 'List configured projects', description: 'List configured projects; no parameters' },
{ action: 'models.list', title: 'Show model preferences', description: 'Show default, favorite, and recent model preferences; no parameters' },
@@ -15,7 +24,7 @@ export const OPENCHAMBER_CONTROL_ACTION_DEFINITIONS = Object.freeze([
{ action: 'schedule.toggle', title: 'Enable or disable a scheduled task', description: 'Enable or disable taskId; requires the disabled boolean' },
]);
export const OPENCHAMBER_CONTROL_ACTIONS = Object.freeze(
const OPENCHAMBER_CONTROL_ACTIONS = Object.freeze(
OPENCHAMBER_CONTROL_ACTION_DEFINITIONS.map(({ action }) => action),
);
@@ -26,3 +35,26 @@ export const OPENCHAMBER_AGENT_TOOL_ACTION_DEFINITIONS = Object.freeze(
export const OPENCHAMBER_AGENT_TOOL_ACTIONS = Object.freeze(
OPENCHAMBER_AGENT_TOOL_ACTION_DEFINITIONS.map(({ action }) => action),
);
export const OPENCHAMBER_WEB_ACTION_DEFINITIONS = Object.freeze([
{ action: 'browser.open', title: 'Open a page in the browser panel', description: 'Open url in the in-app browser panel; use it to look at the running app. Set viewport to mobile, tablet or desktop to lay the page out at that size' },
{ action: 'browser.snapshot', title: 'Read the open page', description: 'Read the open page: url, title, visible text, and interactive elements with the selectors the other browser actions accept. Pass selector to read only that part of a long page. Reports any errors the page logged' },
{ action: 'browser.click', title: 'Click on the open page', description: 'Click an element; give selector, or text to match a link or button by its visible label' },
{ action: 'browser.type', title: 'Type into the open page', description: 'Type value into the field matched by selector; set submit to press Enter afterwards' },
{ action: 'browser.scroll', title: 'Scroll the open page', description: 'Scroll the page; direction is up, down, top, or bottom, or pass selector to bring one element into view' },
{ action: 'browser.back', title: 'Go back in the browser panel', description: 'Return to the previous page in this tab; no parameters' },
{ action: 'browser.forward', title: 'Go forward in the browser panel', description: 'Move forward again in this tab; no parameters' },
{ action: 'browser.inspect', title: 'Read how an element renders', description: 'Read the computed styles of the element matched by selector — colours, fonts, spacing, borders — as the page actually renders them' },
{ action: 'browser.capture', title: 'Save a screenshot of the page', description: 'Save what is currently visible in the browser panel as an image file in the project and return its path, so a change can be shown rather than described. Pass label to name it (for example before-fix); the result reports the page, layout and path to reference in your answer' },
{ action: 'browser.resize', title: 'Change the page viewport', description: 'Lay the open page out at a different size; viewport is mobile, tablet, desktop, or fill to use the whole panel' },
]);
export const OPENCHAMBER_WEB_ACTIONS = Object.freeze(
OPENCHAMBER_WEB_ACTION_DEFINITIONS.map(({ action }) => action),
);
/** Everything the callback route will dispatch, whichever tool asked. */
export const OPENCHAMBER_ALL_ACTIONS = Object.freeze([
...OPENCHAMBER_CONTROL_ACTIONS,
...OPENCHAMBER_WEB_ACTIONS,
]);
@@ -0,0 +1,83 @@
/**
* Where an agent's page screenshots land.
*
* The image is written on the server, next to the code it is evidence for,
* because that is the machine holding the repository — the client that took the
* picture may be somewhere else entirely. A file in the project is also the
* only form of this that survives past the chat: it can be referenced from an
* answer, committed, or attached to a review.
*
* A screenshot nobody can place is not evidence, so the name carries the label
* the agent chose and the moment it was taken, and the caller is handed back
* the page and layout it shows.
*/
import path from 'node:path';
import fsPromises from 'node:fs/promises';
/** Project-relative home for agent screenshots. */
export const SCREENSHOT_DIRECTORY = path.join('.openchamber', 'screenshots');
const MAX_LABEL_LENGTH = 48;
/**
* Turns a label into a filename fragment.
*
* Everything outside a small safe set is dropped rather than escaped: this
* value reaches the filesystem, and a label is a name, never a path. `..`, a
* separator, or a leading dot cannot survive this.
*/
export const screenshotSlug = (label) => {
const slug = String(label ?? '')
.toLowerCase()
.replace(/[^a-z0-9]+/g, '-')
.replace(/^-+|-+$/g, '')
.slice(0, MAX_LABEL_LENGTH)
.replace(/-+$/g, '');
return slug || 'page';
};
/** File-safe timestamp: sorts chronologically and reads as a date. */
const screenshotStamp = (date) => date.toISOString().replace(/[:.]/g, '-').replace('Z', '');
const EXTENSIONS = new Map([
['image/jpeg', '.jpg'],
['image/png', '.png'],
['image/webp', '.webp'],
]);
/**
* Writes one capture into the project and reports where it went.
*
* Returns both the project-relative path — what belongs in an answer or a
* commit — and the absolute one, so a caller that needs the file itself does
* not have to rebuild it.
*/
export const writeScreenshot = async ({
directory,
base64,
mime = 'image/jpeg',
label,
now = new Date(),
fs = fsPromises,
}) => {
if (typeof directory !== 'string' || directory.trim().length === 0) {
throw new Error('A project directory is required to save a screenshot');
}
if (typeof base64 !== 'string' || base64.length === 0) {
throw new Error('The browser returned no image');
}
const extension = EXTENSIONS.get(mime) || '.jpg';
const relativePath = path.join(
SCREENSHOT_DIRECTORY,
`${screenshotSlug(label)}-${screenshotStamp(now)}${extension}`,
);
const absolutePath = path.join(directory, relativePath);
await fs.mkdir(path.dirname(absolutePath), { recursive: true });
await fs.writeFile(absolutePath, Buffer.from(base64, 'base64'));
// Posix separators in the reported path: it is written into Markdown and
// commit messages, where a Windows separator is an escape character.
return { path: relativePath.split(path.sep).join('/'), absolutePath };
};
@@ -0,0 +1,84 @@
import { describe, expect, it } from 'vitest';
import path from 'node:path';
import { SCREENSHOT_DIRECTORY, screenshotSlug, writeScreenshot } from './screenshots.js';
const createFs = () => {
const written = new Map();
const made = [];
return {
written,
made,
mkdir: async (target) => { made.push(target); },
writeFile: async (target, data) => { written.set(target, data); },
};
};
describe('screenshot labels', () => {
it('keeps a readable name', () => {
expect(screenshotSlug('Before fix')).toBe('before-fix');
});
it('never lets a label become a path', () => {
expect(screenshotSlug('../../etc/passwd')).toBe('etc-passwd');
expect(screenshotSlug('/absolute')).toBe('absolute');
expect(screenshotSlug('..')).toBe('page');
expect(screenshotSlug('.hidden')).toBe('hidden');
});
it('falls back to a name rather than an empty one', () => {
expect(screenshotSlug('')).toBe('page');
expect(screenshotSlug('!!!')).toBe('page');
expect(screenshotSlug(undefined)).toBe('page');
});
});
describe('writing a screenshot', () => {
const base64 = Buffer.from('image-bytes').toString('base64');
it('writes into the project and reports a portable relative path', async () => {
const fs = createFs();
const result = await writeScreenshot({
directory: '/work/project',
base64,
mime: 'image/jpeg',
label: 'After fix',
now: new Date('2026-08-13T09:37:00.000Z'),
fs,
});
expect(result.path).toBe('.openchamber/screenshots/after-fix-2026-08-13T09-37-00-000.jpg');
expect(result.path.includes('\\')).toBe(false);
expect(result.absolutePath).toBe(path.join('/work/project', SCREENSHOT_DIRECTORY, 'after-fix-2026-08-13T09-37-00-000.jpg'));
expect(fs.written.get(result.absolutePath).toString()).toBe('image-bytes');
expect(fs.made[0]).toBe(path.join('/work/project', SCREENSHOT_DIRECTORY));
});
it('names the file after the image it actually holds', async () => {
const fs = createFs();
const result = await writeScreenshot({ directory: '/work/project', base64, mime: 'image/png', fs });
expect(result.path.endsWith('.png')).toBe(true);
});
it('refuses to write without a project directory', async () => {
let failed = false;
try {
await writeScreenshot({ directory: '', base64, fs: createFs() });
} catch {
failed = true;
}
expect(failed).toBe(true);
});
it('reports an empty capture instead of writing a zero-byte file', async () => {
const fs = createFs();
let failed = false;
try {
await writeScreenshot({ directory: '/work/project', base64: '', fs });
} catch {
failed = true;
}
expect(failed).toBe(true);
expect(fs.written.size).toBe(0);
});
});
@@ -1,12 +1,14 @@
import path from 'node:path';
import { createOpencodeClient } from '@opencode-ai/sdk/v2';
import { OpenChamberControlError, asControlError } from './error.js';
import { OPENCHAMBER_CONTROL_ACTIONS } from './actions.js';
import { OPENCHAMBER_ALL_ACTIONS } from './actions.js';
import { writeScreenshot } from './screenshots.js';
const DEFAULT_WAIT_TIMEOUT_SECONDS = 600;
const MAX_WAIT_TIMEOUT_SECONDS = 86_400;
const WAIT_POLL_INTERVAL_MS = 500;
const CONTROL_ACTIONS = new Set(OPENCHAMBER_CONTROL_ACTIONS);
// One service, both capabilities: which tool asked is the caller's concern.
const CONTROL_ACTIONS = new Set(OPENCHAMBER_ALL_ACTIONS);
const SCHEDULE_TASK_ID_ACTIONS = new Set([
'schedule.run',
'schedule.delete',
@@ -141,6 +143,7 @@ export const createOpenChamberControlService = (dependencies) => {
waitForOpenCodeReady,
sessionService,
scheduledTaskService,
browserControl = null,
createClient = createOpencodeClient,
sleep = (duration) => new Promise((resolve) => setTimeout(resolve, duration)),
now = Date.now,
@@ -320,11 +323,147 @@ export const createOpenChamberControlService = (dependencies) => {
return publicResult;
};
/**
* Validates browser inputs here rather than in the renderer: an invalid call
* should come back as a usage error the agent can correct, without waking a
* client or waiting for a round trip.
*/
const browserAction = async (action, input, signal, contextDirectory) => {
const parameters = {};
const readViewport = (required) => {
const viewport = asNonEmptyString(input.viewport);
if (!viewport) {
if (required) throw new OpenChamberControlError('viewport is required for browser.resize', 400);
return;
}
if (!['mobile', 'tablet', 'desktop', 'fill'].includes(viewport)) {
throw new OpenChamberControlError('viewport must be mobile, tablet, desktop, or fill', 400);
}
parameters.viewport = viewport;
};
if (action === 'browser.resize') readViewport(true);
if (action === 'browser.capture') {
const label = asNonEmptyString(input.label);
if (label) parameters.label = label;
}
if (action === 'browser.open') {
readViewport(false);
const url = asNonEmptyString(input.url);
if (!url) throw new OpenChamberControlError('url is required for browser.open', 400);
let parsed;
try {
parsed = new URL(url);
} catch {
throw new OpenChamberControlError('url must be an absolute http(s) URL', 400);
}
if (parsed.protocol !== 'http:' && parsed.protocol !== 'https:') {
throw new OpenChamberControlError('url must use http or https', 400);
}
parameters.url = parsed.toString();
}
if (action === 'browser.click') {
const selector = asNonEmptyString(input.selector);
const text = asNonEmptyString(input.text);
if (!selector && !text) {
throw new OpenChamberControlError('browser.click requires selector or text', 400);
}
if (selector) parameters.selector = selector;
if (text) parameters.text = text;
}
if (action === 'browser.snapshot') {
const selector = asNonEmptyString(input.selector);
if (selector) parameters.selector = selector;
}
if (action === 'browser.inspect') {
const selector = asNonEmptyString(input.selector);
if (!selector) throw new OpenChamberControlError('selector is required for browser.inspect', 400);
parameters.selector = selector;
}
if (action === 'browser.type') {
const selector = asNonEmptyString(input.selector);
if (!selector) throw new OpenChamberControlError('selector is required for browser.type', 400);
if (typeof input.value !== 'string') {
throw new OpenChamberControlError('value is required for browser.type', 400);
}
parameters.selector = selector;
parameters.value = input.value;
parameters.submit = input.submit === true;
}
if (action === 'browser.scroll') {
const selector = asNonEmptyString(input.selector);
const direction = asNonEmptyString(input.direction);
if (!selector && !direction) {
throw new OpenChamberControlError('browser.scroll requires direction or selector', 400);
}
if (direction && !['up', 'down', 'top', 'bottom'].includes(direction)) {
throw new OpenChamberControlError('direction must be up, down, top, or bottom', 400);
}
if (selector) parameters.selector = selector;
if (direction) parameters.direction = direction;
}
// Opening a page waits for the navigation to settle, so its budget has to
// exceed the client's own wait; sharing one timeout with the quick actions
// made a slow page indistinguishable from an unreachable browser.
const timeoutMs = action === 'browser.open' ? 45_000 : 20_000;
const result = await browserControl.request(action, parameters, { signal, timeoutMs });
// The image is written here rather than in the renderer: the file belongs
// beside the code it documents, and the client that took it may be on a
// different machine than the repository.
if (action === 'browser.capture') {
const directory = asNonEmptyString(input.directory) || asNonEmptyString(contextDirectory);
if (!directory) {
throw new OpenChamberControlError('directory is required to save a screenshot', 400);
}
const capture = result && typeof result === 'object' ? result : {};
const saved = await writeScreenshot({
directory,
base64: capture.base64,
mime: capture.mime,
label: input.label,
});
// The base64 never goes back to the caller: it is large, and the path is
// what an answer, a commit, or a review can actually use.
return {
path: saved.path,
// Saving the file is only half of showing it. Chat collects the image
// paths written in a finished answer and renders them below it, so the
// agent is told the one thing it cannot infer: that writing the path is
// what puts the picture in front of the user.
hint: `Write ![](${saved.path}) in your reply to show this image to the user; it is rendered under your message.`,
url: capture.url ?? null,
title: capture.title ?? null,
viewport: capture.viewport ?? null,
width: capture.width ?? null,
height: capture.height ?? null,
};
}
return result;
};
const execute = async (action, input = {}, contextDirectory, options = {}) => {
try {
if (!CONTROL_ACTIONS.has(action)) {
throw new OpenChamberControlError(`Unsupported OpenChamber action: ${action || 'missing'}`, 400);
}
if (action.startsWith('browser.')) {
if (!browserControl) {
throw new OpenChamberControlError('The in-app browser is not available on this server', 503);
}
return browserAction(action, input, options.signal, contextDirectory);
}
if (action === 'projects.list') return { projects: await projects() };
if (action === 'models.list') return models();
if (action === 'schedule.status') return scheduledTaskService.status();
@@ -1,5 +1,9 @@
import { describe, expect, it, vi } from 'vitest';
import os from 'node:os';
import path from 'node:path';
import fs from 'node:fs/promises';
import { createOpenChamberControlService } from './service.js';
const createService = (overrides = {}) => {
@@ -261,3 +265,54 @@ describe('OpenChamber control service', () => {
await expect(service.execute('session.delete')).rejects.toThrow('Unsupported OpenChamber action');
});
});
describe('browser capture', () => {
const pixel = 'iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8z8BQDwAEhQGAhKmMIQAAAABJRU5ErkJggg==';
const createBrowserService = async (capture) => {
const directory = await fs.mkdtemp(path.join(os.tmpdir(), 'oc-capture-'));
const request = vi.fn(async () => capture);
const { service } = createService({ browserControl: { request } });
return { service, directory, request };
};
it('saves the image beside the code and hands back a path the answer can use', async () => {
const { service, directory } = await createBrowserService({
base64: pixel,
mime: 'image/png',
url: 'http://localhost:3000/',
title: 'App',
viewport: { mode: 'mobile', width: 390, height: 844 },
width: 390,
height: 844,
});
const result = await service.execute('browser.capture', { label: 'After fix' }, directory);
expect(result.path.startsWith('.openchamber/screenshots/after-fix-')).toBe(true);
expect(result.path.endsWith('.png')).toBe(true);
expect(result.url).toBe('http://localhost:3000/');
expect(result.viewport).toEqual({ mode: 'mobile', width: 390, height: 844 });
// The bytes stay on disk; a tool result is not a place to carry an image.
expect('base64' in result).toBe(false);
const written = await fs.readFile(path.join(directory, result.path));
expect(written.length > 0).toBe(true);
});
it('tells the agent how to actually show the image', async () => {
const { service, directory } = await createBrowserService({ base64: pixel, mime: 'image/png' });
const result = await service.execute('browser.capture', {}, directory);
expect(result.hint).toContain(`![](${result.path})`);
});
it('refuses to capture with no project to save into', async () => {
const { service } = await createBrowserService({ base64: pixel, mime: 'image/png' });
await expect(service.execute('browser.capture', {})).rejects.toThrow(/directory is required/);
});
it('passes a label through to the browser and leaves other actions untouched', async () => {
const { service, directory, request } = await createBrowserService({ base64: pixel, mime: 'image/png' });
await service.execute('browser.capture', { label: 'before' }, directory);
expect(request).toHaveBeenCalledWith('browser.capture', { label: 'before' }, expect.anything());
});
});