Use the Task part's state.metadata.sessionId as the only live identity contract for child sessions. Remove timestamp, status, and ordering heuristics that could misassociate concurrent subagents or repeatedly scan directory sessions while parts stream.
Extract Task metadata parsing and child-summary projection into a focused model with cached projections for unchanged message records. Preserve output and part-level metadata parsing only for legacy persisted records, and keep standalone Task rows visible when sorted activity groups are rendered.
Validated with focused Task and turn projection tests, UI type-check, lint, and dead-code analysis.
* feat(chat): prompt navigator list preview with prompt filtering
The hover preview is now an interactive scrolling mini-list of prompts:
rows render as bordered two-line cards, the highlighted row stays inside
a center dead zone and the list glides only near the window edges, wheel
steps the highlight, and the panel stays open when the pointer moves into
it so a click can be corrected inside the list.
Rail entries are filtered to real prompts: previews are built from
normalized user display parts (synthetic context stripped), fully
synthetic user messages are excluded, and shell-mode messages show their
extracted command via the shared shell bridge helpers.
* fix(chat): render shell command status transitions
The injected /shell text part carries live state in shellAction, which
the render-relevant part comparator ignored — a running→completed update
reached the store but never re-rendered the message row until the next
send. Compare shellAction command/output/status for text parts.
* fix(sync): stream shell bridge part updates while running
Streaming suspension keeps part updates out of the static message records
while an assistant message streams, relying on the live streaming-tail
path to render it. Shell-mode bridge messages are hidden from the
timeline and rendered inside the user row, so they have no live path —
suspension froze their output chunks and left the card without a Show
output action until the run finished. Exempt shell bridges (single bash
tool part parented to a synthetic shell-marker user message) from
suspension; their updates arrive at command-output pace, not delta pace.
* feat(chat): syntax-highlight shell command card
Render the shell-mode command and its output through the shared
WorkerHighlightedCode (Shiki) with bash grammar, matching the bash tool
part presentation, instead of plain pre blocks.
* fix(chat): rework prompt navigator rail into sliding tape with hover preview
- Pin the active indicator to the target during programmatic scrolls so the
scroll spy's intermediate reports don't drag it backwards mid-animation
- Replace the visibility-ratio active-turn picker with a stable reading-line
rule (last turn whose top is above the line), dropping IntersectionObserver
- Replace the list panel with a Codex-style gutter: the whole strip is one
hover/click target mapped to the nearest tick, with a per-prompt preview
card that follows the cursor
- Cap the rail at a fixed window of ticks; hovering the edges carousels
through the rest, with gradient masks hinting at more content
- Render ticks as a tape that glides to keep the active prompt centered,
remounting on history prepend to avoid spurious slide animations
- Keep load-earlier as a compact button aligned over the tick column
* fix(chat): shrink navigator gutter when message column sits under it
On narrow windows the centered message column extends under the rail's
full-width invisible hover zone, which swallowed clicks on the right edge
of user bubbles — including the expand/collapse control. Measure the
column against the gutter and switch to a narrow hit zone when they
overlap.
* fix(sync): stop runaway history auto-load on sessions with empty assistant messages
An assistant message fetched with zero parts (e.g. a run aborted before any
output) was stored as absence — indistinguishable from parts that were never
fetched. getSessionMaterializationStatus therefore reported the session as
never renderable, so the ensure-renderable effects (ChatContainer,
ModelControls) retried syncSession forever; each retry refetched the whole
grown window and fired another background prepend, progressively loading the
entire history of large sessions on open.
Commit an explicit empty [] snapshot for assistant messages so fetched-empty
counts as renderable, while non-assistant messages keep the absent
representation and its no-op commit behavior.
Reproduced and verified headless against a real 857-message session: before,
20 message fetches escalating to limit=857; after, one initial page and a
single progressive-mount prepend.
* feat(chat): add desktop prompt navigator rail
Add a ChatGPT-style right-center prompt marker rail for web/desktop chat
with hover/keyboard preview panel, load-more for partial history (panel only),
Chat setting, and mod+alt+p shortcut. Disabled in VS Code across rail,
shortcut, settings, help, and search surfaces.
Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>
* fix(ui): read promptNavigatorEnabled from getState in shortcut handler
Match the file convention used by other shortcut handlers so the toggle
does not rely on a hook-level selector closure.
Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>
* fix(ui): drop use-no-memo and default prompt navigator off
Remove the project-unprecedented React Compiler opt-out, and ship the
prompt navigator as opt-in to match other recent chat UI toggles.
Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>
Omits the craft-goal command from draft starters in VS Code
Prevents the starter from being resolved or pinnable in that runtime
Keeps non-VS Code behavior unchanged
- the goal dialog renders as the shared MobileOverlayPanel bottom sheet
on mobile instead of a centered dialog
- tapping the target button to ARM keeps the soft keyboard open (the next
message is the objective; same guard as the attachment/mic buttons),
while opening the manage sheet lets the keyboard close as usual
- capacitor: bottom-sheet overlays cap their height by the keyboard inset
AND the top safe area — with the keyboard raised while typing inside a
sheet, 100dvh does not shrink (native resize is off) and the panel could
slide under the notch/status bar; applies to every MobileOverlayPanel
Adds /craft-goal autocomplete and chat handling for starting a Goal crafting session.
Introduces new Magic Prompts content and localized labels/descriptions for Goal crafting.
Migrates desktop draft starters to include Craft a Goal once and persists the migration marker.
Compaction fixes (observed in a real long run):
- the summary message's zeroed tokens froze the goal counter at its
pre-compaction value; segments now close with the previously displayed
total as a continuity floor
- audits and continuations after a summary tail now take execution params
(provider/model/agent/variant) from the newest non-summary assistant
turn instead of inheriting agent 'compaction' and the summarize model
File-backed objectives:
- the objective text lives in <data-dir>/goals/<sessionId>.md, keyed by
session id (one goal per session, a new goal overwrites the file);
metadata carries only an objectiveFile flag so session.updated fanout
stays light, and never a path — ids are pattern-validated before any
filesystem access
- limit raised to 5000 chars, no snapshot field: the UI fetches content
via PUT/GET/DELETE /api/goals/objective/:sessionId (behind the blanket
/api auth gate), writes the file before stamping metadata, and falls
back to an inline objective when the write fails
- the loop reads the file fresh on every tick, so objectives are
live-editable mid-goal; a missing file falls back to the inline text
- scheduled goal tasks write the objective file server-side; VS Code
degrades to the audit note (route unavailable there by design)
Arm the target button in the composer and the next prompt becomes a goal:
the server keeps the session working toward it (idle tick -> small-model
audit -> continuation) until the objective is verifiably complete, blocked,
or out of budget — even with the UI closed.
Server (packages/web/server/lib/session-goal):
- event-driven loop on the global SSE hub; goal state lives in
session.metadata.openchamber.goal (merge-safe patches, stale-write guard
by goal id), so it survives restarts and syncs to every client for free
- the small-model audit (objective + last assistant turn only, language
pinned to the objective) is the sole termination authority; blocked needs
3 consecutive verdicts, audit outages tolerate one unaudited continuation
then stop the goal as resumable-blocked
- hard stops: optional token budget, auto-continuation cap (Resume grants a
fresh allowance), turn errors; user abort pauses the goal instead of
blocking it, and resuming over an aborted tail nudges immediately
- token accounting as a snapshot of the latest turn (input + cache.read +
output), goal-relative via a creation baseline and segmented across
compactions; a compaction summary skips the audit and continues
- continuations reuse the session's own provider/model/agent/variant
UI:
- three-mode target button (arm / disarm / manage dialog), informational
goal strip with inline pause/resume and an Evaluating indicator, sidebar
state glyph, objective length counter (2000-char server clamp),
read-only completed goals
- goal entry points: composer (sessions and drafts), start-new-session-
from-answer dialog, plan implement dialog (plan content becomes the
objective), scheduled tasks (Run as goal + budget)
- Settings -> Chat -> Goal: feature toggle + default token budget with
three-layer parity (web server, client persistence, VS Code bridge);
VS Code renders goal state but hides the entry points (the loop runs in
the web server only)
Notifications: per-turn "ready" notifications are suppressed while a goal
is active; settling sends one final notification (desktop, web-push, APNs
generic titles with the session name as body) honoring the completion
toggle. Error/question/permission notifications are untouched.
Docs: user guide (session-goals) in all 9 locales + sidebar entry,
scheduled-tasks cross-reference, server module DOCUMENTATION.md.
* feat(settings): add editor font size setting for chat input and code editor
Adds an 'Editor font size' control in Settings > Appearance that sets an
absolute px font size for the chat input textarea and the in-app
CodeMirror editor. Mirrors the existing terminalFontSize lifecycle.
- New store field editorFontSize (default 13, clamp 9-32, step 1) in
useUIStore with narrow selectors at each consumer.
- Persistence wired through appearanceAutoSave, desktop + runtime API
types, and persistence.ts read/normalize.
- Settings UI row (NumberInput) with reset to 13, VisibleSetting union
entry, OpenChamberPage registration, and search index entry
appearance.editor-font-size.
- Applied as a post-zoom absolute override on the chat input textarea
and on the CodeMirror theme's content rule, leaving gutter/line-number
chrome at its existing hardcoded sizes (matches terminal scope).
- All 10 locales translated (en, es, fr, ja, ko, pl, pt-BR, uk, zh-CN,
zh-TW); no English placeholders in non-English dictionaries.
Refs #1325
* fix(codemirror): use unitless lineHeight so it scales with editor font size
The & rule in the CodeMirror theme set lineHeight to 1.5rem (~24px),
which does not scale when editorFontSize is increased (e.g., 28-32px).
This causes overlapping lines at larger font sizes.
Change to unitless 1.5, which scales proportionally with whatever fontSize
resolves to (dynamic prop or --text-code fallback). Matches browser best
practice for proportional leading.
Review comment: https://github.com/openchamber/openchamber/pull/2065
---------
Co-authored-by: bashrusakh <bashrusakh@users.noreply.github.com>
Co-authored-by: Bohdan Triapitsyn <artmore@protonmail.com>
React error #31 (Objects are not valid as a React child) was thrown
intermittently when a task/subagent tool returned structured data
(e.g. { TODO: '...' }) in a field that the OpenCode SDK types as a
plain string. Pathological payloads would propagate into JSX children
without runtime validation, white-screening the chat until refresh.
This change adds a single `coerceToText` helper in toolRenderers.tsx
and applies it at every vulnerable JSX expression:
- ToolPart.tsx:1807,1975 {state.error} (typed string, can be object)
- ToolPart.tsx:1825 {q.question} (QuestionCard input cast)
- ToolPart.tsx:1830 {opt.label} (QuestionCard input cast)
- ToolPart.tsx:1848 task tool markdown output
- ToolPart.tsx:1898 ToolScrollableTextOutput entry
- toolRenderers.tsx {todo.content} x4 in renderTodoOutput
renderTodoOutput now also validates the parsed array at the boundary
(JSON.parse result is filtered to objects whose content and status are
runtime strings), so a single bad row no longer poisons the whole
tool output.
Tests: 12 new unit tests in
packages/ui/src/components/chat/message/parts/__tests__/issue-2011-react-error-31.test.ts
covering the {TODO}-key object path, circular references, and
non-string content/status on parsed todos.
Co-authored-by: bashrusakh <bashrusakh@users.noreply.github.com>
* fix(chat): preserve chronological message order during history pagination
The baseDisplayMessages dedup loop iterated from tail to head (newest
to oldest), keeping the newer occurrence of each message ID. During
history pagination (prepend mode), the server returns older messages
that may overlap with the current view at the boundary. The tail-first
iteration discarded the older (prepended) duplicate in favor of the
newer (existing) one, breaking chronological ordering.
Change the loop to iterate head to tail (oldest to newest) so the
first occurrence of each time-sortable message ID is preserved. Remove
the now-unnecessary .reverse() call.
Fixes#2088
* test(chat): add dedup logic coverage for baseDisplayMessages
Covers message ID deduplication in baseDisplayMessages useMemo:
- First-occurrence preservation during dedup
- Input order maintenance
- Empty input, single-element, all-same-ID edge cases
- History pagination prepend scenario with overlapping IDs
---------
Co-authored-by: bashrusakh <bashrusakh@users.noreply.github.com>
The "Add to Context" command and the active-editor pin-selection suggestion both create selection attachments but used the basename only (e.g. assist.ts:47). OpenCode synthesizes its Read call from that filename, so the directory was lost and the model could read or edit the wrong file when names collide.
Use the workspace-relative path (asRelativePath(uri, false)) in both paths so the filename carries the directory and the two paths produce identical filenames, restoring attachment dedup.
Fixes#1914
Tool JSON output now starts with a compact navigable summary view.
Expandable tool output includes quick open-file and diff actions for changed files.
Reasoning headers strip stray HTML comments, and navigation tools stay compact.
The tanstack rows sit in a wrapper offset with transform: translateY(), and a
transformed ancestor becomes the sticky containing block — turn headers stuck
to the wrapper's overscan-dependent top edge instead of the scroll container,
floating over the previous turn. Offset the wrapper with padding-top instead:
identical geometry, sticky computes against the scroll container again, and
the padding only changes when the virtual window shifts, not per scroll frame.
* fix: open mobile model/agent panels on tablet-width Capacitor shells and keep composer taps from dismissing the keyboard
* feat: add iPadOS-style split layout to the Capacitor app
- classify the Capacitor shell as mobile in device detection so shared
surfaces (draft starters, panels) stop falling into tablet branches
- add isIPadApp() and useOrientation() helpers
- iPad: persistent full-height sessions sidebar (mobile sessions surface
inline), Changes/Files in a right sidebar with header shortcut toggles
- animate sidebar open/close like the desktop sidebars and add
finger-sized drag-resize with persisted widths
- anchor the overflow menu and the usage/metadata popover next to their
header buttons regardless of open sidebars
* fix: re-anchor metadata popover on layout shifts and untangle sidebar toggle updates
- recompute the iPad metadata popover anchor via a ResizeObserver on its
wrapper so sidebar toggles/resizes while it is open cannot leave it
misplaced
- move the portrait right-panel close out of the setIpadSidebarOpen
updater into plain sequential state updates
Avoids syncing code line numbers while markdown is still streaming
Keeps code block wrapping and line numbers stable after render
Updates code block layout to support deferred gutter insertion
Normalizes bare ---/+++ headers and paths before rendering
Repairs loosely formatted hunk bodies for patch display
Recounts hunk ranges so diff headers stay accurate
Adds line-number gutters for markdown code blocks
Keeps gutter heights in sync when wrapping or resizing changes
Applies wrap styles directly to pre and code for better overflow handling
Adds a chat code block wrap toggle in markdown code block headers
Persists and restores the setting across desktop/web settings
Adds localized labels and OpenChamber search entry for the new option
Adds a Last turn scope to DiffView that renders OpenCode snapshot diffs from the latest user message summary without re-fetching git contents. The view hides Review in that mode and carries the selected diff scope through main and context-panel navigation.
Connects latest-turn changed-file chips in chat to the snapshot diff view on desktop and mobile, while keeping older turn chips static/read-only to avoid misleading affordances and extra subscriptions. Updates localized labels and empty states plus changelog.
Validation: bun run type-check (packages/ui); bun run lint (packages/ui).
Moves the hidden file input out of the attachment controls so it stays mounted.
Prevents file selections from being lost when the composer variant changes.
Restores reliable local attachment uploads after opening the OS picker.
Keeps mobile composer controls from missing taps during keyboard blur/reflow
Applies the deferred blur behavior to mobile browsers and installed PWAs
Leaves Capacitor behavior unchanged
The dictation overlay is absolutely positioned over the composer, so the
transcript could not expand it — long dictations clipped after two lines.
ComposerDictation now measures the transcript text block (not the flex-1
container, which would feed the composer's own height back and creep a few
px per update) and reports it to ChatInput, which feeds it into the
textarea autosize: same line cap as typing, transcript area scrolls past
it and follows the newest words. Idle/unmount releases the height, and an
idle sibling instance (mobile footer + wrapper engine) can no longer zero
the active one's report.
- Add standalone-only safe-area padding for the composer (bottom floor +
fullscreen top inset) and top toast offset; env() reports 0 on iOS 26
standalone so a fixed floor is required
- Pin the mobile shell to 100lvh: WebKit leaves 100dvh stuck at the
keyboard-shrunk value after dismissal
- Clamp visual-viewport pinning to documentElement.clientHeight to guard
against stale visualViewport metrics
- Defer the composer blur flip (120ms) so taps on composer controls
survive the keyboard-resize reflow; transition the bottom padding so
the late flip reads as a slide, not a dip
- Restore the keyboard after mobile overlays close: MobileOverlayPanel
dispatches synchronous open/close events, ChatInput refocuses within
the same gesture, holds focus through iOS's tap-settle dismissal,
guards the pill collapse via DOM focus, and reveals the composer form
above the keyboard (programmatic focus skips iOS's native reveal)
Adds a load older button when earlier history is available
Preserves scroll position while older messages are loaded
Shows date-grouped messages with clearer per-message timestamps
File references in chat messages can now use the 'path:start-end'
form (e.g. 'src/foo.ts:120-145'). The reference becomes clickable in
the renderer and, on click, the file opens at the start line. Range
selection is intentionally not done at this layer — the
'path:start-end' form is parsed only so the link resolves to the
correct path; navigation jumps to the start line, matching the
behavior of the existing 'path:line' form.
- Extract the file-reference parser to a dedicated module so it can
be unit-tested without pulling in the markdown renderer's worker
dependencies.
- Add a range branch to 'parseFileReference' and update the
block-code path regex to recognize the new form.
- Switch the colon-form regex to a non-greedy path match so
'path:line:col' is no longer mis-parsed as 'path:line' with the
first numeric suffix dropped into the path.
- Add a unit test covering the new and existing parser forms.
Keeps the textarea reference available during mobile viewport adjustments
Ensures the composer scrolls back into view after keyboard interactions
Updates text selection menu dependencies to include the current session
Adds a server-side "small model" capability: direct, cheap LLM calls that
reuse the user's existing OpenCode provider logins — the mechanism OpenCode
uses internally for titles and summaries but does not expose through the
SDK or plugins. Zero new dependencies; plain fetch with per-provider wire
formats, credentials never leave the server.
Core (packages/web/server/lib/small-model):
- Resolution mirrors OpenCode's session scoping: explicit settings override
→ small_model from the OpenCode config → family scan within the session's
provider → the session's own model. The global provider scan only serves
callers without a session context, and background callers forbid it
entirely (restrictToPreferredProvider), so conversation content never
reaches a provider the user didn't pick — explicit choices excepted.
- Per-provider auth replicating OpenCode's plugin loaders: GitHub Copilot
(device token as bearer, no exchange), ChatGPT plan via the codex
Responses API (single-flight OAuth refresh written back to auth.json),
Anthropic messages, Google generateContent, generic OpenAI-compatible.
- OpenCode's free models (opencode/big-pickle, *-free) are never called
directly; unauthenticated providers are skipped by design.
- Prompt clamping to the model's catalog context limit; thinking disabled
where a wire switch exists (Z.AI/GLM, MiniMax-M3, Gemini Flash); robust
content parsing with a clear error when a thinking model spends its whol
budget on reasoning.
- Settings → Sessions gains a Small Model group: use-default checkbox plus
an override picker limited to authenticated providers, persisted with
web/desktop/VS Code sanitization parity.
Consumers:
- Session assist: a server-side watcher on the global SSE hub generates a
short recap and one suggested follow-up after a session idles quietly fo
a minute, stored on session metadata (openchamber.assist). Freshness is
keyed to the last assistant message id, so new activity invalidates the
payload everywhere with no extra writes. The chat shows the recap under
the last message after five quiet minutes and the suggestion as a
dismissible chip above the composer (tap fills the input, never sends).
Gated by a new Chat setting (default on) that is a hard generation
switch. Language is anchored to the conversation itself, with a
script-mismatch guard against model/backend language hallucination.
- TTS: a third input mode, summarized — long replies are condensed to
spoken prose before playback on any TTS engine.
- Git: commit-message and PR generation moved off the active chat session
onto the small model fed with real diffs and the commit list (bodies
included), with a session-transport fallback for free-model-only setups.
- Notes: Add to notes distills long selections into 1-3 dense sentences
preserving exact identifiers, with verbatim fallback on failure.
Fixes along the way:
- The global event watcher now starts unconditionally; it was gated behind
the desktop-notify env, leaving the server-side event hub dead in
packaged apps.
- OpenCode re-emits message.updated for old user messages after idle; the
watcher no longer mistakes those for new activity.
- Session metadata merges from a fresh read right before the PATCH, so
writes made during the generation window (suggestion dismissals, review
links) are preserved; the assist runtime stops during graceful shutdown.
Mobile browsers don't shrink the layout for the keyboard, so the
fullscreen composer is now pinned to the visual viewport (fixed at its
offset and height, tracked as the browser pans) instead of overflowing
underneath it; the draft screen's normal composer gets the same pinning
anchored to the visible bottom via a rAF tracker, since Safari's own
focused-field reveal proved unreliable there after leaving fullscreen.
The draft and empty-session roots drop their transform-gpu (a transform
would make them the containing block for the pinned form), the app
header hides while the browser fullscreen composer is up (the form can't
out-stack it from inside the composer wrapper's stacking context), and
leaving fullscreen nudges the still-focused field back into view.
Draft starter chips now hide while the keyboard is open in browsers too,
via an oc-browser-keyboard-open root class driven by composer focus.
Command, file/agent, skill, and snippet autocompletes now stop at the top
of the chat area (Capacitor) or the visible viewport edge (mobile
browsers, which pan the page for the keyboard) and may grow that far,
measured live across keyboard settles and viewport changes. On mobile the
keyboard-hint footer and description lines are gone, rows center their
icons, and list overscroll no longer bounces the page behind. Selecting a
command no longer dismisses the keyboard (its rows now block the tap's
focus steal like the composer buttons), and the dead dismissKeyboard
option is removed.
Starter chips route their text into the submit as an explicit override;
the previous flow staged it in the textarea, which doesn't exist while
the mobile composer is collapsed into the pill, so the submit read an
empty snapshot and silently bailed, leaving the command text sitting in
the input.
The pill's expand focused the textarea from a rAF, outside the user
gesture — mobile browsers only show the soft keyboard for synchronous
focus, so the composer expanded silently. The expand now flushes the
render and focuses in the same gesture, and preventScroll applies only
inside the Capacitor shell so browsers keep their native reveal that
lifts the field above the keyboard (same for the post-overlay keyboard
restore). Also drafts the [Unreleased] changelog entries for everything
since v1.13.9.
The button's visibility mixed the sync meta with the prefetch-cache hint;
a stale prefetch entry (cursor recorded at the initial page) could keep
the affordance alive after the user had already loaded to the top. Sync
now exposes an explicit isComplete (positive confirmation from a fetch,
distinct from !hasMore on unpopulated meta) and it overrides the prefetch
hint, which stays in effect only before the meta knows anything.
Drop the animated pill/composer morph in favor of instant swaps that are
synchronized with the keyboard choreography: a new oc:keyboard-intent
event collapses the composer (flushSync) before the hide compensation is
measured, so keyboard travel and composer height change land as a single
chat motion on both iOS and Android (Android also gains keyboard signals
and deterministic re-pins around its native resize). The WKWebView caret
is hidden during the transition so it no longer flies to its new position.
Draft screen: starter chips hide instantly while the keyboard is up and
the centered title rides the keyboard shift compensation instead of
double-jumping; the composer drag handle also works in dictation mode;
the highlight mirror is disabled on mobile so the caret matches the text.
Fixes: worktree discovery and the GitHub auth probe now wait for the
runtime connection (no more empty branch pickers / stale auth on cold
start), worktree discovery merges per project instead of clobbering the
persisted map, the cross-project session list resets on instance switch
(with an in-flight load guard) so no stale sessions linger, and mobile
overlay content contains its overscroll instead of bouncing the page.
Mobile composer redesign: when the keyboard is closed the input collapses
into a narrow pill (sessions, attach, placeholder, mic) with a round
new-session button that fades away on the draft screen. Model and agent
selectors move into a row above the textarea; the draft project/branch
pickers and the attachment menu become searchable bottom sheets reusing
MobileOverlayPanel; a drag handle (also available while dictating) swipes
the composer into and out of a fullscreen mode.
Keyboard-lifecycle hardening: composer controls (agent cycle, dictation
and its overlay controls) no longer steal focus and dismiss the keyboard;
overlays reopen the keyboard on close via a debounced restore chain that
survives menu-to-picker handoffs and skips the native file picker; open
overlays and dictation keep the composer expanded. Dictation starts
directly from the pill and its overlay content fades in after the shape
settles. The keyboard slide compensates the pill-to-full height change in
one motion, and the mobile highlight mirror is disabled so the caret
always matches the text layout.
Complete rebuild of voice input on a server-authoritative streaming
architecture, replacing the legacy Web Speech / whole-blob / WASM engines
and the dead voice-agent layer (~4k lines removed).
Speech-to-text (dictation):
- Client streams 16 kHz mono PCM16 chunks over /api/dictation/ws with
seq/ack ordering; buffered audio is retained and replayed on reconnect
- Server transcribes and streams live partial transcripts back;
segments auto-commit every ~15s with silence suppression and adaptive
finalization timeouts
- Local provider (default, zero config): sherpa-onnx models in a forked
worker process — auto-download with progress, staged extraction with
verification, corrupt-model auto-recovery, idle shutdown after 5 min
- Model catalog with settings picker (accuracy/speed ratings, sizes,
download/delete): Parakeet TDT v2 (English) and v3 (25 European
languages, auto-detected), Whisper base and tiny (multilingual, light)
- OpenAI-compatible provider for any Whisper endpoint
- Composer overlay with live transcript, volume meter, timer, and
cancel / insert / insert-and-send actions; failed transcriptions keep
their audio for retry or accepting the partial text as-is
- Configurable keyboard shortcut (default mod+alt+v) toggles dictation;
Enter confirms and Escape cancels while recording
- Overlay is pixel-aligned with the composer (measured footer height,
matching paddings/typography/gaps) — no layout shift when toggling
Text-to-speech:
- Local Kokoro provider (English, 11 voices) synthesized in the same
worker via /api/dictation/tts/speak, managed by the shared model
pipeline; sentence-pipelined playback keeps time-to-first-audio at
~1 sentence regardless of message length, and stop cancels in-flight
synthesis
- Sanitizer keeps inline-code content (strips backticks only), reads
interword slashes aloud, and removes only absolute file paths
Settings:
- Voice page unified: a single read-aloud toggle owns all playback
options (the confusing "Enable Voice Mode" is gone); a new "Enable
voice input" toggle (default on, persisted to settings.json) hides
the composer mic entirely when disabled
Mobile and transport:
- iOS/Android microphone permissions added (dictation was previously
impossible on mobile)
- Fixed Android WebSocket upgrades: the Capacitor WebView origin
(https://localhost) was missing from the packaged-client allowlist,
403-ing every WS connection — root cause of the old mobile SSE lock,
which is now removed for all transports
Security and conventions:
- All HTTP routes sit behind the global /api auth gate; the WS upgrade
explicitly validates the UI session and origin, with oc_url_token
narrowly allowlisted and covered by tests; the dictation socket mints
a fresh URL token before connecting
- Routes register before the generic OpenCode proxy; the client goes
through runtimeFetch/getRuntimeUrlResolver, and runtime switches
reset the dictation socket
- VS Code deliberately reports dictation as unavailable (no server
process in that runtime)
CI: workflow Node bumped 20 -> 22 to match the repo engines and fix
better-sqlite3 installs broken by node-gyp@latest on Node 20.
New dependency: sherpa-onnx-node (prebuilt N-API; macOS/Linux x64+arm64,
Windows x64 — Windows-on-ARM falls back to the OpenAI-compatible provider)
- Migrate sidebar session groups, git changes panel, virtualized code
blocks, and JSON tree viewer from virtua to @tanstack/react-virtual;
virtua remains only inside the Pierre diff viewer integration
- Sidebar: preserve scroll position when virtualization enables
mid-session (enable only once the ancestor scroll element is resolved,
seed initial offset from its live scrollTop, render plain rows for the
single pre-paint frame); disable native scroll anchoring on the
sessions scroller; keep row spacing identical between plain and
virtualized modes; absolute row positioning so variable-height rows
cannot drift past the container
- Chat: expand tool/thinking blocks downward by only adjusting scroll
for rows growing above the viewport; raise the desktop history-load
lead to 1.5 viewports so prepends land above the visible area
- Git changes: compute the prefetch window from the first visible row,
skipping overscan rows above the viewport
- Sidebar rows: make the whole highlighted row area clickable, guarded
against double-firing from interactive children
- Replace virtua with @tanstack/react-virtual for chat history on all
surfaces: bottom anchoring (anchorTo: end), key-stable prepend
preservation, and native iOS touch/momentum deferral live in the core
- Patch virtual-core to clamp the render range to real scroll bounds
during transient adjustments
- Rows render in normal flow inside a translated wrapper so sticky user
headers keep working; measurement snapshots cached per session
- Pre-write container height in scrollToFn so the browser cannot clamp
anchor corrections to the stale height; hold the prepend anchor for up
to 180 frames on mobile while fresh rows settle (cancelled by user
input; desktop relies on core anchoring alone)
- Adaptive row-size estimate from per-session measured averages; disable
reveal fade-in for virtualized history rows
- Mobile loads older history only through an explicit localized top
button: no scroll-position trigger and no post-mount background
prepend, so every insert happens from a resting state; a quiet-window
hold defers any stray prepend commit while a touch gesture is active
- Desktop/VS Code keep the seamless scroll-up trigger and progressive
background prepend
- Mobile: defeat iOS momentum scroll when compensating history prepend
(overflow toggle + short rAF watchdog); disable history virtualization
and post-paint background prepends; preload Markdown renderer and use
plain-text Suspense fallback to avoid first-frame geometry shifts
- Desktop: stop double-compensating prepends on the virtualized list -
virtua shift owns the adjustment; remove sticky-anchor heuristics that
misfired as failed restores
- Sync: skip no-op store writes when messages/parts are unchanged