Commit Graph
22 Commits
Author SHA1 Message Date
Isaac Sanchez-HawkinsandIsaac Sanchez 65e288bb71 fix: respect session prefetch page size (#1332)
Co-authored-by: Isaac Sanchez <isanchez-hawkins@arize.com>
2026-05-19 17:39:37 +03:00
vhqtvnandvhqtvn ff35f40b43 fix: resilient reconnect — preserve state on fetch fail, pause when offline (#1308)
* fix: preserve state when reconnect-time fetches fail

Several client API methods swallowed fetch/SDK errors and returned an
empty value (`[]`, `{}`), which was indistinguishable from a successful
"server says nothing here" response. Reconnect resync paths trusted that
empty result as authoritative and deleted local state — so after a
network blip (sleep/wake, wifi reconnect, tunnel switch), the UI could
show:

- sessions stuck on the "running" indicator (status never cleared)
- pending permission prompts disappearing from the UI
- pending question prompts disappearing from the UI

and only a page reload would recover. A related case: `listAgents`
silently returning `[]` defeated the 3-attempt retry loop in
`useAgentsStore` because the loop never saw an error.

The systematic fix:

- `getSessionStatusForDirectory` now returns `null` on fetch failure
  (vs the previous `{}`); the reconnect resync treats only a non-null
  response as authoritative — candidates missing from the response are
  written as `{type: "idle"}`, candidates after a failure are left
  untouched.
- `listPendingPermissions`, `listPendingQuestions`, and `listAgents`
  now throw on SDK/network failure. The pre-existing outer try/catch
  blocks in `resyncBlockingRequestsForDirectory` and the retry loop in
  `useAgentsStore` were already in the right shape — they just never
  fired because no exception was thrown. A small `formatSdkError`
  helper renders the SDK `{data, error}` shape into the thrown message.
- `permissionStore.setSessionAutoAccept` catches the new throw and
  falls back to whatever sync-store snapshots provide; the next SSE
  event or reconnect resync will catch up anything missed.

AGENTS.md gets a new "Distinguish fetch failure from empty success"
subsection documenting the principle (throw vs `T | null` patterns,
when to pick which, the retry-loop trap) so this doesn't regress.

Adds 3 regression tests covering the resync paths: existing
questions/permissions are preserved when the corresponding `list*`
method throws, and a permission-fetch failure does not block the
question block from running (verifies per-block try/catch isolation).

* fix: pause reconnect loop when offline or hidden

The SSE/WebSocket reconnect loop retried indefinitely with no awareness
of whether the browser was online or whether the tab was even visible.
Three issues compounded:

- No `online`/`offline` event handling. With a foreground tab on a dead
  network, we'd hit the server every ~5s forever, and on network
  recovery we'd wait up to ~5s for the next probe instead of reacting
  to the `online` event.
- No visibility awareness. A backgrounded PWA on a flaky link kept
  probing at the same rate as a foreground tab. The browser does
  throttle hidden-tab timers, but the intent wasn't expressed in code.
- The "exponential backoff" math
  `min(5000, max(retryDelayMs, 250) * (failures <= 1 ? 1 : 2))`
  re-initialized `retryDelayMs` to 250 every iteration, so the cap of
  5s was never reached — we waited 500ms forever after the second
  failure. Not actually exponential.

Now:

- `online` event aborts the current attempt (if disconnected) and
  cuts inter-attempt waits short. `offline` event aborts so the loop
  enters the slow-probe path immediately.
- `computeRetryDelay` returns the long cap (60s) when `navigator.onLine`
  is false or the tab is hidden; the short cap (5s) when foreground +
  online. The `online` event is the expected recovery path; the 60s cap
  is a fallback for browsers that miss the event.
- Real exponential growth: `BASE * 2^min(failures-1, 8)`, clamped.
- New `waitForRetry` helper interrupts on `online`,
  visibility-becomes-visible, and abort signal — so visibility/network
  recovery doesn't wait out the rest of the current sleep.

AGENTS.md gets a "Reconnect-loop pacing" subsection alongside the
fetch-failure rule, since they're the same family of resilience
concerns.

One regression test: simulates offline + failed first attempt + `online`
event after the failure; verifies the next attempt fires within seconds
instead of waiting the full 60s offline cap.

* fix: long-cap backoff for permanent 4xx server errors

Before this commit the reconnect loop didn't distinguish HTTP error
types. A stuck-path client (wrong URL after server upgrade) or an
expired-auth client (stale token) would hit the server at the normal
5-second cap forever — ~12 reqs/min, indefinitely, with no path to
recovery besides the user reloading.

Now the catch block extracts an HTTP status (looking on `error.status`
and `error.response.status` — the SDK exposes both depending on the
code path) and overrides the backoff:

- 4xx other than 408/429 → use the long cap (60s) immediately.
  Blind retries won't fix wrong path / bad auth / forbidden, so don't
  pound the server. waitForRetry's `online` / visibility-visible
  interrupters still apply — when an operator fixes the server-side
  config and the client comes back to foreground, recovery is prompt.
- 408 (Request Timeout) and 429 (Too Many Requests) → normal
  exponential path. Those are retryable in spirit.
- 5xx / network / unknown → normal exponential path. Unchanged.

AGENTS.md gets a new bullet under "Reconnect-loop pacing" covering
this — the rule fits naturally alongside the existing `navigator.onLine`
and visibility signals.

Two regression tests:
- A 404-throwing SDK doesn't fire a second attempt within 250ms (proves
  we left the exponential path). After `online` interrupts the wait,
  subsequent attempts fire promptly — proves the override doesn't break
  recovery once the underlying problem is fixed.
- A 429-throwing SDK recovers within 2s — proves 429 still hits the
  fast exponential path and isn't caught by the permanent-error branch.

---------

Co-authored-by: vhqtvn <8930337+vhqtvn@users.noreply.github.com>
2026-05-18 17:47:20 +03:00
Isaac Sanchez-HawkinsandIsaac Sanchez e50a484633 fix(sync): drop orphan session parts (#1183)
* fix(sync): drop orphan session parts

* fix(sync): guard missing part cache

---------

Co-authored-by: Isaac Sanchez <isanchez-hawkins@arize.com>
2026-05-12 11:05:37 +03:00
Isaac Sanchez-HawkinsandIsaac Sanchez eb5b1de9b7 test(sync): shorten websocket fallback test (#1211)
* test(sync): shorten websocket fallback test

* test(sync): avoid duplicate fallback cleanup

---------

Co-authored-by: Isaac Sanchez <isanchez-hawkins@arize.com>
2026-05-12 10:57:27 +03:00
14c0bfe0bc fix(sync): skip duplicate status events (#1152)
* fix(sync): skip duplicate status events

* test(sync): cover duplicate idle statuses

---------

Co-authored-by: Isaac Sanchez <isanchez-hawkins@arize.com>
Co-authored-by: Bohdan Triapitsyn <artmore@protonmail.com>
2026-05-08 15:56:18 +03:00
Isaac Sanchez-HawkinsandIsaac Sanchez 47fc6a4606 fix(sync): compare retry status metadata (#1141)
* fix(sync): compare retry status metadata

* fix(sync): compare status fields directly

---------

Co-authored-by: Isaac Sanchez <isanchez-hawkins@arize.com>
2026-05-08 15:12:08 +03:00
Isaac Sanchez-HawkinsandIsaac Sanchez a3a82166f1 fix(sync): update request arrays immutably (#1139)
* fix(sync): update request arrays immutably

* test(sync): cover rejected question updates

---------

Co-authored-by: Isaac Sanchez <isanchez-hawkins@arize.com>
2026-05-08 15:07:40 +03:00
Bohdan Triapitsyn e892346c6b refactor: stabilize live chat sync materialization (#1132)
Canonicalize session message/part materialization across load, prefetch, reconnect, and recovery paths so OpenChamber restores session snapshots through one consistent merge flow.

Preserve live assistant streaming text when stale or delayed snapshots arrive, while still replacing optimistic user parts with confirmed server snapshots to avoid duplicated user messages.

Narrow recovery triggers to explicit incomplete snapshot signals instead of broad session-event fallbacks, reducing unnecessary session refetches during active streaming.

Keep turn windowing aligned with parented assistant replies and add regression coverage for materialization gaps, stale snapshot protection, optimistic user replacement, reconnect recovery, and turn grouping.
2026-05-07 18:57:44 +03:00
Bohdan Triapitsyn ff830d3812 fix: protect live session sync caches 2026-05-07 12:30:30 +03:00
2b0e1ef42e fix(sync): preserve pending questions across session switch and directory eviction (closes #918) (#1103)
* fix(sync): preserve pending questions across session switch and directory eviction

Closes #918, completes the gap left by #909.

The 'agent question disappears after switching session / coming back
later' bug had two root causes that #909 only partially addressed:

1. Directory-eviction TTL (20 min) silently dropped child stores that
   held pending questions/permissions. The discard wasn't gated on
   in-flight blocking-request state, so any 'question.asked' event
   that arrived during the eviction-then-rehydrate window was routed
   to a non-existent store and silently lost.

2. PR #909 re-fetches listPendingQuestions/Permissions only on SSE
   reconnect. Switching sessions within the same socket — including
   navigating back to a directory whose child store was rebuilt after
   eviction — left the UI relying on store state that may have missed
   events that fired while a different session was active.

Three edits, in src/sync:

- eviction.ts / types.ts / child-store.ts: add hasPendingBlockingRequests
  to EvictPlan + DisposeCheck and never evict a directory whose store
  carries a non-empty state.question or state.permission record.
- sync-context.tsx: extract resyncBlockingRequestsForDirectory from the
  reconnect path and call it on currentSessionId changes (debounced
  250ms), reusing PR #909's signature-based merge so concurrent SSE
  updates aren't clobbered.
- __tests__/eviction.test.ts, __tests__/session-switch-resync.test.ts:
  new unit coverage for the eviction guard and resync semantics
  (deduped fetch per switch, in-flight SSE preservation, stale entry
  cleanup, unknown-session filtering).

* fix(sync): refresh store before blocking request resync

---------

Co-authored-by: Alexander Busse <alex@ableph.net>
Co-authored-by: Bohdan Triapitsyn <artmore@protonmail.com>
2026-05-05 19:39:35 +03:00
Bohdan Triapitsyn 5614012acb fix(ui): keep streaming deltas through pipeline 2026-05-03 13:42:50 +03:00
jwcrystal 9424cff02c fix: reconnect SSE immediately on OS wake-from-sleep (#1066)
* fix: reconnect SSE immediately on OS wake-from-sleep

When the desktop app resumes from OS sleep, TCP connections are dead
but timers were paused during sleep so the heartbeat watchdog doesn't
fire until ~30s after wake.

Add Electron powerMonitor.resume → renderer notification → event-pipeline
immediate abort, cutting reconnection delay from ~30s to ~0ms.

Changes:
- electron/main.mjs: import powerMonitor, emit openchamber:system-resume
  to all renderer windows on OS resume
- ui/sync/event-pipeline.ts: listen for openchamber:system-resume, set
  attemptAbortReason and abort the active SSE/WS attempt to trigger
  immediate reconnection with retryDelayMs=0 and lastEventId preservation

* fix: reconnect SSE immediately on OS wake-from-sleep

When the desktop app resumes from OS sleep, TCP connections are dead
but timers were paused during sleep so the heartbeat watchdog doesn't
fire until ~30s after wake.

Add Electron powerMonitor.resume → renderer notification → event-pipeline
immediate abort, cutting reconnection delay from ~30s to ~0ms.

Changes:
- electron/main.mjs: import powerMonitor, emit openchamber:system-resume
  to all renderer windows on OS resume
- ui/sync/event-pipeline.ts: listen for openchamber:system-resume via
  globalThis.window, set attemptAbortReason and abort the active SSE/WS
  attempt to trigger immediate reconnection with retryDelayMs=0 and
  lastEventId preservation
- Test: event-pipeline-resume.test.js verifies abort → reconnect flow
2026-04-29 12:19:31 +03:00
Bohdan Triapitsyn 9d9af7e263 Improve event stream resilience 2026-04-24 12:30:11 +03:00
Bohdan Triapitsyn f24e6de21b fix: improve event stream reconnect reliability
Recover stalled event streams without dropping the session
Wait briefly for reconnection before showing connection lost errors
Persist Electron server logs for easier disconnect debugging
2026-04-22 21:03:02 +03:00
YifanandBohdan Triapitsyn bee9d19f3a feat(web): add WebSocket transport for message event streaming with SSE fallback (#764)
* feat: add websocket message stream transport

* fix: avoid false missing session directories in sidebar

* fix: re-probe project root session directories

* refactor: use button group for message stream transport

* fix: resolve chat input hook dependency warning

---------

Co-authored-by: Bohdan Triapitsyn <artmore@protonmail.com>
2026-04-17 11:07:26 +03:00
Bohdan Triapitsyn 2cedc443a5 fix(sync): keep deltas after initial part.updated coalescing 2026-04-16 19:16:32 +03:00
Bohdan Triapitsyn af7cf6ec9c fix(sync): prefer freshest session status and persist active-now list 2026-04-16 19:16:32 +03:00
Bohdan Triapitsyn ddc1039d1c fix(sync): unify live session truth across chat and sidebar 2026-04-16 19:16:32 +03:00
jwcrystalandBohdan Triapitsyn ea5c19e934 fix(sync): deduplicate overlapping delta after coalesced part.updated (#916)
* fix: hide archived section and empty folders when no sessions remain

- Only push archived group in useSessionGrouping when there are archived
  sessions, preventing an empty archived section from rendering
- Hide empty folders in archived bucket via shouldKeepFolder check in
  SessionGroupSection (folders with no sessions and no content in
  children are filtered out)
- Always filter folders through shouldKeepFolder, not just during search

* perf: memoize archived folder filtering

* fix(sync): deduplicate overlapping delta after coalesced part.updated

When message.part.updated coalesces in the event pipeline and a
message.part.delta for the same part arrives in the same flush window,
the reducer appends the delta verbatim to the already-complete field
value, producing duplicated text in tool output and assistant messages.

Add targeted overlap reconciliation: only when a part.updated replaces
an existing part with overlapping string content, mark the next delta
for that field as dedupe-eligible. Normal streaming deltas remain
untouched (pure append).

Covers four cases:
- Full overlap: delta already present -> no-op
- Partial overlap: only non-overlapping suffix appended
- No overlap: unchanged append behavior
- Legitimate repeated output (ha + ha -> haha): preserved

---------

Co-authored-by: Bohdan Triapitsyn <artmore@protonmail.com>
2026-04-15 10:12:30 +03:00
jwcrystalandBohdan Triapitsyn 1656c3bb93 perf(sync): optimize multi-session event pipeline with per-directory queues and delta coalescing (#908)
* fix: hide archived section and empty folders when no sessions remain

- Only push archived group in useSessionGrouping when there are archived
  sessions, preventing an empty archived section from rendering
- Hide empty folders in archived bucket via shouldKeepFolder check in
  SessionGroupSection (folders with no sessions and no content in
  children are filtered out)
- Always filter folders through shouldKeepFolder, not just during search

* perf: memoize archived folder filtering

* perf(sync): per-directory event queues to eliminate cross-session HoL blocking

The SSE event pipeline previously used a single global queue and a single
flush timer shared across all directories. Under concurrent multi-session
workloads, a busy directory's delta storm would block other directories'
status and state events from reaching the UI until the next flush tick,
producing the "multi-session latency" symptom users report.

Split the queue into one DirectoryQueue per directory, each with its own
coalesce map, stale-delta set, and flush timer. Directories flush
independently so a busy directory can no longer starve a quiet one. Coalesce
keys are now scoped to a single directory's queue, so the directory prefix
is removed from the key strings.

Cross-directory behavior only; same-directory multi-session behavior is
unchanged (React 18 auto-batching still collapses a single directory's
flush into one render).

* perf(sync): coalesce consecutive message.part.delta events per flush window

Within a 16ms flush window, consecutive delta events for the same
(messageID, partID, field) tuple are string-concatenated into a single
accumulated delta rather than being queued individually.

This directly addresses same-project multi-session workloads — most
notably parent sessions with subagent tasks (child sessions share the
same directory queue). Both parties stream deltas concurrently, which
previously multiplied raw event count proportionally to the number of
active sessions. Coalescing can reduce queue depth by 10-100x during
active streaming.

Safety: verified against event-reducer.ts — the delta handler is a pure
string append (existingValue + props.delta) with no per-event side
effects (no time.updated, no notifications, no diff calculations). The
merged result is semantically identical to applying each delta separately.

The staleDeltas skip mechanism is unaffected: accumulated delta payloads
retain their type and identifiers, so message.part.updated supersession
still works correctly.

* test(sync): cover per-directory queues and delta coalescing

Extend event-pipeline.test.js with behavioural coverage for both
optimizations landed in 98d013a and 258acf0:

P1 (per-directory queues)
- Delivers events from two directories without loss
- Keeps distinct sessionIDs in the same directory as independent coalesce
  slots (session.status is not overwritten across sessions)
- Collapses repeated session.status for the same session down to latest

Option C (delta coalescing)
- Accumulates consecutive deltas for the same (messageID, partID, field)
  into a single dispatched event with concatenated content
- Does not merge deltas across different fields on the same part
- Does not merge deltas across different parts on the same message
- Does not merge deltas across different directories (per-dir queues)
- Skips accumulated deltas when message.part.updated is coalesced onto
  an earlier update, proving staleDeltas still works with C
- Leaves non-delta coalescing (session.status replace semantics) intact

All 13 tests pass under bun:test.

Also adds event-pipeline.bench.js, a runnable synthetic benchmark that
reports delta reduction and byte integrity across 8 workload scenarios
from "single session, 500 tokens" up to "10 projects × 5 sessions ×
1000 tokens". Run with:

  bun packages/ui/src/sync/__tests__/event-pipeline.bench.js

Current numbers on this machine: 99.5% - 99.9% delta event reduction
with full byte-level integrity (concatenated delta bytes always equal
the input total).

* fix(sync): remove staleDeltas — it silently drops delta events

---------

Co-authored-by: Bohdan Triapitsyn <artmore@protonmail.com>
2026-04-14 20:08:40 +03:00
jwcrystal 964c209ff6 fix(sync): remove stale-delta skip and add parts-gap recovery (#889)
The pipeline's stale-delta mechanism incorrectly marked all
message.part.delta events as stale when a message.part.updated
coalesced, regardless of queue position. This caused valid streaming
deltas to be silently dropped, resulting in blank or incomplete
assistant messages.

Additionally, when part events were dropped by the reducer (missing
parts array or partID not found), there was no recovery path — the
state stayed permanently out of sync until the next SSE reconnect
or manual refresh.

Also discovered: message.updated that successfully writes an assistant
message but has empty parts would render a blank bubble, with no
repair triggered since repair only ran on reducer return false.

Changes:
- Remove staleDeltas Set and deltaKey from event-pipeline.ts
- Coalesce still replaces same-key events, but deltas are never skipped
- Add enqueuePartsRepair + repairSessionParts to sync-context.tsx
  (5s cooldown, deduped, async SDK re-fetch)
- Trigger repair on reducer return false for part events
- Trigger repair on message.updated return true with empty parts
- Add sync debug.ts with gated diagnostic logging
- Add pipeline coalescing tests
2026-04-11 23:40:47 +03:00
Dave Otero dd0587501f fix: recover SSE directory routing in sync pipeline (#830) 2026-04-03 11:46:19 +03:00