fix: resilient reconnect — preserve state on fetch fail, pause when offline (#1308)

* fix: preserve state when reconnect-time fetches fail

Several client API methods swallowed fetch/SDK errors and returned an
empty value (`[]`, `{}`), which was indistinguishable from a successful
"server says nothing here" response. Reconnect resync paths trusted that
empty result as authoritative and deleted local state — so after a
network blip (sleep/wake, wifi reconnect, tunnel switch), the UI could
show:

- sessions stuck on the "running" indicator (status never cleared)
- pending permission prompts disappearing from the UI
- pending question prompts disappearing from the UI

and only a page reload would recover. A related case: `listAgents`
silently returning `[]` defeated the 3-attempt retry loop in
`useAgentsStore` because the loop never saw an error.

The systematic fix:

- `getSessionStatusForDirectory` now returns `null` on fetch failure
  (vs the previous `{}`); the reconnect resync treats only a non-null
  response as authoritative — candidates missing from the response are
  written as `{type: "idle"}`, candidates after a failure are left
  untouched.
- `listPendingPermissions`, `listPendingQuestions`, and `listAgents`
  now throw on SDK/network failure. The pre-existing outer try/catch
  blocks in `resyncBlockingRequestsForDirectory` and the retry loop in
  `useAgentsStore` were already in the right shape — they just never
  fired because no exception was thrown. A small `formatSdkError`
  helper renders the SDK `{data, error}` shape into the thrown message.
- `permissionStore.setSessionAutoAccept` catches the new throw and
  falls back to whatever sync-store snapshots provide; the next SSE
  event or reconnect resync will catch up anything missed.

AGENTS.md gets a new "Distinguish fetch failure from empty success"
subsection documenting the principle (throw vs `T | null` patterns,
when to pick which, the retry-loop trap) so this doesn't regress.

Adds 3 regression tests covering the resync paths: existing
questions/permissions are preserved when the corresponding `list*`
method throws, and a permission-fetch failure does not block the
question block from running (verifies per-block try/catch isolation).

* fix: pause reconnect loop when offline or hidden

The SSE/WebSocket reconnect loop retried indefinitely with no awareness
of whether the browser was online or whether the tab was even visible.
Three issues compounded:

- No `online`/`offline` event handling. With a foreground tab on a dead
  network, we'd hit the server every ~5s forever, and on network
  recovery we'd wait up to ~5s for the next probe instead of reacting
  to the `online` event.
- No visibility awareness. A backgrounded PWA on a flaky link kept
  probing at the same rate as a foreground tab. The browser does
  throttle hidden-tab timers, but the intent wasn't expressed in code.
- The "exponential backoff" math
  `min(5000, max(retryDelayMs, 250) * (failures <= 1 ? 1 : 2))`
  re-initialized `retryDelayMs` to 250 every iteration, so the cap of
  5s was never reached — we waited 500ms forever after the second
  failure. Not actually exponential.

Now:

- `online` event aborts the current attempt (if disconnected) and
  cuts inter-attempt waits short. `offline` event aborts so the loop
  enters the slow-probe path immediately.
- `computeRetryDelay` returns the long cap (60s) when `navigator.onLine`
  is false or the tab is hidden; the short cap (5s) when foreground +
  online. The `online` event is the expected recovery path; the 60s cap
  is a fallback for browsers that miss the event.
- Real exponential growth: `BASE * 2^min(failures-1, 8)`, clamped.
- New `waitForRetry` helper interrupts on `online`,
  visibility-becomes-visible, and abort signal — so visibility/network
  recovery doesn't wait out the rest of the current sleep.

AGENTS.md gets a "Reconnect-loop pacing" subsection alongside the
fetch-failure rule, since they're the same family of resilience
concerns.

One regression test: simulates offline + failed first attempt + `online`
event after the failure; verifies the next attempt fires within seconds
instead of waiting the full 60s offline cap.

* fix: long-cap backoff for permanent 4xx server errors

Before this commit the reconnect loop didn't distinguish HTTP error
types. A stuck-path client (wrong URL after server upgrade) or an
expired-auth client (stale token) would hit the server at the normal
5-second cap forever — ~12 reqs/min, indefinitely, with no path to
recovery besides the user reloading.

Now the catch block extracts an HTTP status (looking on `error.status`
and `error.response.status` — the SDK exposes both depending on the
code path) and overrides the backoff:

- 4xx other than 408/429 → use the long cap (60s) immediately.
  Blind retries won't fix wrong path / bad auth / forbidden, so don't
  pound the server. waitForRetry's `online` / visibility-visible
  interrupters still apply — when an operator fixes the server-side
  config and the client comes back to foreground, recovery is prompt.
- 408 (Request Timeout) and 429 (Too Many Requests) → normal
  exponential path. Those are retryable in spirit.
- 5xx / network / unknown → normal exponential path. Unchanged.

AGENTS.md gets a new bullet under "Reconnect-loop pacing" covering
this — the rule fits naturally alongside the existing `navigator.onLine`
and visibility signals.

Two regression tests:
- A 404-throwing SDK doesn't fire a second attempt within 250ms (proves
  we left the exponential path). After `online` interrupts the wait,
  subsequent attempts fire promptly — proves the override doesn't break
  recovery once the underlying problem is fixed.
- A 429-throwing SDK recovers within 2s — proves 429 still hits the
  fast exponential path and isn't caught by the permanent-error branch.

---------

Co-authored-by: vhqtvn <8930337+vhqtvn@users.noreply.github.com>
This commit is contained in:
vhqtvn
2026-05-18 17:47:20 +03:00
committed by GitHub
co-authored by vhqtvn
parent d5cdf464fa
commit ff35f40b43
8 changed files with 598 additions and 52 deletions
+138 -4
View File
@@ -31,6 +31,16 @@ const DEFAULT_RECONNECT_DELAY_MS = 250
const DEFAULT_HEARTBEAT_TIMEOUT_MS = 30_000
const WS_FALLBACK_WINDOW_MS = 60_000
const DEFAULT_WS_READY_TIMEOUT_MS = 2_000
// Retry pacing. Visible+online tabs probe quickly so the user sees connection
// recovery in under a second of real outage; hidden/offline tabs back off
// further so a backgrounded PWA on a flaky link doesn't burn battery probing
// a dead network every few seconds. The browser would throttle hidden-tab
// timers anyway, but this keeps the intent explicit and shrinks server load
// from idle tabs.
const RETRY_BACKOFF_BASE_MS = 250
const RETRY_BACKOFF_CAP_VISIBLE_MS = 5_000
const RETRY_BACKOFF_CAP_HIDDEN_OR_OFFLINE_MS = 60_000
const RETRY_BACKOFF_MAX_EXPONENT = 8
const ABSOLUTE_URL_PATTERN = /^[a-zA-Z][a-zA-Z\d+\-.]*:\/\//
export type EventPipelineInput = {
@@ -261,6 +271,93 @@ export function createEventPipeline(input: EventPipelineInput) {
error instanceof DOMException && error.name === "AbortError" ||
(typeof error === "object" && error !== null && (error as { name?: string }).name === "AbortError")
const isOffline = (): boolean =>
typeof navigator === "object" && navigator !== null && navigator.onLine === false
const isHidden = (): boolean =>
typeof document !== "undefined" && document.visibilityState !== "visible"
// Extract an HTTP status code from anywhere it might be hiding on the
// error object. The SDK's unwrap pattern stashes it on `.status`; raw
// fetch failures may carry `.response.status`; some SDKs also use `.code`.
const extractStatus = (error: unknown): number | undefined => {
if (!error || typeof error !== "object") return undefined
const direct = (error as { status?: unknown }).status
if (typeof direct === "number") return direct
const fromResponse = (error as { response?: { status?: unknown } }).response?.status
if (typeof fromResponse === "number") return fromResponse
return undefined
}
// 4xx errors don't recover from blind retry — wrong path, expired auth,
// bad request body. Keep retrying anyway (a remote reconfigure or reauth
// can fix the underlying problem) but at the long cap so we're not
// hammering the server at 5s intervals indefinitely. 408 (timeout) and
// 429 (rate limit) are retryable in spirit — let them through to the
// normal exponential path.
const isPermanentHttpStatus = (status: number): boolean => {
if (status < 400 || status >= 500) return false
if (status === 408 || status === 429) return false
return true
}
/**
* Wait between reconnect attempts. Resolves early when:
* - the browser fires `online` (network came back — probe immediately),
* - the tab becomes visible (user came back — probe immediately),
* - the pipeline is being torn down (cleanup aborts).
* Otherwise resolves after `ms` like a plain timer.
*/
const waitForRetry = (ms: number) => new Promise<void>((resolve) => {
if (ms <= 0 || abort.signal.aborted) {
resolve()
return
}
const cleanup = () => {
if (timer !== undefined) {
clearTimeout(timer)
timer = undefined
}
if (typeof globalThis.window !== "undefined") {
globalThis.window.removeEventListener("online", onInterrupt)
}
if (typeof document !== "undefined") {
document.removeEventListener("visibilitychange", onVisibilityInterrupt)
}
abort.signal.removeEventListener("abort", onInterrupt)
}
const onInterrupt = () => {
cleanup()
resolve()
}
const onVisibilityInterrupt = () => {
if (typeof document !== "undefined" && document.visibilityState === "visible") {
onInterrupt()
}
}
let timer: ReturnType<typeof setTimeout> | undefined = setTimeout(onInterrupt, ms)
if (typeof globalThis.window !== "undefined") {
globalThis.window.addEventListener("online", onInterrupt, { once: true })
}
if (typeof document !== "undefined") {
document.addEventListener("visibilitychange", onVisibilityInterrupt)
}
abort.signal.addEventListener("abort", onInterrupt, { once: true })
})
const computeRetryDelay = (failures: number): number => {
if (failures <= 0) return 0
// Offline: don't spin probing a dead network. Use the long cap and rely on
// waitForRetry to resolve early when the `online` event fires. The cap is
// also a fallback for browsers that miss `online`.
if (isOffline()) return RETRY_BACKOFF_CAP_HIDDEN_OR_OFFLINE_MS
const cap = isHidden() ? RETRY_BACKOFF_CAP_HIDDEN_OR_OFFLINE_MS : RETRY_BACKOFF_CAP_VISIBLE_MS
const exponent = Math.min(failures - 1, RETRY_BACKOFF_MAX_EXPONENT)
return Math.min(cap, RETRY_BACKOFF_BASE_MS * 2 ** exponent)
}
let streamErrorLogged = false
let attempt: AbortController | undefined
let lastEventAt = Date.now()
@@ -600,9 +697,24 @@ export function createEventPipeline(input: EventPipelineInput) {
: `${currentTransport}_error:unknown`
notifyDisconnected(reason)
// Backoff so a hard-down server doesn't spin the browser event loop.
// Cap at 5s; reset occurs in markConnected().
retryDelayMs = Math.min(5_000, Math.max(retryDelayMs, 250) * (consecutiveFailures <= 1 ? 1 : 2))
// Exponential backoff so a hard-down server / dead network doesn't
// spin the event loop. Caps lower (5s) when the user is foreground
// and the browser thinks it's online; caps higher (60s) when hidden
// or offline so a backgrounded PWA on a flaky link doesn't burn
// battery. waitForRetry below resolves early on `online` or
// visibility-visible so recovery is still under a second.
//
// Override for permanent 4xx errors: stuck-path / bad-auth scenarios
// won't recover from blind retry. Use the long cap immediately so
// the client doesn't pound the server log at 12 reqs/min. The
// waitForRetry interrupters still apply, so a fix on the other end
// followed by `online`/visibility recovery probes promptly.
const status = extractStatus(error)
if (status !== undefined && isPermanentHttpStatus(status)) {
retryDelayMs = RETRY_BACKOFF_CAP_HIDDEN_OR_OFFLINE_MS
} else {
retryDelayMs = computeRetryDelay(consecutiveFailures)
}
}
} finally {
abort.signal.removeEventListener("abort", onAbort)
@@ -617,7 +729,7 @@ export function createEventPipeline(input: EventPipelineInput) {
attemptAbortReason = null
}
if (retryDelayMs > 0) {
await wait(retryDelayMs)
await waitForRetry(retryDelayMs)
}
}
})().finally(flushAll)
@@ -642,6 +754,24 @@ export function createEventPipeline(input: EventPipelineInput) {
attempt?.abort()
}
// Browser told us the network is back. If we're already in a disconnected
// cycle, abort the (stale) attempt and let the loop probe immediately;
// waitForRetry also resolves early on `online`, so any inter-attempt sleep
// ends now. Guard on `disconnected` so a spurious `online` from the browser
// doesn't disrupt a healthy connection.
const onOnline = () => {
if (!disconnected) return
attempt?.abort()
}
// Browser told us we're offline. Abort the current attempt — its socket /
// fetch will throw soon anyway, this just stops sooner. computeRetryDelay
// then returns the long cap so we wait for `online` instead of hammering
// a dead network.
const onOffline = () => {
attempt?.abort()
}
if (typeof document !== "undefined") {
document.addEventListener("visibilitychange", onVisibility)
window.addEventListener("pageshow", onPageShow)
@@ -651,6 +781,8 @@ export function createEventPipeline(input: EventPipelineInput) {
// test environments can replace globalThis.window with a stub.
if (typeof globalThis.window !== "undefined") {
globalThis.window.addEventListener("openchamber:system-resume", onSystemResume)
globalThis.window.addEventListener("online", onOnline)
globalThis.window.addEventListener("offline", onOffline)
}
const cleanup = () => {
@@ -660,6 +792,8 @@ export function createEventPipeline(input: EventPipelineInput) {
}
if (typeof globalThis.window !== "undefined") {
globalThis.window.removeEventListener("openchamber:system-resume", onSystemResume)
globalThis.window.removeEventListener("online", onOnline)
globalThis.window.removeEventListener("offline", onOffline)
}
abort.abort()
flushAll()