fix: resilient reconnect — preserve state on fetch fail, pause when offline (#1308)

* fix: preserve state when reconnect-time fetches fail

Several client API methods swallowed fetch/SDK errors and returned an
empty value (`[]`, `{}`), which was indistinguishable from a successful
"server says nothing here" response. Reconnect resync paths trusted that
empty result as authoritative and deleted local state — so after a
network blip (sleep/wake, wifi reconnect, tunnel switch), the UI could
show:

- sessions stuck on the "running" indicator (status never cleared)
- pending permission prompts disappearing from the UI
- pending question prompts disappearing from the UI

and only a page reload would recover. A related case: `listAgents`
silently returning `[]` defeated the 3-attempt retry loop in
`useAgentsStore` because the loop never saw an error.

The systematic fix:

- `getSessionStatusForDirectory` now returns `null` on fetch failure
  (vs the previous `{}`); the reconnect resync treats only a non-null
  response as authoritative — candidates missing from the response are
  written as `{type: "idle"}`, candidates after a failure are left
  untouched.
- `listPendingPermissions`, `listPendingQuestions`, and `listAgents`
  now throw on SDK/network failure. The pre-existing outer try/catch
  blocks in `resyncBlockingRequestsForDirectory` and the retry loop in
  `useAgentsStore` were already in the right shape — they just never
  fired because no exception was thrown. A small `formatSdkError`
  helper renders the SDK `{data, error}` shape into the thrown message.
- `permissionStore.setSessionAutoAccept` catches the new throw and
  falls back to whatever sync-store snapshots provide; the next SSE
  event or reconnect resync will catch up anything missed.

AGENTS.md gets a new "Distinguish fetch failure from empty success"
subsection documenting the principle (throw vs `T | null` patterns,
when to pick which, the retry-loop trap) so this doesn't regress.

Adds 3 regression tests covering the resync paths: existing
questions/permissions are preserved when the corresponding `list*`
method throws, and a permission-fetch failure does not block the
question block from running (verifies per-block try/catch isolation).

* fix: pause reconnect loop when offline or hidden

The SSE/WebSocket reconnect loop retried indefinitely with no awareness
of whether the browser was online or whether the tab was even visible.
Three issues compounded:

- No `online`/`offline` event handling. With a foreground tab on a dead
  network, we'd hit the server every ~5s forever, and on network
  recovery we'd wait up to ~5s for the next probe instead of reacting
  to the `online` event.
- No visibility awareness. A backgrounded PWA on a flaky link kept
  probing at the same rate as a foreground tab. The browser does
  throttle hidden-tab timers, but the intent wasn't expressed in code.
- The "exponential backoff" math
  `min(5000, max(retryDelayMs, 250) * (failures <= 1 ? 1 : 2))`
  re-initialized `retryDelayMs` to 250 every iteration, so the cap of
  5s was never reached — we waited 500ms forever after the second
  failure. Not actually exponential.

Now:

- `online` event aborts the current attempt (if disconnected) and
  cuts inter-attempt waits short. `offline` event aborts so the loop
  enters the slow-probe path immediately.
- `computeRetryDelay` returns the long cap (60s) when `navigator.onLine`
  is false or the tab is hidden; the short cap (5s) when foreground +
  online. The `online` event is the expected recovery path; the 60s cap
  is a fallback for browsers that miss the event.
- Real exponential growth: `BASE * 2^min(failures-1, 8)`, clamped.
- New `waitForRetry` helper interrupts on `online`,
  visibility-becomes-visible, and abort signal — so visibility/network
  recovery doesn't wait out the rest of the current sleep.

AGENTS.md gets a "Reconnect-loop pacing" subsection alongside the
fetch-failure rule, since they're the same family of resilience
concerns.

One regression test: simulates offline + failed first attempt + `online`
event after the failure; verifies the next attempt fires within seconds
instead of waiting the full 60s offline cap.

* fix: long-cap backoff for permanent 4xx server errors

Before this commit the reconnect loop didn't distinguish HTTP error
types. A stuck-path client (wrong URL after server upgrade) or an
expired-auth client (stale token) would hit the server at the normal
5-second cap forever — ~12 reqs/min, indefinitely, with no path to
recovery besides the user reloading.

Now the catch block extracts an HTTP status (looking on `error.status`
and `error.response.status` — the SDK exposes both depending on the
code path) and overrides the backoff:

- 4xx other than 408/429 → use the long cap (60s) immediately.
  Blind retries won't fix wrong path / bad auth / forbidden, so don't
  pound the server. waitForRetry's `online` / visibility-visible
  interrupters still apply — when an operator fixes the server-side
  config and the client comes back to foreground, recovery is prompt.
- 408 (Request Timeout) and 429 (Too Many Requests) → normal
  exponential path. Those are retryable in spirit.
- 5xx / network / unknown → normal exponential path. Unchanged.

AGENTS.md gets a new bullet under "Reconnect-loop pacing" covering
this — the rule fits naturally alongside the existing `navigator.onLine`
and visibility signals.

Two regression tests:
- A 404-throwing SDK doesn't fire a second attempt within 250ms (proves
  we left the exponential path). After `online` interrupts the wait,
  subsequent attempts fire promptly — proves the override doesn't break
  recovery once the underlying problem is fixed.
- A 429-throwing SDK recovers within 2s — proves 429 still hits the
  fast exponential path and isn't caught by the permanent-error branch.

---------

Co-authored-by: vhqtvn <8930337+vhqtvn@users.noreply.github.com>
This commit is contained in:
vhqtvn
2026-05-18 17:47:20 +03:00
committed by GitHub
co-authored by vhqtvn
parent d5cdf464fa
commit ff35f40b43
8 changed files with 598 additions and 52 deletions
+65 -25
View File
@@ -26,6 +26,25 @@ import {
// Use relative path by default (works with both dev and nginx proxy server)
// Can be overridden with VITE_OPENCODE_URL for absolute URLs in special deployments
const DEFAULT_BASE_URL = import.meta.env.VITE_OPENCODE_URL || "/api";
/**
* Render an SDK error payload into a short string for Error messages.
* The SDK returns `{data, error}` shape without throwing on non-2xx; methods
* that need to signal failure (so callers can preserve state instead of
* conflating failure with an empty success) wrap the error with this helper.
*/
function formatSdkError(error: unknown): string {
if (error instanceof Error) return error.message;
if (typeof error === "string") return error;
if (error && typeof error === "object" && "message" in error && typeof (error as { message: unknown }).message === "string") {
return (error as { message: string }).message;
}
try {
return JSON.stringify(error);
} catch {
return String(error);
}
}
const ABSOLUTE_URL_PATTERN = /^[a-zA-Z][a-zA-Z\d+\-.]*:\/\//;
const ID_RANDOM_CHARS = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz";
const ID_RANDOM_LENGTH = 14;
@@ -951,12 +970,20 @@ class OpencodeService {
async getSessionStatus(): Promise<
Record<string, { type: "idle" | "busy" | "retry"; attempt?: number; message?: string; next?: number }>
> {
return this.getSessionStatusForDirectory(this.currentDirectory ?? null);
return (await this.getSessionStatusForDirectory(this.currentDirectory ?? null)) ?? {};
}
/**
* Returns the upstream `/session/status` map, or `null` if the fetch failed.
*
* `null` vs `{}` matters for reconnect resync: the server omits idle sessions
* from the response, so an empty `{}` means "everything is idle" and a candidate
* missing from the response is authoritatively idle. A network/HTTP failure must
* not be conflated with that — return `null` so the caller can preserve state.
*/
async getSessionStatusForDirectory(
directory: string | null | undefined
): Promise<Record<string, { type: "idle" | "busy" | "retry"; attempt?: number; message?: string; next?: number }>> {
): Promise<Record<string, { type: "idle" | "busy" | "retry"; attempt?: number; message?: string; next?: number }> | null> {
try {
const base = this.baseUrl.replace(/\/$/, "");
const url = new URL(`${base}/session/status`);
@@ -974,12 +1001,12 @@ class OpencodeService {
});
if (!response.ok) {
return {};
return null;
}
const data = await response.json().catch(() => null);
if (!data || typeof data !== "object") {
return {};
return null;
}
return data as Record<
@@ -987,14 +1014,14 @@ class OpencodeService {
{ type: "idle" | "busy" | "retry"; attempt?: number; message?: string; next?: number }
>;
} catch {
return {};
return null;
}
}
async getGlobalSessionStatus(): Promise<
Record<string, { type: "idle" | "busy" | "retry"; attempt?: number; message?: string; next?: number }>
> {
return this.getSessionStatusForDirectory(null);
return (await this.getSessionStatusForDirectory(null)) ?? {};
}
/**
@@ -1059,17 +1086,22 @@ class OpencodeService {
return result.data || false;
}
/**
* Throws on fetch/SDK failure. Callers that drive authoritative state from
* the result (e.g. reconnect resync) must let the throw propagate so they
* can preserve existing state instead of conflating "fetch failed" with
* "server returned no pending permissions".
*/
async listPendingPermissions(options?: { directories?: Array<string | null | undefined> }): Promise<PermissionRequest[]> {
const fetches: Array<Promise<PermissionRequest[]>> = [];
const fetchForDirectory = async (directory?: string | null): Promise<PermissionRequest[]> => {
try {
const trimmed = typeof directory === 'string' ? directory.trim() : '';
const result = await this.client.permission.list(trimmed ? { directory: trimmed } : undefined);
return (result.data || []) as unknown as PermissionRequest[];
} catch {
return [];
const trimmed = typeof directory === 'string' ? directory.trim() : '';
const result = await this.client.permission.list(trimmed ? { directory: trimmed } : undefined);
if (result.error) {
throw new Error(`permission.list failed: ${formatSdkError(result.error)}`);
}
return (result.data || []) as unknown as PermissionRequest[];
};
// Try unscoped first (server may return global pending items).
@@ -1133,17 +1165,21 @@ class OpencodeService {
return result.data || false;
}
/**
* Throws on fetch/SDK failure. See {@link listPendingPermissions} for
* rationale — resync paths preserve state on throw via outer try/catch
* instead of conflating failure with an empty server response.
*/
async listPendingQuestions(options?: { directories?: Array<string | null | undefined> }): Promise<QuestionRequest[]> {
const fetches: Array<Promise<QuestionRequest[]>> = [];
const fetchForDirectory = async (directory?: string | null): Promise<QuestionRequest[]> => {
try {
const trimmed = typeof directory === 'string' ? directory.trim() : '';
const result = await this.client.question.list(trimmed ? { directory: trimmed } : undefined);
return (result.data || []) as unknown as QuestionRequest[];
} catch {
return [];
const trimmed = typeof directory === 'string' ? directory.trim() : '';
const result = await this.client.question.list(trimmed ? { directory: trimmed } : undefined);
if (result.error) {
throw new Error(`question.list failed: ${formatSdkError(result.error)}`);
}
return (result.data || []) as unknown as QuestionRequest[];
};
// Try unscoped first (server may return global pending items).
@@ -1257,15 +1293,19 @@ class OpencodeService {
}
// Agent Management
/**
* Throws on fetch/SDK failure so caller-side retry loops (see
* useAgentsStore) can observe failure and retry; silently returning an
* empty list would defeat retries and clear the cached agent list.
*/
async listAgents(): Promise<Agent[]> {
try {
const response = await this.client.app.agents(
this.currentDirectory ? { directory: this.currentDirectory } : undefined
);
return response.data || [];
} catch {
return [];
const response = await this.client.app.agents(
this.currentDirectory ? { directory: this.currentDirectory } : undefined
);
if (response.error) {
throw new Error(`app.agents failed: ${formatSdkError(response.error)}`);
}
return response.data || [];
}
// SSE infrastructure removed — EventPipeline in sync/event-pipeline.ts handles