fix(scheduled-tasks): prevent dual-server double dispatch of daily tasks (#2713)

* fix(scheduled-tasks): claim schedule occurrences across server instances

Two OpenChamber servers sharing project config each armed timers and both
dispatched the same daily/weekly/cron/once slot (#2710). Claim the occurrence
in shared config under a cross-process write lock before creating a session.

Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>

* fix(scheduled-tasks): harden occurrence claim failure and lock ownership

Address PR review blockers: release running-slot bookkeeping when claim
throws, avoid silently dropping an armed occurrence after a due-slack sync,
verify lock-file ownership on release, and cover real on-disk lock behavior.

Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>

* fix(scheduled-tasks): always release running slot on state-write failures

Wrap runTask bookkeeping in finally so claim, manual-start, and completion
lock timeouts cannot stuck-run a task; drop the diskNext claim guard that
suppressed later occurrences; recover unparseable locks via mtime age.

Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>

* fix(scheduled-tasks): stop re-arming past nextRunAt and clear stuck running

Only schedule future nextRunAt values so once-task losers and claim-failed
paths cannot spin delay-0 retries. Clear past once nextRunAt on claim, and on
completion-write failure retry terminal status so manual runNow still returns
the session instead of a hard 500 with lastStatus stuck running.

Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>

* fix(scheduled-tasks): release write chain on lock acquire timeout

withProjectWriteLock left the in-process promise chain pending when
acquireProjectFileLock timed out, wedging every later project write and
stranding runTask before finally. Always release the chain; surface
persistError on run; record once claim failures in task state.

Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>

---------

Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>
This commit is contained in:
Serhii Dziupin
2026-08-13 16:14:42 +03:00
committed by GitHub
co-authored by Serhii Dziupin
parent 86e6a2ae76
commit e99f6560be
20 changed files with 1436 additions and 117 deletions
@@ -9,6 +9,55 @@ Server-owned scheduled task runtime and routes for OpenChamber-only automation.
- Runtime orchestration and execution is owned by `packages/web/server/lib/scheduled-tasks/runtime.js`.
- This module is OpenChamber feature logic; it is intentionally separate from OpenCode proxy/runtime internals.
## Cross-instance occurrence claiming
Multiple OpenChamber server processes can share the same on-disk project config
(for example CLI `serve` on port 3000 and the Electron desktop server on port
57123). Each process keeps its own timers, so without coordination a daily (or
weekly / cron / once) slot would dispatch twice.
Before a **scheduled** run creates a session, the runtime claims the occurrence
in shared project config under the project write lock:
- Writes `state.lastScheduledFor` to the armed `nextRunAt` timestamp and advances
`state.nextRunAt` to the following occurrence.
- A second instance that loses the claim skips session creation and reschedules
from the winner's persisted `nextRunAt`.
- Project config writes also take a cross-process `.json.lock` file so the
read-modify-write is serialized across processes, not only within one process.
- Lock timeout / filesystem errors on claim, manual-start, or completion state
writes always release the in-process running slot (via `finally`) and best-effort
re-arm the **next future** occurrence; they must not leave the task permanently
"running" or reject unhandled from the queue pump.
- Project write locks release the in-process promise chain even when
`acquireProjectFileLock` times out, so a later write for the same project can
proceed after the on-disk lock is cleared (a hung chain would permanently wedge
every mutating API and strand `runTask` before its `finally`).
- Re-arm helpers only schedule a persisted `nextRunAt` when it is still in the
future. A past slot (common for `once` after claim, which cannot advance
`nextRunAt`) falls back to `computeNextRunAt` — which returns null for a
consumed/past once occurrence — so a losing instance stops instead of
spinning delay-0 timers against the project lock.
- On claim lock/fs failure, best-effort persist `lastStatus: error` + `lastError`
when nobody else claimed the occurrence, so a past `once` task is not left
enabled-but-inert with only a warn log. Recurring schedules still re-arm the
next slot.
- On completion-write failure after a session already ran, in-memory status is
set to a terminal value and a single persist retry is attempted so
`lastStatus` does not stay `running`. Manual `runNow` still returns the
`sessionID` as a successful dispatch (`ok` follows run status, with
`persistError` set) rather than a hard 500; the run API and Scheduled Tasks
UI surface `persistError` as a warning toast.
- The claim predicate rejects a duplicate solely via `lastScheduledFor` within
slack of this occurrence. It does not consult advanced on-disk `nextRunAt`
(that field is routinely overwritten by a second instance syncing inside
`TASK_DUE_SLACK_MS`, including on later days when `lastScheduledFor` is already
set from a prior claim).
- Claiming always writes `nextRunAt` (including `undefined`) so a past once-slot
is cleared when there is no following occurrence.
Manual `runNow` does not claim a schedule occurrence.
## Files
- `packages/web/server/lib/scheduled-tasks/runtime.js`