fix(scheduled-tasks): prevent dual-server double dispatch of daily tasks (#2713)
* fix(scheduled-tasks): claim schedule occurrences across server instances Two OpenChamber servers sharing project config each armed timers and both dispatched the same daily/weekly/cron/once slot (#2710). Claim the occurrence in shared config under a cross-process write lock before creating a session. Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com> * fix(scheduled-tasks): harden occurrence claim failure and lock ownership Address PR review blockers: release running-slot bookkeeping when claim throws, avoid silently dropping an armed occurrence after a due-slack sync, verify lock-file ownership on release, and cover real on-disk lock behavior. Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com> * fix(scheduled-tasks): always release running slot on state-write failures Wrap runTask bookkeeping in finally so claim, manual-start, and completion lock timeouts cannot stuck-run a task; drop the diskNext claim guard that suppressed later occurrences; recover unparseable locks via mtime age. Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com> * fix(scheduled-tasks): stop re-arming past nextRunAt and clear stuck running Only schedule future nextRunAt values so once-task losers and claim-failed paths cannot spin delay-0 retries. Clear past once nextRunAt on claim, and on completion-write failure retry terminal status so manual runNow still returns the session instead of a hard 500 with lastStatus stuck running. Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com> * fix(scheduled-tasks): release write chain on lock acquire timeout withProjectWriteLock left the in-process promise chain pending when acquireProjectFileLock timed out, wedging every later project write and stranding runTask before finally. Always release the chain; surface persistError on run; record once claim failures in task state. Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com> --------- Co-authored-by: Serhii Dziupin <makeittech@users.noreply.github.com>
This commit is contained in:
committed by
GitHub
co-authored by
Serhii Dziupin
parent
86e6a2ae76
commit
e99f6560be
@@ -9,6 +9,55 @@ Server-owned scheduled task runtime and routes for OpenChamber-only automation.
|
||||
- Runtime orchestration and execution is owned by `packages/web/server/lib/scheduled-tasks/runtime.js`.
|
||||
- This module is OpenChamber feature logic; it is intentionally separate from OpenCode proxy/runtime internals.
|
||||
|
||||
## Cross-instance occurrence claiming
|
||||
|
||||
Multiple OpenChamber server processes can share the same on-disk project config
|
||||
(for example CLI `serve` on port 3000 and the Electron desktop server on port
|
||||
57123). Each process keeps its own timers, so without coordination a daily (or
|
||||
weekly / cron / once) slot would dispatch twice.
|
||||
|
||||
Before a **scheduled** run creates a session, the runtime claims the occurrence
|
||||
in shared project config under the project write lock:
|
||||
|
||||
- Writes `state.lastScheduledFor` to the armed `nextRunAt` timestamp and advances
|
||||
`state.nextRunAt` to the following occurrence.
|
||||
- A second instance that loses the claim skips session creation and reschedules
|
||||
from the winner's persisted `nextRunAt`.
|
||||
- Project config writes also take a cross-process `.json.lock` file so the
|
||||
read-modify-write is serialized across processes, not only within one process.
|
||||
- Lock timeout / filesystem errors on claim, manual-start, or completion state
|
||||
writes always release the in-process running slot (via `finally`) and best-effort
|
||||
re-arm the **next future** occurrence; they must not leave the task permanently
|
||||
"running" or reject unhandled from the queue pump.
|
||||
- Project write locks release the in-process promise chain even when
|
||||
`acquireProjectFileLock` times out, so a later write for the same project can
|
||||
proceed after the on-disk lock is cleared (a hung chain would permanently wedge
|
||||
every mutating API and strand `runTask` before its `finally`).
|
||||
- Re-arm helpers only schedule a persisted `nextRunAt` when it is still in the
|
||||
future. A past slot (common for `once` after claim, which cannot advance
|
||||
`nextRunAt`) falls back to `computeNextRunAt` — which returns null for a
|
||||
consumed/past once occurrence — so a losing instance stops instead of
|
||||
spinning delay-0 timers against the project lock.
|
||||
- On claim lock/fs failure, best-effort persist `lastStatus: error` + `lastError`
|
||||
when nobody else claimed the occurrence, so a past `once` task is not left
|
||||
enabled-but-inert with only a warn log. Recurring schedules still re-arm the
|
||||
next slot.
|
||||
- On completion-write failure after a session already ran, in-memory status is
|
||||
set to a terminal value and a single persist retry is attempted so
|
||||
`lastStatus` does not stay `running`. Manual `runNow` still returns the
|
||||
`sessionID` as a successful dispatch (`ok` follows run status, with
|
||||
`persistError` set) rather than a hard 500; the run API and Scheduled Tasks
|
||||
UI surface `persistError` as a warning toast.
|
||||
- The claim predicate rejects a duplicate solely via `lastScheduledFor` within
|
||||
slack of this occurrence. It does not consult advanced on-disk `nextRunAt`
|
||||
(that field is routinely overwritten by a second instance syncing inside
|
||||
`TASK_DUE_SLACK_MS`, including on later days when `lastScheduledFor` is already
|
||||
set from a prior claim).
|
||||
- Claiming always writes `nextRunAt` (including `undefined`) so a past once-slot
|
||||
is cleared when there is no following occurrence.
|
||||
|
||||
Manual `runNow` does not claim a schedule occurrence.
|
||||
|
||||
## Files
|
||||
|
||||
- `packages/web/server/lib/scheduled-tasks/runtime.js`
|
||||
|
||||
Reference in New Issue
Block a user