mirror of
https://github.com/multica-ai/multica.git
synced 2026-07-31 00:40:46 +02:00
* fix(daemon): isolate runtime poll & heartbeat schedules per runtime A daemon serving multiple workspaces ran a single round-robin poll loop and a single HTTP heartbeat loop across every registered runtime. A 30s HTTP timeout for any one runtime serialized that delay across all the others — observed in production as one workspace's runtimes wedging every other workspace's runtimes on the same daemon. This change: - Replaces the shared runtime-set channel with a multi-subscriber watcher so taskWakeupLoop, heartbeatLoop, and pollLoop can each react to runtime-set changes independently. - Splits heartbeatLoop and pollLoop into supervisor + per-runtime worker goroutines. Each runtime owns its claim cadence and its heartbeat ticker, so a slow request on one runtime no longer blocks any other. - Stagers the per-runtime heartbeat first tick by a jittered delay up to one full interval to avoid a thundering herd at startup. - Sizes the WS writer channel to scale with the runtime count (max(16, 2*N)) so a full per-runtime heartbeat batch always fits; the previous fixed 8-slot buffer dropped heartbeats whenever a daemon watched more than ~8 runtimes. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): acquire execution slot only after ClaimTask, drain pollers before taskWG Two issues from review on the previous commit: 1. Acquiring the shared task slot before ClaimTask reintroduced the very head-of-line blocking the refactor was meant to remove. With MaxConcurrentTasks=1, a slow claim on one runtime parked the only slot for the duration of the HTTP timeout (up to 30s), starving every other runtime's claim attempts. Slots are now acquired after the claim returns a task; other runtimes' pollers stay free to claim. The already-dispatched task waits for a slot under MaxConcurrentTasks bounds, which is the same backpressure shape we had before. 2. pollLoop's shutdown path called taskWG.Wait immediately after cancelling pollers, but a poller could still be between ClaimTask returning a task and taskWG.Add(1). When taskWG's counter is zero that races with Wait — undefined sync.WaitGroup misuse, sometimes panic. Added a pollerWG so the supervisor blocks until every poller goroutine has actually returned before reaching taskWG.Wait. Tests: - TestRunRuntimePollerIsolatesSlowRuntime now uses MaxConcurrentTasks=1 (was 4) so it would have failed under the old slot-before-claim path. - New TestPollLoopShutdownWaitsForPollersBeforeTaskWG drives the exact race window — claim returns a task at the same moment shutdown fires — under -race. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): acquire slot before ClaimTask so capacity-waiters never enter dispatched The previous commit moved slot acquisition AFTER ClaimTask to address a review concern about head-of-line blocking with MaxConcurrentTasks=1. That introduced a strictly worse failure mode: server-side ClaimTask flips the task to `dispatched` immediately (agent.sql:174-176), and the runtime sweeper fails any task in `dispatched` for >300s with `failed/timeout` (runtime_sweeper.go:25-28). When local execution capacity is full and the next claimed task can't acquire a slot within 5 minutes, the user sees the exact failure this issue is fixing — `dispatched_at` set, `started_at` NULL, `failure_reason=timeout`. Reverted to slot-before-claim. The trade-off is the original review concern: with MaxConcurrentTasks=1 and a slow ClaimTask, other runtimes' claims are delayed by up to client.Timeout=30s. That's a 30s polling delay, not a failure — server-side those tasks remain `queued` (no timeout in that state) until a slot frees. 30s ≪ 300s, so other runtimes' tasks cannot get sweeper-failed because of this. The pollerWG fix from the previous commit (avoiding sync.WaitGroup misuse on shutdown) is preserved. Tests: - TestRunRuntimePollerIsolatesSlowRuntime: MaxConcurrentTasks back to 4 (the pre-issue baseline) — the headroom case where slot-before- claim still gives full per-runtime isolation. - New TestRunRuntimePollerSkipsClaimWhenAtCapacity: holds the only slot and verifies the poller never calls ClaimTask while sem is empty. The previous "claim first" path would have failed this. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: multica-agent <github@multica.ai>
10 KiB
10 KiB