mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-11 16:36:32 +02:00
* feat(runtime): unbind agents on runtime delete instead of destroying them Deleting a runtime archived its agents and then hard-deleted the rows, so the agents and every conversation with them disappeared — while the confirmation dialog said "archive", which a user reasonably reads as recoverable. Retiring a laptop is an ordinary action; losing the agents configured on it is not an ordinary consequence. An agent is now a persistent business object and a runtime is replaceable execution capacity: deleting a runtime unbinds its agents. `runtime_id IS NULL` means unbound — orthogonal to archived — and the agent keeps its instructions, skills, chats, labels, channel installations, autopilots and task history. service.AgentReadiness already refused an agent with no runtime, so the scheduling safety gate needed no change. Two columns become nullable, not one. Without `agent_task_queue.runtime_id`, deleting the runtime still cascades the task history away (and task_message / task_usage / task_token with it), so the agents would survive with no record of anything they did — the same class of loss. A NOT VALID CHECK keeps NULL confined to history: an active task must always have a runtime, so claim / dispatch / delivery-CAS paths can never observe one without. It is written against completed_at rather than a status list so a future non-terminal status fails closed instead of slipping through. Two prerequisites this depends on: - 'deferred' (migration 128) was missing from CancelAgentTasksByRuntimeOrAgent. It went unnoticed because the delete used to cascade those rows away; with the new CHECK it would abort the delete and make the runtime undeletable. - The channel-installation / label / chat-pin / invocation-target / draft-restore cleanups were scoped to "archived agents on this runtime". Archived user agents now survive, so that scope is narrowed to kind='system' — otherwise the fix would produce a subtler loss: agent alive, configuration wiped. Also removes the squad guard that refused (409) when an active squad's leader was an archived agent on the runtime, plus the archived-squad delete that existed only to get past squad.leader_id's RESTRICT FK. The leader is no longer deleted, so nothing needs to be given up to retire a machine. Autopilots are no longer paused either: their assignee survives, and a rebind restores them without the owner having to remember to re-enable. Reason codes: an unbound agent reports agent_runtime_required, not runtime_offline. The copy for runtime_offline tells users to reconnect a machine; an unbound agent has no machine to reconnect, and the fix is to bind a runtime. Chat's bare 409 string gains the same code so the composer can offer that action. API: agents gain runtime_bound. runtime_id stays a string (empty when unbound) so installed clients keep parsing and no gated two-release rollout is needed. The confirmed-delete endpoint is /unbind-agents-and-delete; /archive-agents-and-delete still routes to it, and the compared expected_active_agent_ids set is unchanged — widening it would 409 every older client forever. Co-authored-by: multica-agent <github@multica.ai> * fix: make runtime unbinding recoverable Co-authored-by: multica-agent <github@multica.ai> * fix: address runtime unbind review nits Co-authored-by: multica-agent <github@multica.ai> * fix: resolve runtime unbind review blockers Co-authored-by: multica-agent <github@multica.ai> * fix(migrations): renumber runtime unbind after main merge Co-authored-by: multica-agent <github@multica.ai> * test(daemon): avoid late-request lease flake Co-authored-by: multica-agent <github@multica.ai> * test(autopilots): bind validation fixture runtime Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai>
31 lines
1.2 KiB
SQL
31 lines
1.2 KiB
SQL
-- Reverts 251_agent_runtime_unbind.up.sql.
|
|
--
|
|
-- Restoring NOT NULL requires that no unbound row exists. This down migration
|
|
-- deliberately does NOT delete unbound agents to make itself succeed: deleting
|
|
-- them is precisely the data loss the up migration removes, and an operator
|
|
-- rolling a schema back is not asking to destroy every agent whose machine was
|
|
-- retired.
|
|
--
|
|
-- So this fails loudly while unbound rows are present. To roll back, first
|
|
-- rebind the affected agents (PATCH /api/agents/:id with a runtime_id) and
|
|
-- re-point or accept the loss of history rows whose runtime is gone:
|
|
--
|
|
-- SELECT id, name FROM agent WHERE runtime_id IS NULL;
|
|
-- SELECT count(*) FROM agent_task_queue WHERE runtime_id IS NULL;
|
|
--
|
|
-- Note that a rolled-back application binary still contains the old
|
|
-- archive-then-hard-delete runtime teardown, so runtime deletion must not be
|
|
-- exercised during a rollback window.
|
|
|
|
ALTER TABLE autopilot
|
|
DROP COLUMN IF EXISTS pause_reason;
|
|
|
|
ALTER TABLE agent_task_queue
|
|
DROP CONSTRAINT IF EXISTS agent_task_queue_active_requires_runtime;
|
|
|
|
ALTER TABLE agent_task_queue
|
|
ALTER COLUMN runtime_id SET NOT NULL;
|
|
|
|
ALTER TABLE agent
|
|
ALTER COLUMN runtime_id SET NOT NULL;
|