mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-04 17:18:35 +02:00
* fix: ingest feishu media as chat attachments * fix: ingest feishu post embedded media * fix(lark): make inbound media retries safe * fix lark media resource limit * fix(lark): move inbound media off ack path * fix(channel): make inbound media runs durable * fix(channel): close enqueue-vs-append race on media deferral EnqueueChatTask read the session-wide media deadline in one statement and sealed the input batch in a later one. Under READ COMMITTED a media message committing between the two got sealed into a task the deadline read had already decided was 'queued', so the daemon could claim it before its attachment bound — the agent received the bare placeholder, and the later media-ready promotion was a no-op against a non-deferred task. After the seal, re-derive the deferral from the sealed batch itself in the same transaction (DeferChatTaskForSealedPendingMedia): if any sealed message still carries an unexpired media marker, flip the task to deferred with fire_at aligned to the latest marker. The existing post-commit promote fence already covers the opposite direction (marker cleared mid-transaction). Adds a deterministic regression test that injects the media append between the deadline read and the seal via a wrapped pgx.Tx. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): keep committed chat task out of enqueue error path The post-commit media-ready fence returned its error from EnqueueChatTask even though the deferred task was already durably committed. The router flush treats any enqueue error as "no task exists": it clears the typing indicator and logs an enqueue failure while the run still happens at its fire_at deadline. Log the fence failure instead — the claim-path deferred promoter re-queues the task regardless. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): cap global media resolution concurrency Media jobs were serialized per session but unbounded across sessions: a burst could open arbitrarily many concurrent 45s Lark downloads, and each unknown-length upload may buffer up to the 100 MiB resource cap in memory. Gate resolveAndBindMedia behind a global slot semaphore (default 8, RouterConfig.MediaConcurrency). Per-session ordering is unchanged; on shutdown a job cancelled while waiting for a slot proceeds straight to the bounded DB finalize so marker clearing stays prompt. Also document that the per-message media budget spans queue/slot waits (it must match the persisted fire_at) and why timed-out uploads cannot leak unbounded orphans. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(chat): keep channel-sealed user messages on task cancel Sealing the channel input batch stamps task_id onto channel user messages, which exposed them to the cancel draft-restore path: an empty-transcript cancel would DeleteUserChatMessageByTask the sealed Feishu/Slack messages and detach their attachments. Those messages are the durable record of what the platform sender wrote — the sender has no Multica composer to restore a draft into. Gate the restore-delete on ChatSessionHasChannelBinding in both the synchronous finalize and the deferred finalize (the latter covers markers left by an older replica during a rolling deploy); a bound session now settles as "Stopped." instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): skip the media pipeline for messages without media Every inbound message on a Media-enabled platform persisted a 45s media deadline and queued a resolution job, so a plain text message could wait behind the global media semaphore (its task deferred while other sessions download 100 MiB videos) and a crash between append and clear delayed a pure-text run to the full 45s fallback. Add MediaResolver.HasMedia — a pure in-memory probe the Router calls on the ACK path — and only persist the deadline / enqueue the job when the message actually references platform media. The Feishu resolver decodes the already-received payload and reports standalone image or video keys and post-embedded img/media spans. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(chat): gate cancel restore on immutable channel provenance The previous guard keyed the cancel restore-delete off ChatSessionHasChannelBinding, but a binding only proves routing exists right now: archiving a session and rebinding an installation both delete the binding while preserving chat history, so a still-cancellable sealed task could again restore-delete the original inbound messages. Persist provenance on the message instead: migration 203 adds chat_message.channel_ingested, stamped inside the channel append transaction and never mutated, and both cancel finalize paths now gate on TaskHasChannelIngestedMessages over the task's sealed batch. The binding-existence query is removed. Regression tests cover ingest -> archive/unbind -> cancel for a queued and a started task. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): reclaim media uploads that never gain an attachment row Deadline expiry dropped already-resolved refs and a BindMedia failure was log-only, leaving uploaded objects with no attachment row and no reclaim path — the dedup mark commits with the message before media runs, so a redelivery is dropped as a duplicate and never re-resolves (and thus never overwrites) those keys, and workspace/session deletion only enumerates the attachment table. Add MediaResolver.DiscardMedia — a best-effort delete by StorageKey — and call it from both failure paths in resolveAndBindMedia. The Feishu resolver forwards to the storage backend's Delete. Tests cover a partial upload discarded at the deadline, discard on bind failure, and key-level deletion in the resolver. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(server): refresh comments stale after detached media ingestion Channel tasks now seal a self-owned input batch, media ingestion is no longer out of scope for the flattener, and MediaRefs are filled by the detached resolver after append rather than by feishuChannel pre-engine. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(chat): stop keying channel empty-completion silence off chat_input_task_id Sealing gave channel tasks a self input-owner, which broke writeChatCompletionOutcome's discriminator: it treated any owned task as direct, so an empty channel completion wrote the no_response fallback row and the outbound patcher — which forwards any non-empty chat:done content verbatim — pushed the English fallback body to Feishu/Slack, violating the MUL-4351 contract. Silence is now decided by the immutable channel_ingested provenance of the task's input batch, looked up by the batch OWNER id (chat_input_task_id): auto-retry clones inherit the owner while their sealed messages stay tagged with the parent's id, so keying off the task's own id would misread a channel retry as direct. The cancel-path provenance gates switch to the same owner key via chatInputOwnerID. chat_input_task_id is back to meaning only "input batch owner". Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(migrations): renumber to 203/204 after upstream took 202 Upstream main merged 202_runtime_profile_add_qwen while this branch held 202/203, tripping TestMigrationNumericPrefixesStayUniqueAfterLegacySet on the CI merge tree. channel_media_pending becomes 203 and channel_ingested becomes 204; no content changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channels): gate outbound delivery on channel provenance, not owner Merging main brought #5645 (keep direct chat replies in Multica), whose outbound gate assumed channel tasks leave chat_input_task_id NULL. Sealed channel tasks own an input batch too, so on the merge tree every channel reply and failure notice was classified as direct and silently dropped — agents stopped replying in Feishu/Slack. Both outbound gates now call engine.TaskInputIsChannelIngested: a NULL owner keeps #5645's deliver-by-default for pre-sealing tasks, an owned batch delivers only when it carries the immutable channel_ingested stamp (keyed by the owner id, so auto-retry clones inherit the verdict). Direct replies stay in Multica; sealed channel replies reach the platform. Tests cover both directions on both platforms. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): discard media orphans on a fresh context, after finalize DiscardMedia shared finalizeCtx with BindMedia, so a bind that failed because the finalize deadline expired handed the storage deletes an already-dead context — the compensation silently no-opped and the orphans leaked anyway. The deadline path also ran S3 deletes before the marker clear, eating the same 5s budget the user-facing bind/promotion needed. Collect the refs from both failure paths, run bind + promotion on the finalize budget first, then delete on a fresh discard context. The bind-failure test now pins that discard receives a live context. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): compensate result-uncertain media uploads and commits The compensation protocol treated "the call returned an error" as "the side effect did not happen", which is wrong in both directions across the result-uncertain windows: - An upload error can follow a server-side write (lost response, deadline mid-write). The attempted key never reached the router, so nothing could reclaim it and dedup guarantees no re-resolve. The resolver now idempotently deletes the deterministic key on a fresh budget right at the failure site. - A commit error is not a rollback guarantee: a lost ack can report failure after Postgres durably committed the attachment rows, and the router's discard would then delete objects those rows reference. BindMediaRefs now converges the ambiguity on a fresh budget — any of the batch's URLs present proves the atomic commit landed (bind reports success); none proves the rollback (discard stays safe); a failed verification returns ErrMediaBindResultUnknown and the router keeps the uploads, preferring a rare orphan over a broken attachment. Fault-injection coverage: an upload error deletes the attempted key; a lost-ack commit keeps the bound attachment and reports success; a verified rollback stays a discardable error; the router keeps uploads on the unknown-outcome sentinel. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(migrations): renumber to 207/208 after upstream took 203-206 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(channel): note DiscardMedia self-invocation and the unknown-outcome skip Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(migrations): renumber to 212/213 after upstream took 207-211 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(channel): replace inline media compensation with an intent ledger and reconciler Inline best-effort compensation cannot answer "did my side effect happen?" at the moment it needs the answer — the DELETE/PUT reordering and the empty-read-vs-in-flight-COMMIT gaps were both instances of the same two-system atomicity problem. Persist the intent instead and let an asynchronous reconciler settle it: - channel_media_pending_object (migration 214; claim index 215 as its own single-statement CONCURRENTLY migration): a state machine row ('pending' -> 'deleting') with lease, attempt, and backoff columns. - The resolver upserts the row BEFORE each PUT, state-guarded so a key the reconciler owns is never resurrected (the resource is skipped). ObjectURL is a pure function of configuration, so the row carries the attachment URL pre-upload. - BindMediaRefs deletes the batch's rows INSIDE the attachment-insert transaction: commit landed <=> intents gone, atomically, so an ambiguous COMMIT never needs adjudication. A key already claimed to 'deleting' is skipped (placeholder stays). - Nothing is ever deleted inline. The reconciler — an independent worker so storage latency cannot starve other sweepers — claims due rows ('pending' past the settle delay, or expired leases) under a fresh lease, checks for a durable attachment reference only AFTER the claim (race-free: bind can no longer succeed on the key), deletes unreferenced objects outside any transaction, and backs off failed deletes with attempt-based retry. Crash windows converge for free. - The settle delay is a fixed constant carrying NO correctness weight; invariant tests pin it at >=10x every pipeline budget. Metrics cover deletes, referenced clears, delete failures, and ledger backlog. Removed: MediaResolver.DiscardMedia, ErrMediaBindResultUnknown, the post-commit verification, and both router discard branches. Tests: intent-before-upload ordering; upload error leaves the row and deletes nothing; bind-wins vs reconciler-wins on the same key; lost-ack and rolled-back commit injections (intent cleared iff the attachment landed); reconciler three-state settle; expired-lease reclaim; delete failure backoff and retry; settle invariants. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): never build or sweep the media reconciler without storage store is nil when S3 is unconfigured AND the local upload dir fails to initialize, but the reconciler was constructed unconditionally and main only gates the goroutine on the reconciler pointer — the first unreferenced ledger row (rows can pre-exist from a boot where storage worked) would nil-pointer panic a bare goroutine and take down the process. Construct the reconciler only when a storage backend exists, and guard RunOnce defensively: with no deleter it skips the sweep without claiming, so rows are not stranded in 'deleting' until lease expiry. Test covers the pre-existing-row + missing-storage boot. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(migrations): renumber to 213-216 after upstream took 212 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(channel): remove the dead pre-resolved MediaRefs ingress path lark.InboundMessage.MediaRefs and the resolver's early-returns for pre-populated refs were vestiges of the pre-detached synchronous design — no producer fills them before the router anymore. Worse, the intent ledger made the path actively misleading: refs arriving without ledger rows would be silently skipped at bind (with a log blaming the reconciler), contradicting the field's "already persisted" contract. Delete the field, its channelMessageFromLark mapping, and both early-returns; channel.InboundMessage.MediaRefs is now documented as what it actually is — ResolveMedia's output channel, always empty on ingress, attachable only through a claimed ledger intent. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): enforce workspace tenancy on every ledger query The intent-ledger upsert's conflict branch guarded only on state, so a cross-workspace storage_key collision could rewrite the row's workspace/message/url ownership; release and delete keyed on (storage_key, lease_token) alone. The derived key embeds the workspace UUID so none of this is reachable today — but tenancy must be enforced by the workspace column in every query, never derived from the key string (MUL-3515 rule, restated in this PR's review). The upsert now updates only within the same workspace (a cross-tenant conflict updates nothing, returns no row, and the resolver skips the upload — the fail-safe direction), and release/delete take (workspace_id, storage_key, lease_token). Tests pin that a foreign workspace can neither steal, release, nor delete a row. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(migrations): build the ledger primary key via a concurrent index storage_key TEXT PRIMARY KEY created its unique index implicitly at CREATE TABLE, against the repo convention that every migration index — including a new table's unique index — is built CONCURRENTLY in its own single-statement migration (the exact three-step pattern client_usage_daily shipped in 207-209). The table now declares storage_key NOT NULL, 216 builds the unique index concurrently, and 217 attaches the primary key USING INDEX; the claim index moves to 218. ON CONFLICT (storage_key) still resolves against the constraint, and the full down/up round-trip is verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): bound each reconciler object delete with its own timeout DeleteObject ran on the worker-lifetime context and the SDK's default HTTP client has no overall request timeout, so one black-holed connection would wedge the sequential sweep loop — and with it every later batch and the backlog gauge — forever; a single-replica deployment has no other worker to reclaim the lease. Each delete now gets a 30s timeout (well under the 2min lease), and a timed-out delete takes the existing release/backoff path. Covered by a blocking-deleter test with an injectable timeout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): anchor the media deadline to the DB clock and bound queue waits by it Two deadline gaps from review: - The persisted marker was an application-clock timestamp compared against SQL now() everywhere it is read, so a skewed app node could shrink the fallback window and hand the agent a placeholder before the resolver's local budget ended. The append transaction now anchors a relative budget (MediaPendingSeconds) with now() + make_interval, writer and readers sharing one clock; the local resolve budget stays monotonic app-side. A DB test pins that the remaining budget measured by the DB clock equals the requested one. - enqueueMedia's waits (per-session order, global slot) only watched shutdown, so in a burst an already-expired job kept its goroutine and payload until it reached the front. Both waits now also watch the message's deadline; on expiry the job skips the resolver entirely and runs only the empty finalize (marker clear + promotion), which also unblocks the session's later messages. Covered by a queued-expiry test that finalizes while the only slot is deterministically held. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(migrations): renumber to 216-221 after upstream took 213-215 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): start the local media budget before the append transaction The DB anchors the durable fallback at insert-time now(), but the local monotonic budget started only after AppendMessage returned — so the resolver outlived the fallback by the append/commit latency, a window where the deferred task is already claimable while the resolver still runs and the agent reads a placeholder that binds moments later. Capture the local deadline before calling AppendMessage, restoring the ordering local-gives-up <= durable-fallback-fires. A slow-append test pins that the resolver's context deadline is measured from the pre-append instant. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(migrations): renumber to 224-229 after upstream took 216-223 Verified against the merged tree: the numeric-prefix uniqueness test passes and the full migration set applies cleanly from scratch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): heartbeat the reconciler lease per row One claim covers up to 50 rows under a single 2-minute lease, but the batch is processed sequentially and each delete may run its full 30s timeout — a few stalled deletes could outlive the lease mid-batch, letting another replica reclaim the tail: duplicate concurrent deletes, inflated attempt/backoff on rows whose owner was alive, and skewed metrics. The lease is now renewed before EACH row's settle work, so it only ever needs to cover one row's worst case (invariant-tested: lease >= 2x the per-delete timeout). A renewal that matches no row means another worker reclaimed it after a genuine expiry — the row is skipped, leaving the new owner's state untouched. Test simulates a mid-batch reclaim and pins the skip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): dedup post media resources and make local writes atomic A rich post may reference the same image_key/file_key in several spans. The object key derives from (message, type, key), so duplicates uploaded to the SAME key twice: LocalStorage.UploadStream truncated the destination up front and removed it outright on a copy error, so a second failing attempt destroyed the object the first success had produced — leaving an attachment row pointing at nothing. A second succeeding attempt instead produced two attachment rows for one object. Collapse duplicate spans by (fetch type, platform key) before the upload loop, and write local uploads through a temp file renamed into place so a failed write can only discard its own temp file. Tests cover a duplicated span uploading once and a failed re-upload leaving the previous object intact. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): fence late-materializing PUTs with a tombstone schedule A DELETE cannot be ordered against a PUT the client already abandoned: the store may materialize the object after the delete completes. The reconciler cleared the ledger row right after deleting, so such an object had no row and nothing to reclaim it — which made the settle delay the de-facto correctness barrier for the PUT/DELETE race, exactly what the design says it must not be. The row is now kept as a tombstone ('tombstoned' state, migration 226's CHECK) and re-deleted on a widening schedule (15m, 1h, 6h, 24h, the pass index carried in last_error), so a late materialization is reclaimed by a later pass; only after the schedule is exhausted is the row dropped. Claim, heartbeat, lease, and tenancy predicates are unchanged — a tombstone is claimed exactly like any other due row. A separate gauge reports tombstones so they cannot be mistaken for a backlog of objects awaiting reclaim, and the header comment now states precisely what state fences (bind/commit) versus what the schedule fences (late PUTs). Tests: the reviewer's interleaving — DELETE completes, the abandoned PUT materializes right after, and the object is gone by the end of the schedule — plus a full schedule walk asserting the object is counted once and the row clears at the end. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): tombstones must re-delete, not re-ask the reference question A tombstone revisit ran the same reference check as a first settle, so an attachment carrying the same URL — a re-ingested copy of the object — sent the row down the "referenced, keep it" branch: the object was kept and the row cleared, abandoning the re-delete schedule that fences the ORIGINAL object against an abandoned PUT. A tombstone has already been judged unreferenced and deleted; it exists only to re-delete whatever materializes later, so it now goes straight to the delete + schedule tail (extracted as settleDeletedObject, shared with the first settle). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(channel): keep the tombstone schedule position in its own column The re-delete pass index was encoded into last_error, which the failure path also writes: one failed re-delete erased the position and restarted the walk. A store failing intermittently could therefore keep a tombstone alive indefinitely — every recovery would resume at pass 1 and the row would never reach the end of the schedule to be dropped. tombstone_pass is now its own column (the table is introduced in this PR, so migration 226 carries it), advanced only by a successful delete, and the tombstone write clears the now-stale last_error. Test walks the schedule across a failed re-delete and asserts it resumes rather than restarts, and that the row still terminates. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(lark): derive media object keys per chat message The object key was derived from the platform message alone, so a second ingest of the same Feishu message reused the first ingest's ledger row. That row can be a tombstone (up to ~31h while the re-delete schedule runs), and the intent upsert refuses anything that has left 'pending', so the second ingest skipped the upload and silently produced a placeholder with no attachment. A re-ingest is reachable: the inbound dedup claim is reclaimable once 60s stale and the dedup row is only vacuumed after 24h. Keying on the chat message the object will attach to keeps the two ingests independent, and nothing leaks: each one's objects are covered by its own ledger row. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(storage): route both local upload paths through one atomic write UploadStream wrote through a temp file and renamed into place, but the buffered Upload path still truncated the destination up front — the destructive shape the stream path exists to avoid, one caller away from coming back. Both now share writeAtomic, which also restores the 0644 the direct write used (CreateTemp makes files 0600). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(channel): gofmt the media-pending append fields Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(storage): keep the local upload chmod best-effort The rename-into-place rewrite made a failed chmod fail the whole upload. CreateTemp's 0600 has to be widened to the 0644 the direct write used, but an upload dir on a mount that ignores chmod (SMB/NFS/FUSE) accepted the old direct write fine — turning those deployments' uploads into hard errors would be a regression for a cosmetic property. Log and continue. Tests pin 0644 on both upload paths, and that a failed buffered upload leaves no temp litter and no damage to a previous object. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(channel): never re-delete an object an attachment references The tombstone pass skipped the reference check and deleted unconditionally, so a durable attachment carrying that URL lost the only object it can read — the dangling attachment the intent ledger exists to prevent, and the opposite of the posture every other path here takes ("a reclaimable orphan beats a broken attachment"). The check now runs on every pass. A positive result on a tombstone is unreachable by design — keys are per (chat message, resource) and a bind cannot attach a key that has left 'pending' — so reaching it means an invariant broke: keep the object, clear the row, log it, and count it on a dedicated reconciler_tombstone_referenced_total counter. The test's contract is flipped to assert the referenced object survives. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(storage): make the local staging file reclaimable after a crash os.CreateTemp's random suffix meant a crash between the staging write and the rename left a file nothing could name: the ledger records only the final storage key, and DeleteObject removed only the object and its sidecar. Each leftover can approach the 100 MiB resource cap and they accumulate without bound. The staging path is now derived from the object key, so DeleteObject removes it alongside the object — which makes the media reconciler reclaim it too, since the intent row is written before the upload. Opening it 0644 directly also drops the chmod the previous commit had to make best-effort. Both read paths refuse the staging name (keys come from the request URL, and a half-written body should not be readable); a user-supplied ".tmp" extension is unaffected, since object keys are generated and never dot-prefixed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(channel): renumber the media migrations after merging main main took 224 (agent_task_session_rollout_missing), so the ledger group moves to 225-230 and the cross-references inside the table migration follow. main's CompleteTask also grew a sessionRolloutMissing parameter; the three call sites this PR added to chat_input_ownership_test.go pass false. Verified the way the numbering is meant to be verified: full migration set applied from scratch on the merged tree, and the whole server suite run against that database. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(channel): let Postgres compute every reconciler deadline The reconciler built settle cutoffs, lease expiry, backoff and re-delete times from the process clock and compared them against the database's now(). A replica whose clock had drifted would therefore settle rows whose upload was still in flight (the object is deleted and the bind then refuses to attach — media silently lost), hand out leases that are born expired (rows churn between workers, attempt/backoff inflate), or compress the tombstone schedule that fences a late-materializing PUT. The four settle queries now take durations and derive their timestamps from now(), so every replica reads one clock. The parameter types are the guard: an app-side timestamp can no longer be passed. Test asserts the persisted lease, backoff and re-delete deadlines all track the database's now(). The generated code also picks up main's new agent_task_queue column in the two RETURNING task.* queries this PR adds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(lark): drop unrelated gofmt-only churn from this PR Six files carried whitespace/comment-reformatting with no functional change, unrelated to the inbound media pipeline. Reverting them to the base revision keeps the diff focused on the feature (75 -> 69 files): server/internal/service/empty_claim_cache.go server/internal/integrations/lark/markdown_detect.go server/internal/integrations/lark/ws_chunk_assembler.go server/internal/integrations/lark/ws_chunk_assembler_test.go server/internal/integrations/lark/ws_frame_test.go server/internal/integrations/lark/registration_test.go Verified: `git diff -w` against these files was already empty, so no behavior is affected. go vet clean; tests covering these files pass. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai>
588 lines
26 KiB
SQL
588 lines
26 KiB
SQL
-- name: CreateChatSession :one
|
|
INSERT INTO chat_session (workspace_id, agent_id, creator_id, title, runtime_id, is_agent_intro, project_id)
|
|
VALUES ($1, $2, $3, $4, (SELECT runtime_id FROM agent WHERE id = $2), $5, sqlc.narg('project_id'))
|
|
RETURNING *;
|
|
|
|
-- name: ClearChatSessionProjectByProject :exec
|
|
-- Project references are intentionally soft (no database FK). Keep chat
|
|
-- history while removing the context selection when a project is deleted.
|
|
-- Do not touch updated_at: context cleanup is not chat activity.
|
|
UPDATE chat_session
|
|
SET project_id = NULL
|
|
WHERE project_id = $1 AND workspace_id = $2;
|
|
|
|
-- name: GetChatSession :one
|
|
SELECT * FROM chat_session
|
|
WHERE id = $1;
|
|
|
|
-- name: GetChatSessionInWorkspace :one
|
|
SELECT * FROM chat_session
|
|
WHERE id = $1 AND workspace_id = $2;
|
|
|
|
-- name: ListChatSessionsByCreator :many
|
|
-- IM-style list: each active session with its unread *count* (assistant
|
|
-- messages after the read cursor), a preview of the latest message, and
|
|
-- ordered by most-recent activity so a new reply bumps a session to the top.
|
|
SELECT cs.*,
|
|
(SELECT count(*) FROM chat_message m
|
|
WHERE m.chat_session_id = cs.id
|
|
AND m.role = 'assistant'
|
|
AND m.created_at > cs.last_read_at)::int AS unread_count,
|
|
COALESCE(lm.content, '') AS last_message_content,
|
|
COALESCE(lm.role, '') AS last_message_role,
|
|
lm.created_at AS last_message_at,
|
|
lm.failure_reason AS last_message_failure_reason,
|
|
COALESCE(lm.message_kind, '') AS last_message_kind
|
|
FROM chat_session cs
|
|
LEFT JOIN LATERAL (
|
|
SELECT content, role, created_at, failure_reason, message_kind
|
|
FROM chat_message m
|
|
WHERE m.chat_session_id = cs.id
|
|
ORDER BY m.created_at DESC
|
|
LIMIT 1
|
|
) lm ON true
|
|
WHERE cs.workspace_id = $1 AND cs.creator_id = $2 AND cs.status = 'active'
|
|
ORDER BY (cs.pinned_at IS NOT NULL) DESC, cs.pinned_at DESC, COALESCE(lm.created_at, cs.updated_at) DESC;
|
|
|
|
-- name: ListAllChatSessionsByCreator :many
|
|
-- Unlike ListChatSessionsByCreator this returns archived sessions too (for the
|
|
-- "Archived" view), so unread must be forced to 0 for archived rows: archiving
|
|
-- deliberately does NOT advance last_read_at (so unarchive can restore the true
|
|
-- unread state), but an archived session is read-only and hidden from history,
|
|
-- so any residual unread is uncleanable and must not light up any badge. Gating
|
|
-- on status here is the single source of truth for all unread surfaces (FAB,
|
|
-- sidebar Chat tab, chat-window header) — see MUL-4360.
|
|
SELECT cs.*,
|
|
CASE WHEN cs.status = 'archived' THEN 0
|
|
ELSE (SELECT count(*) FROM chat_message m
|
|
WHERE m.chat_session_id = cs.id
|
|
AND m.role = 'assistant'
|
|
AND m.created_at > cs.last_read_at)
|
|
END::int AS unread_count,
|
|
COALESCE(lm.content, '') AS last_message_content,
|
|
COALESCE(lm.role, '') AS last_message_role,
|
|
lm.created_at AS last_message_at,
|
|
lm.failure_reason AS last_message_failure_reason,
|
|
COALESCE(lm.message_kind, '') AS last_message_kind
|
|
FROM chat_session cs
|
|
LEFT JOIN LATERAL (
|
|
SELECT content, role, created_at, failure_reason, message_kind
|
|
FROM chat_message m
|
|
WHERE m.chat_session_id = cs.id
|
|
ORDER BY m.created_at DESC
|
|
LIMIT 1
|
|
) lm ON true
|
|
WHERE cs.workspace_id = $1 AND cs.creator_id = $2
|
|
ORDER BY (cs.pinned_at IS NOT NULL) DESC, cs.pinned_at DESC, COALESCE(lm.created_at, cs.updated_at) DESC;
|
|
|
|
-- name: UpdateChatSessionTitle :one
|
|
UPDATE chat_session SET title = $2, updated_at = now()
|
|
WHERE id = $1
|
|
RETURNING *;
|
|
|
|
-- name: UpdateChatSessionProject :one
|
|
-- Project context is user-editable session metadata. Do not touch updated_at:
|
|
-- changing context is not conversation activity and must not reorder history.
|
|
UPDATE chat_session
|
|
SET project_id = sqlc.narg('project_id')
|
|
WHERE id = sqlc.arg('id') AND workspace_id = sqlc.arg('workspace_id')
|
|
RETURNING *;
|
|
|
|
-- name: UpdateChatSessionTitleIfCurrent :one
|
|
-- Compare-and-swap the title: only overwrite it when it still equals the
|
|
-- value the caller observed (@expected_title). This is the idempotency /
|
|
-- no-clobber guard behind LLM auto-titling (MUL-4295): the async generator
|
|
-- captures the session's current (default/original) title before calling the
|
|
-- model, and this write lands only if a manual rename or a competing writer
|
|
-- has not changed the title in the meantime. A mismatch returns pgx.ErrNoRows
|
|
-- (zero rows updated), which the caller treats as "someone renamed it — leave
|
|
-- it alone", NOT as an error.
|
|
UPDATE chat_session SET title = @new_title, updated_at = now()
|
|
WHERE id = @id AND title = @expected_title
|
|
RETURNING *;
|
|
|
|
-- name: SetChatSessionPinned :one
|
|
-- Pin/unpin a chat. Deliberately does NOT touch updated_at: pinning is a
|
|
-- list-ordering preference, not activity, so it must not bump the session's
|
|
-- last-activity sort key (which would make an unpinned chat jump the list).
|
|
-- pinned = true stamps pinned_at only when it was NULL, so re-pinning keeps
|
|
-- the original pin order; pinned = false clears it.
|
|
UPDATE chat_session
|
|
SET pinned_at = CASE WHEN @pinned::bool THEN COALESCE(pinned_at, now()) ELSE NULL END
|
|
WHERE id = $1
|
|
RETURNING *;
|
|
|
|
-- name: SetChatSessionArchived :one
|
|
-- Archive/unarchive a chat session by flipping status between 'active' and
|
|
-- 'archived'. Bumps updated_at so the row re-sorts on the receiving list. The
|
|
-- send-message path refuses archived sessions (see SendChatMessage), so the
|
|
-- conversation is effectively read-only until it is unarchived.
|
|
UPDATE chat_session
|
|
SET status = CASE WHEN @archived::bool THEN 'archived' ELSE 'active' END,
|
|
updated_at = now()
|
|
WHERE id = $1
|
|
RETURNING *;
|
|
|
|
-- name: UpdateChatSessionSession :exec
|
|
-- Updates the resume pointer for a chat session. Empty/NULL inputs are
|
|
-- ignored via COALESCE so a task that completes without a session_id (e.g.
|
|
-- the agent crashed before establishing one) cannot wipe out a previously
|
|
-- recorded resume pointer. This makes the chat memory robust against
|
|
-- intermittent agent failures.
|
|
UPDATE chat_session
|
|
SET session_id = COALESCE(sqlc.narg('session_id'), session_id),
|
|
work_dir = COALESCE(sqlc.narg('work_dir'), work_dir),
|
|
runtime_id = COALESCE(sqlc.narg('runtime_id'), runtime_id),
|
|
updated_at = now()
|
|
WHERE id = sqlc.arg('id');
|
|
|
|
-- name: LockChatSessionForDelete :one
|
|
-- Acquires an exclusive (FOR UPDATE) row lock on chat_session(id). Used by
|
|
-- the delete path so that a concurrent SendChatMessage cannot enqueue a new
|
|
-- agent_task_queue row referencing this session between our cancel and
|
|
-- delete steps. The FK from agent_task_queue.chat_session_id takes a
|
|
-- KEY SHARE lock on the parent row during INSERT validation, which
|
|
-- conflicts with FOR UPDATE — concurrent inserts block here and then fail
|
|
-- their FK check after we commit the delete.
|
|
SELECT id FROM chat_session
|
|
WHERE id = $1
|
|
FOR UPDATE;
|
|
|
|
-- name: LockChatSessionForRuntimeBind :one
|
|
-- Acquires an exclusive (FOR UPDATE) row lock on chat_session(id), serialising
|
|
-- "which runtime does this session execute on" against "enqueue the next task".
|
|
--
|
|
-- Both SendDirectChatMessage and the agent-builder runtime switch take this lock
|
|
-- for their whole transaction. Without it the two are a read-then-write race: a
|
|
-- send reads the carrier agent's runtime_id, the switch then passes its
|
|
-- pending-task check and rebinds the carrier, and the send finally inserts a task
|
|
-- still stamped with the pre-switch runtime — so the user is told the switch
|
|
-- succeeded while their message runs on the old runtime (MUL-5163).
|
|
--
|
|
-- The lock alone is not sufficient: the send path must also re-read the agent
|
|
-- INSIDE the locked transaction, because a send blocked at INSERT would otherwise
|
|
-- resume and write the runtime_id it read before blocking.
|
|
--
|
|
-- Same row and same lock mode as LockChatSessionForDelete, and both take it as
|
|
-- their first statement, so the delete path and this one cannot deadlock.
|
|
SELECT id FROM chat_session
|
|
WHERE id = $1
|
|
FOR UPDATE;
|
|
|
|
-- name: DeleteChatSession :exec
|
|
-- Hard delete. chat_message rows cascade via FK ON DELETE CASCADE; the
|
|
-- chat_session_id on agent_task_queue is set NULL by FK so completed/failed
|
|
-- task history survives the session being removed. Callers MUST run inside
|
|
-- the same transaction that holds LockChatSessionForDelete and that has
|
|
-- already cancelled any in-flight tasks (see CancelAgentTasksByChatSession)
|
|
-- so the daemon does not keep running work whose result has nowhere to
|
|
-- land. workspace_id in the WHERE clause is a SQL-layer tenant guard; see
|
|
-- DeleteIssue.
|
|
DELETE FROM chat_session WHERE id = $1 AND workspace_id = $2;
|
|
|
|
-- name: TouchChatSession :exec
|
|
UPDATE chat_session SET updated_at = now()
|
|
WHERE id = $1;
|
|
|
|
-- name: CreateChatMessage :one
|
|
-- message_kind defaults to 'message' via COALESCE so every existing caller
|
|
-- (which omits it) keeps writing ordinary messages; the empty-reply path passes
|
|
-- 'no_response' to mark a visible turn with no text output (MUL-4351).
|
|
INSERT INTO chat_message (
|
|
chat_session_id, role, content, task_id, failure_reason, elapsed_ms,
|
|
message_kind, channel_media_pending_until, channel_ingested
|
|
)
|
|
VALUES (
|
|
$1, $2, $3, sqlc.narg(task_id), sqlc.narg(failure_reason), sqlc.narg(elapsed_ms),
|
|
COALESCE(sqlc.narg(message_kind)::text, 'message'),
|
|
-- The media deadline is DB-clock time: every consumer compares it against
|
|
-- SQL now() (GetChannelMediaPendingUntil, the deferred promote, the
|
|
-- trailing-message guard), so the writer must use the same clock. The
|
|
-- caller passes a relative budget in seconds; an application-clock
|
|
-- timestamp here would let a skewed app node shrink or stretch the
|
|
-- fallback window.
|
|
CASE WHEN sqlc.narg(channel_media_pending_secs)::float8 IS NULL THEN NULL
|
|
ELSE now() + make_interval(secs => sqlc.narg(channel_media_pending_secs)::float8) END,
|
|
COALESCE(sqlc.narg(channel_ingested)::boolean, FALSE)
|
|
)
|
|
RETURNING *;
|
|
|
|
-- name: TaskHasChannelIngestedMessages :one
|
|
-- Immutable channel provenance for a task's user-message input batch:
|
|
-- channel_ingested is stamped inside the channel append transaction and never
|
|
-- mutated afterwards, so it survives session archiving and installation
|
|
-- rebinds that delete the channel_chat_session_binding row. Callers pass the
|
|
-- batch OWNER id (chat_input_task_id, which auto-retry clones inherit), not
|
|
-- necessarily the task's own id. The cancel restore-delete and the
|
|
-- empty-completion silent-drop both gate on this — a channel sender has no
|
|
-- Multica composer for a restored draft, and the no_response fallback body
|
|
-- must never be pushed to an external channel.
|
|
SELECT EXISTS (
|
|
SELECT 1 FROM chat_message
|
|
WHERE task_id = $1
|
|
AND role = 'user'
|
|
AND channel_ingested
|
|
) AS channel_ingested;
|
|
|
|
-- name: GetChannelMediaPendingUntil :one
|
|
-- The latest unexpired media deadline gates a channel task. Using a durable
|
|
-- task fire_at means a process restart still produces the placeholder fallback.
|
|
SELECT channel_media_pending_until
|
|
FROM chat_message
|
|
WHERE chat_session_id = $1
|
|
AND role = 'user'
|
|
AND channel_media_pending_until > now()
|
|
ORDER BY channel_media_pending_until DESC
|
|
LIMIT 1;
|
|
|
|
-- name: ClearChatMessageChannelMediaPending :exec
|
|
UPDATE chat_message
|
|
SET channel_media_pending_until = NULL
|
|
WHERE id = $1 AND chat_session_id = $2;
|
|
|
|
-- name: LinkChatMessageToTask :exec
|
|
UPDATE chat_message
|
|
SET task_id = $2
|
|
WHERE id = $1 AND role = 'user';
|
|
|
|
-- name: LinkUnownedChannelChatMessagesToTask :exec
|
|
-- Seals the trailing channel-message batch to its task. The task row and these
|
|
-- links are committed together, so an older in-flight task cannot absorb a
|
|
-- newer media message and a later assistant row cannot hide that message.
|
|
UPDATE chat_message AS message
|
|
SET task_id = @task_id
|
|
WHERE message.chat_session_id = @chat_session_id
|
|
AND message.role = 'user'
|
|
AND message.task_id IS NULL
|
|
AND NOT EXISTS (
|
|
SELECT 1
|
|
FROM chat_message AS prior
|
|
WHERE prior.chat_session_id = @chat_session_id
|
|
AND prior.role != 'user'
|
|
AND (prior.created_at, prior.id) > (message.created_at, message.id)
|
|
);
|
|
|
|
-- name: DeferChatTaskForSealedPendingMedia :one
|
|
-- Closes the enqueue-vs-append race: under READ COMMITTED a media message can
|
|
-- commit between GetChannelMediaPendingUntil and the batch seal above, landing
|
|
-- an unexpired media marker inside a task the deadline read decided was
|
|
-- 'queued'. Re-derive the deferral from the sealed batch itself, in the same
|
|
-- transaction, so a task is never claimable while its own input still has an
|
|
-- unexpired marker. No row (ErrNoRows) means no correction was needed.
|
|
UPDATE agent_task_queue AS task
|
|
SET status = 'deferred', fire_at = pending.max_until
|
|
FROM (
|
|
SELECT max(message.channel_media_pending_until) AS max_until
|
|
FROM chat_message AS message
|
|
WHERE message.task_id = @task_id
|
|
AND message.role = 'user'
|
|
AND message.channel_media_pending_until > now()
|
|
) AS pending
|
|
WHERE task.id = @task_id
|
|
AND pending.max_until IS NOT NULL
|
|
AND (task.fire_at IS NULL OR task.fire_at < pending.max_until)
|
|
RETURNING task.*;
|
|
|
|
-- name: DeleteUserChatMessageByTask :one
|
|
DELETE FROM chat_message
|
|
WHERE task_id = $1 AND role = 'user'
|
|
RETURNING *;
|
|
|
|
-- name: ListChatMessages :many
|
|
SELECT * FROM chat_message
|
|
WHERE chat_session_id = $1
|
|
ORDER BY created_at ASC, id ASC;
|
|
|
|
-- name: ListChatInputMessages :many
|
|
-- Loads the immutable user-message input batch owned by a direct-chat task.
|
|
-- The caller passes the task's chat_input_task_id (itself for an original send,
|
|
-- the root task for an auto-retry child), so a claim reads exactly the messages
|
|
-- the user sent for this turn — and never absorbs a message that arrived after
|
|
-- the batch was sealed, no matter what the assistant wrote or when. Only used
|
|
-- for new task-owned direct-chat tasks; legacy/channel (chat_input_task_id
|
|
-- NULL) tasks keep using ListChatMessages + trailingUserMessages.
|
|
SELECT * FROM chat_message
|
|
WHERE task_id = $1 AND role = 'user'
|
|
ORDER BY created_at ASC, id ASC;
|
|
|
|
-- name: ListChatMessagesPage :many
|
|
SELECT * FROM chat_message
|
|
WHERE chat_session_id = $1
|
|
AND (
|
|
sqlc.narg('before_created_at')::timestamptz IS NULL
|
|
OR (created_at, id) < (sqlc.narg('before_created_at')::timestamptz, sqlc.narg('before_id')::uuid)
|
|
)
|
|
ORDER BY created_at DESC, id DESC
|
|
LIMIT $2;
|
|
|
|
-- name: GetChatMessage :one
|
|
SELECT * FROM chat_message
|
|
WHERE id = $1;
|
|
|
|
-- name: CreateChatTask :one
|
|
-- The chat sender (initiator) is a direct_human originator and accountable;
|
|
-- attribution provenance is stamped so this path is not a NULL-source enqueue
|
|
-- bypass (MUL-4302 §2).
|
|
INSERT INTO agent_task_queue (
|
|
agent_id, runtime_id, issue_id, status, priority, chat_session_id,
|
|
initiator_user_id, originator_user_id, accountable_user_id, force_fresh_session, runtime_mcp_overlay,
|
|
runtime_connected_apps, originator_source, trigger_evidence_kind, trigger_evidence_ref_id,
|
|
fire_at
|
|
)
|
|
VALUES (
|
|
$1, $2, NULL,
|
|
CASE WHEN sqlc.narg('fire_at')::timestamptz IS NULL THEN 'queued' ELSE 'deferred' END,
|
|
$3, $4, $5,
|
|
sqlc.narg(originator_user_id),
|
|
sqlc.narg(accountable_user_id),
|
|
COALESCE(sqlc.narg('force_fresh_session')::boolean, FALSE),
|
|
sqlc.narg(runtime_mcp_overlay),
|
|
sqlc.narg(runtime_connected_apps),
|
|
sqlc.narg(originator_source),
|
|
sqlc.narg(trigger_evidence_kind),
|
|
sqlc.narg(trigger_evidence_ref_id),
|
|
sqlc.narg('fire_at')::timestamptz
|
|
)
|
|
RETURNING *;
|
|
|
|
-- name: PromoteChannelChatTasksIfMediaReady :many
|
|
-- Media completion may race with the 3s run batcher. Promote every original
|
|
-- channel task waiting for this session only after all unexpired media markers
|
|
-- are gone; retry/escalation/direct-chat deferred tasks are excluded.
|
|
UPDATE agent_task_queue AS task
|
|
SET status = 'queued', fire_at = NULL
|
|
WHERE task.chat_session_id = @chat_session_id
|
|
AND task.status = 'deferred'
|
|
AND task.issue_id IS NULL
|
|
AND task.parent_task_id IS NULL
|
|
AND task.escalation_for_task_id IS NULL
|
|
AND NOT EXISTS (
|
|
SELECT 1
|
|
FROM chat_message AS message
|
|
WHERE message.chat_session_id = @chat_session_id
|
|
AND message.role = 'user'
|
|
AND message.channel_media_pending_until > now()
|
|
)
|
|
RETURNING task.*;
|
|
|
|
-- name: SetChatTaskInputOwnerSelf :one
|
|
-- Stamps a freshly-created direct-chat task as the owner of its own input batch
|
|
-- (chat_input_task_id = id), so a later claim loads exactly the user messages
|
|
-- tagged with this task id (ListChatInputMessages) rather than scanning trailing
|
|
-- history. Runs in the same transaction as CreateChatTask + message ownership:
|
|
-- direct-send inserts one owned message, while channel enqueue seals its
|
|
-- trailing unowned batch. Legacy tasks keep chat_input_task_id NULL and retain
|
|
-- the trailing-history fallback during rolling deploys.
|
|
UPDATE agent_task_queue
|
|
SET chat_input_task_id = id
|
|
WHERE id = $1
|
|
RETURNING *;
|
|
|
|
-- name: GetLastChatTaskSession :one
|
|
-- Returns the most recent task in this chat session that managed to record a
|
|
-- session_id. Includes both completed and failed tasks: even a failed task
|
|
-- may have established a real agent session before failing, and we'd rather
|
|
-- resume there than start over and lose conversation memory. Used as a
|
|
-- fallback when chat_session.session_id is NULL. Resume-unsafe failures are
|
|
-- excluded because replaying those sessions deterministically reproduces the
|
|
-- same terminal state. Keep this list in sync with resumeUnsafeFailureReason
|
|
-- and GetLastTaskSession.
|
|
SELECT session_id, work_dir, runtime_id FROM agent_task_queue
|
|
WHERE chat_session_id = $1
|
|
AND (
|
|
status = 'completed'
|
|
OR (
|
|
status = 'failed'
|
|
AND COALESCE(failure_reason, '') NOT IN ('iteration_limit', 'agent_fallback_message', 'api_invalid_request', 'codex_semantic_inactivity', 'agent_error.context_overflow')
|
|
AND NOT (COALESCE(error, '') ILIKE '%400%' AND COALESCE(error, '') ILIKE '%invalid_request_error%')
|
|
)
|
|
)
|
|
AND session_id IS NOT NULL
|
|
ORDER BY completed_at DESC
|
|
LIMIT 1;
|
|
|
|
-- name: GetPendingChatTask :one
|
|
-- Returns the most recent in-flight task for a chat session, if any.
|
|
-- Used by the frontend to recover pending state after refresh / reopen.
|
|
-- created_at is the anchor for the chat StatusPill timer (it computes
|
|
-- elapsed = now - task.created_at), so the pill survives refresh / reopen
|
|
-- without "resetting to 0s".
|
|
SELECT id, status, created_at FROM agent_task_queue
|
|
WHERE chat_session_id = $1 AND status IN ('queued', 'dispatched', 'running', 'waiting_local_directory')
|
|
ORDER BY created_at DESC
|
|
LIMIT 1;
|
|
|
|
-- name: ListPendingChatTasksByCreator :many
|
|
-- Aggregate view of all in-flight chat tasks owned by a given creator in a
|
|
-- workspace. Drives the FAB's "running" indicator when the chat window is
|
|
-- closed and no single session's query is active.
|
|
--
|
|
-- Returns cs.agent_id so the handler can filter tasks belonging to private
|
|
-- agents the caller has lost access to using the already-loaded `allowed`
|
|
-- set — no second ListAllChatSessionsByCreator scan on the hot path.
|
|
--
|
|
-- atq.chat_session_id IS NOT NULL is redundant given the JOIN, but stated
|
|
-- explicitly so the planner can prove the query predicate is a subset of the
|
|
-- idx_agent_task_queue_chat_pending_v2 partial-index predicate and use it.
|
|
SELECT atq.id AS task_id, atq.status, atq.chat_session_id, cs.agent_id
|
|
FROM agent_task_queue atq
|
|
JOIN chat_session cs ON cs.id = atq.chat_session_id
|
|
WHERE atq.chat_session_id IS NOT NULL
|
|
AND atq.status IN ('queued', 'dispatched', 'running', 'waiting_local_directory')
|
|
AND cs.workspace_id = $1
|
|
AND cs.creator_id = $2
|
|
ORDER BY atq.created_at DESC;
|
|
|
|
-- name: HasPendingChatTasksByCreator :one
|
|
-- Boolean fast-path for the FAB's "running" indicator. Returns a single
|
|
-- EXISTS row instead of the full task list, so the planner can stop at the
|
|
-- first matching in-flight task (LIMIT 1 semantics via EXISTS).
|
|
--
|
|
-- Permission filtering is baked into the query: agent_id = ANY($3) restricts
|
|
-- the result to the agents the caller may currently see, so a member who lost
|
|
-- access to a private agent never gets a true from a task they can no longer
|
|
-- reach. The handler must pass its resolved accessible-agent id set as $3;
|
|
-- an empty array yields false.
|
|
SELECT EXISTS (
|
|
SELECT 1
|
|
FROM agent_task_queue atq
|
|
JOIN chat_session cs ON cs.id = atq.chat_session_id
|
|
WHERE atq.chat_session_id IS NOT NULL
|
|
AND atq.status IN ('queued', 'dispatched', 'running', 'waiting_local_directory')
|
|
AND cs.workspace_id = sqlc.arg(workspace_id)
|
|
AND cs.creator_id = sqlc.arg(creator_id)
|
|
AND cs.agent_id = ANY(sqlc.arg(agent_ids)::uuid[])
|
|
) AS has_pending;
|
|
|
|
-- name: MarkChatSessionRead :exec
|
|
-- Advances the read cursor to now, dropping the session's unread_count to 0.
|
|
UPDATE chat_session SET last_read_at = now()
|
|
WHERE id = $1;
|
|
|
|
-- name: GetMostRecentUserChatMessage :one
|
|
-- Returns the most recent role='user' message in a session. Used by the
|
|
-- Lark `/issue` command parser: when the user types `/issue` with no
|
|
-- title, the spec falls back to "use the previous user message as the
|
|
-- title". Bot replies (role='assistant') are excluded — only human
|
|
-- input qualifies as a fallback title source.
|
|
SELECT * FROM chat_message
|
|
WHERE chat_session_id = $1 AND role = 'user'
|
|
ORDER BY created_at DESC
|
|
LIMIT 1;
|
|
|
|
-- name: ChatSessionHasUserMessage :one
|
|
-- Reports whether a session has any human (role='user') message yet. Used to
|
|
-- scope the is_agent_intro self-introduction prompt to the very first,
|
|
-- server-driven turn: an intro session starts with zero user messages, so the
|
|
-- opening run gets the "introduce yourself" prompt. Once the creator replies,
|
|
-- later turns in the same session must fall back to the normal reply prompt
|
|
-- instead of repeating the introduction every turn (MUL-4259).
|
|
SELECT EXISTS (
|
|
SELECT 1 FROM chat_message
|
|
WHERE chat_session_id = $1 AND role = 'user'
|
|
) AS has_user_message;
|
|
|
|
-- name: CreateChatDraftRestore :one
|
|
-- Persists the deferred-cancellation draft restore (#5219) in the same tx
|
|
-- that deletes the triggering user message: the chat:cancel_finalized
|
|
-- broadcast is best-effort, so an offline client recovers the draft from
|
|
-- this row instead. id is the deleted message's id.
|
|
INSERT INTO chat_draft_restore (id, chat_session_id, task_id, content, attachment_ids)
|
|
VALUES ($1, $2, $3, $4, $5)
|
|
RETURNING *;
|
|
|
|
-- name: ListChatDraftRestoresBySession :many
|
|
SELECT * FROM chat_draft_restore
|
|
WHERE chat_session_id = $1
|
|
ORDER BY created_at ASC;
|
|
|
|
-- name: DeleteChatDraftRestore :execrows
|
|
-- Idempotent consume: deleting an already-consumed restore matches no row.
|
|
DELETE FROM chat_draft_restore
|
|
WHERE id = $1 AND chat_session_id = $2;
|
|
|
|
-- name: DeleteChatDraftRestoresBySession :exec
|
|
-- chat_draft_restore carries no chat_session FK (MUL-3515), so DeleteChatSession
|
|
-- prunes its pending restores in the same tx that deletes the session.
|
|
DELETE FROM chat_draft_restore
|
|
WHERE chat_session_id = $1;
|
|
|
|
-- The chat_session row lock is the mutual-exclusion protocol between the
|
|
-- draft-restore writer (FinalizeDeferredCancelledChat) and every deleter of a
|
|
-- chat_session or one of its cascade parents. Without an FK, an INSERT into
|
|
-- chat_draft_restore takes no lock on its session, so a prune-then-delete would
|
|
-- otherwise miss a restore committed after its snapshot and strand it — with the
|
|
-- user's prompt in it — forever (#5219).
|
|
--
|
|
-- The contract, held by all five paths below:
|
|
-- deleter: lock the sessions FOR UPDATE -> prune restores -> delete the parent
|
|
-- finalizer: lock the session FOR UPDATE -> insert the restore
|
|
-- Whoever locks first wins: a finalizer that got there first commits its row
|
|
-- before the deleter's prune statement takes its snapshot, so the prune sweeps
|
|
-- it; a deleter that got there first leaves no session for the finalizer to
|
|
-- lock, so it never inserts.
|
|
--
|
|
-- Lock order is chat_session -> agent_task_queue everywhere (the finalizer locks
|
|
-- the session before claiming its task) so the deleters' cascade into
|
|
-- agent_task_queue cannot deadlock against it.
|
|
--
|
|
-- The single-session delete path needs no new query: it already holds
|
|
-- LockChatSessionForDelete (same FOR UPDATE row lock) across its prune.
|
|
|
|
-- name: LockChatSessionForTask :one
|
|
-- The finalizer's half of the protocol. No rows means the session is already
|
|
-- gone (its cascade NULLs agent_task_queue.chat_session_id), so there is nothing
|
|
-- to lock and nothing to restore into.
|
|
SELECT cs.id
|
|
FROM agent_task_queue t
|
|
JOIN chat_session cs ON cs.id = t.chat_session_id
|
|
WHERE t.id = $1
|
|
FOR UPDATE OF cs;
|
|
|
|
-- name: LockChatSessionsByWorkspace :many
|
|
-- ORDER BY id: a stable lock order keeps two concurrent deleters from
|
|
-- deadlocking against each other.
|
|
SELECT id FROM chat_session
|
|
WHERE workspace_id = $1
|
|
ORDER BY id
|
|
FOR UPDATE;
|
|
|
|
-- name: LockChatSessionsByArchivedRuntimeAgents :many
|
|
SELECT cs.id FROM chat_session cs
|
|
JOIN agent a ON a.id = cs.agent_id
|
|
WHERE a.runtime_id = $1 AND a.archived_at IS NOT NULL
|
|
ORDER BY cs.id
|
|
FOR UPDATE OF cs;
|
|
|
|
-- name: LockChatSessionsBySystemRuntimeAgents :many
|
|
SELECT cs.id FROM chat_session cs
|
|
JOIN agent a ON a.id = cs.agent_id
|
|
WHERE a.runtime_id = $1 AND a.kind = 'system'
|
|
ORDER BY cs.id
|
|
FOR UPDATE OF cs;
|
|
|
|
-- name: DeleteChatDraftRestoresByArchivedRuntimeAgents :exec
|
|
-- chat_session cascades from agent, so hard-deleting a runtime's archived agents
|
|
-- silently drops their sessions — and, without an FK, would strand the pending
|
|
-- restores (which still hold the user's prompt text) forever. Prune them in the
|
|
-- same tx, BEFORE the agent rows go: the join below needs them. Mirrors
|
|
-- DeleteChannelInstallationsByArchivedRuntimeAgents.
|
|
DELETE FROM chat_draft_restore
|
|
WHERE chat_session_id IN (
|
|
SELECT cs.id FROM chat_session cs
|
|
JOIN agent a ON a.id = cs.agent_id
|
|
WHERE a.runtime_id = $1 AND a.archived_at IS NOT NULL
|
|
);
|
|
|
|
-- name: DeleteChatDraftRestoresBySystemRuntimeAgents :exec
|
|
-- Same cascade, for the system agents a runtime teardown also hard-deletes
|
|
-- (DeleteSystemAgentsByRuntime). Split from the archived-agent prune because the
|
|
-- runtime-profile teardown deletes only archived agents: pruning system-agent
|
|
-- sessions there would destroy restores whose session survives.
|
|
DELETE FROM chat_draft_restore
|
|
WHERE chat_session_id IN (
|
|
SELECT cs.id FROM chat_session cs
|
|
JOIN agent a ON a.id = cs.agent_id
|
|
WHERE a.runtime_id = $1 AND a.kind = 'system'
|
|
);
|