Commit Graph

778 Commits

Author SHA1 Message Date
Lambda
2cbd72e800 merge: bring origin/main into codex/mika-issue-first-onboarding
Every conflict was i18n or a file this branch had already rewritten.

Helper-flow templates (create-agent-guide-issue, helper-instructions,
helper-starter-prompts) came back as modify/delete: this branch removed
the flow, main only retranslated them. Deleted. Same call for
install-runtime-issue.ts and welcome-after-onboarding.test.tsx, which
keep the Mika step over main's Helper step.

The locale bundles were merged per key rather than per side, because
main's e3bf9ccd4 renamed the filed unit (issue -> 任务 / タスク / 태스크)
and gave the execution run a distinct form (zh `task`, ja 作業, ko 작업).
Taking either side wholesale would have reverted one of the two changes,
so this branch's newly written strings were re-rendered in main's new
vocabulary instead — including the strings that merged cleanly and so
never surfaced as conflicts.

Key sets follow the auto-merged en bundle: keys main added under the
removed Helper flow are dropped, keys main genuinely added are taken,
and the renamed pickers.*.clear_action is gone. Parity is exact.

Also moves the agents instructions-tab onto the role-named type scale.
apps/web/app/type-scale.test.ts had been failing since this branch added
the system-instructions panel with text-sm/text-xs.

Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 16:43:23 +08:00
Multica Eve
86f255658f MUL-5730 feat(server): expose chat audience without bloating prompts (#6428)
Tell a chat agent whether it is answering in a shared room or a 1:1, without
splitting the cached brief prefix: the room shape rides the existing channel
binding row to the claim response, and is stated once per turn in the chat
prompt rather than in the static runtime brief.

Originally contributed as #6390; re-rolled to move the audience fact to the
per-turn prompt and to minimize the copy it renders.

Co-authored-by: Seacen Zhao <xichangzhao@outlook.com>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 16:12:17 +08:00
Bohan Jiang
c1bd4f995c fix(agent): fail a run whose provider session ran out of context (MUL-5739) (#6422)
Detect provider context exhaustion from Claude Code's structured
terminal_reason rather than from its error prose, so a saturated session is
retired and the next task cold-starts instead of resuming a conversation
that can no longer answer.

Refs #6402 — does not close it: the captured 2.1.220/2.1.221 frames report
is_error alongside terminal_reason, so the specific `completed` escape the
issue reports is not reproduced. The structured check is the load-bearing
fix and is proven present on the reported CLI version; the composite text
predicate is a bounded backstop for uncaptured shapes.
2026-08-05 15:41:15 +08:00
Multica Eve
b3fa17f0e0 fix(chat): preserve follow-up transcript order (#6419)
Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 15:39:55 +08:00
Multica Eve
73283556d7 MUL-5750: fix(chat): keep idle sends out of follow-up queue (#6418)
* fix(chat): keep idle sends out of follow-up queue

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): preserve the positional queue head

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): polish deferred queue states

Co-authored-by: multica-agent <github@multica.ai>

* fix(migrations): resolve pending index prefix collision

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 15:06:42 +08:00
Lambda
3879564fd7 feat(mika): use the unicorn emoji as Mika's placeholder avatar
Mika shipped with a hand-rolled data-URI SVG — a sparkle glyph on a dark
rounded square — which only that one constant knew how to produce. Agents
already have an emoji avatar convention (`emoji:` marker in avatar_url,
owned by agentEmojiAvatarPrefix on the server and parseAvatarEmoji on the
client), so this reuses it instead: ActorAvatar renders the emoji as text and
no surface needs to special-case her.

The two cards that stand in for Mika before the agent row exists move with
it — the onboarding intro card and the Runtimes "Start with Mika" card. Both
were drawing the same sparkle mark, and leaving them would mean a member sees
one face during onboarding and a different one the moment Mika is created.
They now share MIKA_PLACEHOLDER_EMOJI, which carries a pointer to the server
constant so the two cannot drift apart silently.

The dark square went with the sparkle: it was built to frame a white line-art
glyph, and an emoji on it reads badly. These use bg-muted, matching how
ActorAvatar already frames an emoji avatar.

Placeholder until Mika has real artwork.

Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 14:04:19 +08:00
TeAmo
a4dc54f2ed MUL-5570 / MUL-5493: feat(chat): add a managed follow-up queue (#6211)
* feat(chat): queue follow-up messages during active runs

* fix(chat): preserve queued message compatibility

* fix(chat): harden queued task consistency

* fix(chat): protect legacy and mobile queue state

* fix(mobile): refresh stop handler on task promotion

* fix(mobile): preserve queued chat state

* fix(chat): hide queued prompts from legacy clients

* test(chat): read queued channel input from paged transcript

* fix(chat): protect legacy and mobile queue state

* fix(mobile): refresh stop handler on task promotion

* fix(mobile): preserve queued chat state

* fix(chat): hide queued prompts from legacy clients

* test(chat): read queued channel input from paged transcript

* fix(chat): hide queued inputs from paged transcripts

* test: prove queued messages do not consume cursor pages

---------

Co-authored-by: TeAmo <liu.junhui3@iwhalecloud.com>
2026-08-05 13:15:57 +08:00
Multica Eve
d60bc63f9a Revert "MUL-5516: fix(issues): re-trigger blocked issues when resumed (#6155)" (#6377)
This reverts commit 738217c275.

Re-firing an assigned agent when an Issue leaves `blocked` is not the
product behaviour we want: `blocked` is a human-attention state, and
clearing it is not by itself a signal to start work. Reverted on product
grounds ahead of the v0.4.18 release, not because of a defect.

Restores the pre-#6155 behaviour exactly: only `backlog` -> active
promotes a parked Issue into a run, and the write/preview trigger probes
go back to being separate.

Requested by Bohan on MUL-5716.

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 17:53:03 +08:00
Naiyuan Qing
9fb46adc43 fix(agent-builder): serialise draft autosave against session delete (MUL-5642) (#6376)
SaveAgentBuilderDraft read the chat session, then upserted agent_builder_draft
as a separate statement, with no lock between them. DeleteChatSession takes
LockChatSessionForDelete for its whole transaction, so a save could pass its
checks, block on nothing, and land its INSERT after the delete committed.
agent_builder_draft carries no chat_session FK (repo rule), so nothing rejected
that write.

The surviving row held the configuration the user had just confirmed discarding.
It is invisible to the UI — the drafts list joins through chat_session — and no
prune can reach it: DeleteAgentBuilderDraft and the runtime teardown both key off
a session that no longer exists, leaving only the workspace teardown. The client
autosaves on an 800ms debounce and the conversation is addressable by URL, so a
second tab can autosave at any moment while this one discards.

Add LockChatSessionForDraftWrite, the same row and lock mode the delete and
runtime-bind paths take, and run the upsert in a transaction that acquires it
first and re-reads the session under it. Existence and status are the only two
things a concurrent writer can change, and both are now decided inside the lock;
workspace, creator and carrier are immutable for a session and stay on the
cheap unlocked read. Either ordering is now correct: the save commits first and
the delete prunes it, or the delete commits first and the save returns 404.

The same lock closes the archive variant, where the last autosave after
"create agent" could write a draft onto an already read-only session.

Both regression tests drive the interleaving deterministically — hold the
session row, prove the save blocks, then commit — and fail on the pre-fix
handler with a 204 that writes the orphan.

Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 17:52:33 +08:00
Bohan Jiang
eea687d461 MUL-5686: fix(chat): resume the cancelled turn's session instead of starting cold (#6352)
* fix(chat): resume the cancelled turn's session instead of starting cold

Stopping a chat turn the agent had already begun answering, then sending
the next message, produced a reply with no memory of the conversation.

The cancelled turn's provider session is real and recorded — the daemon
pins it onto the task row mid-flight and cancellation keeps it there —
but nothing could hand it back:

  - GetLastChatTaskSession / GetLastTaskSession only considered
    'completed' and 'failed' rows, so a cancelled row's session was
    invisible to resume resolution;
  - chat_session.session_id is written only by CompleteTask / FailTask,
    and a cancelled task reaches neither: the daemon discards its result
    and sends a cancel-ack. On a chat whose first turn was cancelled the
    pointer therefore stayed NULL and the next turn started cold, since
    buildChatPrompt injects only the current user message.

Let cancelled rows into both resume lookups, and advance the chat-level
pointer at cancel time so a provider that mints a new session id per
resume does not rewind past the cancelled exchange. The mid-flight pin
may now also fill an EMPTY session slot on a just-cancelled row, which
is what a Codex pin waiting on its rollout needs when the cancel wins
the race; occupied slots and completed/failed rows stay untouchable.

Retired sessions, poisoned-failure filters and the rollout-present guard
are unchanged. A transcript killed mid-tool-call that the provider later
refuses is still caught by taskfailure.UnresumableHistory, which retires
the session and starts the next turn fresh.

Fixes #6340

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): close the two cancel/pin races on the chat resume pointer

Review found the pointer advance had the same shape of hole it was meant
to fix, on both sides of the cancel.

1. The status flip and the pointer advance were separate statements. In
   the gap the row is already `cancelled` while the pointer still names
   the PREVIOUS turn, so a follow-up the user had queued could be claimed
   there and resume the older session — and a failed pointer write only
   logged, still reporting the cancel as successful. Both now commit in
   one transaction and the error propagates.

2. The mid-flight pin may land AFTER the cancel (for Codex it waits for
   the rollout), so the cancel transaction sees no session to publish and
   the pin only fills the task row. On a chat that already had history
   the stale pointer kept shadowing it, which is precisely the case the
   pin change was added for.

Both paths now run one guarded statement,
AdvanceCancelledChatSessionPointer. It reads the task row itself rather
than trusting an in-memory copy, ignores anything that is not a cancelled
chat task, and refuses to move the pointer when a NEWER task on the chat
already recorded a session — so a straggler pin cannot drag the
conversation back onto the interrupted turn.

Regression tests, both failing before this commit: a cancel whose pointer
write is blocked must not expose `cancelled` to another connection, and a
late pin on an already-cancelled row must reach the next real claim. The
newer-turn guard is covered too.

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): make the pin atomic and restore the chat->task lock order

Second review round found the remaining half of the same class of bug.

1. PinTaskSession still committed the session onto the task row and
   advanced the chat pointer as two statements, and the pointer failure
   only logged while the endpoint still answered 204. A follow-up claimed
   in that window resumes the previous turn — the exact case the pin
   change was added for. Both writes now share one transaction and every
   failure is reported.

2. Adding the pointer write gave the cancel transaction two rows to hold,
   and it took them in the wrong order: agent_task_queue first, then
   chat_session. DeleteChatSession takes them the other way round, so the
   two could deadlock (40P01, and runInTx has no deadlock retry). The
   repo's documented order is chat_session -> agent_task_queue; both the
   cancel and the pin now open with LockChatSessionForTask, the same
   helper FinalizeDeferredCancelledChat uses. ErrNoRows there means a
   non-chat task or an already-deleted session — nothing to lock and
   nothing to advance.

Regression tests, all three failing before this commit (the concurrency
one with a real `deadlock detected`): the pin must not expose a session
on the task row while its pointer write is blocked; a cancel waiting on
the chat session must hold no lock on the task row (FOR UPDATE NOWAIT
probe); and cancel racing a chat delete must never come back with 40P01.

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): put every chat-task terminal write behind the same lock order

The cancel and pin paths took chat_session before agent_task_queue; the
terminal reports still took them the other way round, so the two could
deadlock (40P01) — reproduced deterministically as cancel-vs-complete.
Choosing per-path was never going to work: either every writer that holds
both rows agrees on an order or none of them are safe.

CompleteTask, FailTask and the cancelled-chat finalize now open with the
same LockChatSessionForTask the cancel, pin, DeleteChatSession and
FinalizeDeferredCancelledChat paths use, via a shared helper that
documents the invariant in one place.

Also de-flakes the concurrency tests this PR added. They held a row lock
and then read other rows on a FRESH pooled connection; the suite shares
one database, so a sibling package's DDL could queue an ACCESS EXCLUSIVE
request in between, park our later ACCESS SHARE request behind it, and
wedge the package until the 10-minute timeout (observed in a parallel
`go test ./internal/...` run). Those transactions now take their table
locks up front and read on the connection that already holds them, and
every racing call is bounded so a stall fails loudly instead of hanging.

Regression tests, all failing before this commit: complete and fail must
hold no lock on the task row while waiting for the chat session (FOR
UPDATE NOWAIT probe), and cancel-vs-complete plus pin-vs-fail must never
come back with 40P01.

Co-authored-by: multica-agent <github@multica.ai>

* test(chat): fail the race tests on any unexpected error, bound every call

Review nits on the concurrency tests.

Matching only *pgconn.PgError(40P01) let the pin side through: a pin that
loses a deadlock is reported by the HTTP handler as a plain 500 with no
PgError to unwrap, so the check skipped exactly the failure the test
exists for. Both racers now fail on anything except the one benign
outcome — whoever finalized the task first leaves the loser matching no
row (pgx.ErrNoRows).

The remaining racing calls still used an unbounded context.Background(),
which this PR had already claimed were bounded. They now share the same
raceTimeout as the rest.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 16:47:17 +08:00
Lambda
3bbc511f32 feat(onboarding): pick the runtime and the model from dropdowns
Nine runtimes as cards took three rows and left nowhere to put a second
decision. Two dropdowns fit both on one screen, and the model is a real
choice the card grid had no room for.

Both controls already existed on the agents surface — RuntimePicker and
ModelDropdown, the same pair create-agent-dialog uses — so this reuses them
rather than growing a parallel picker. ModelDropdown owns its own discovery,
grouping and unsupported-runtime states, and selecting a different runtime
clears the model because models are per-runtime.

The model now reaches the agent: POST /api/agents/mika takes an optional
model, CreateSystemUserAgent writes it to the column the agent table already
had, and empty still means "whatever the runtime defaults to" — which is what
every deployment without per-agent model support gets anyway.

Also drops the "No local runtime, or prefer a remote computer?" note. It said
in twenty-five words what the Skip button next to it says by existing, and it
appeared on all three phases.

currentUserId comes in as a prop rather than from the auth store: reading the
store inside the step broke six tests that render it without one, the same
coupling the header slot avoided.

Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 15:59:56 +08:00
Lambda
1c21379aca feat(onboarding): let quick-action chips carry the opening's examples
Chat now renders agent-suggested follow-up actions as buttons under a reply
(MUL-5149), and the onboarding kickoff qualifies for that suggestion pass
with no extra wiring — it is a direct chat turn on a web session with a
non-empty reply.

That made the opening's fourth beat redundant: Mika wrote a three-to-five
line menu of example tasks, then three chips appeared underneath offering
the same thing. The prose menu is the worse half — a member has to retype a
line they read, but can send a button — so the beat is gone and the opening
budget drops with it. Measured on a real local run: the first reply went
from 307 to 208 characters, and the chips still arrived 21s after the
kickoff.

The questionnaire profile stays in the kickoff. The suggestion pass resumes
the same provider session, so the profile now steers the chips as well as
the reply; only the sentence naming its purpose changes.

Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 15:59:56 +08:00
Lambda
1f81c988e3 feat(onboarding): Mika issue-first onboarding
Replaces the starter-agent welcome with one conversation: onboarding creates
Mika, the workspace's built-in Chief of Staff, and opens a real chat whose
first turn is a product-authored kickoff hidden from the transcript. Every
workspace — first and subsequent — is created through this flow.

Mika is a system agent, not an agent-template instance. Her product prompt is
//go:embed-ed and composed at claim time, so a release updates it without
touching any workspace's row; the row holds only the workspace's own notes.
Creation is server-owned and idempotent under a per-workspace advisory lock,
and archiving a system agent is rejected.

This is the pre-merge half of the branch, squashed while rebasing onto main.
Replaying its fifteen commits individually meant re-deriving each one against
a main they were never written for; the net change reconciles against today's
main in four files, so it is reconciled once, here.

Three of those four are main moving under the branch: MUL-5573 took
quick-actions generation server-side and dropped QuickActionsDisabled /
RegenerateQuickActionsFor from the task payload and the SendDirectChatMessage
signature, so this takes main's shape and keeps only the onboarding entry
point. The fourth keeps main's OnboardingLogoutButton wrapper around the
flow's new mode/onCancel props.

Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 15:59:56 +08:00
Bohan Jiang
757b09c801 fix(server): remove the last language seeds from the LLM prompts (MUL-5689) (#6348)
* fix(server): drop the Chinese seed from the chat title prompt (MUL-5689)

chatTitleSystemPrompt carried literal Chinese in two places: a rule that
spelled out "Chinese input → Chinese title", and a formatting example
listing "标题:" as a prefix not to use. It is the same defect as the
quick-actions label rule — a prompt that names a language, even only as a
formatting example, reads as permission to answer in it.

Chat titles are generated from the user's opening message alone, so they
have neither the ALREADY SUGGESTED feedback loop nor an agent reply to be
pulled by. Removing the seed is the whole fix; the language rule stays,
now stated without naming a language.

Nothing is lost by dropping the "标题:" example: chatTitleLabelPrefixes
already strips 标题/题目/主题 (and the English forms) from the model's
output, which is where that guarantee actually lives.

Co-authored-by: multica-agent <github@multica.ai>

* fix(server): correct two inaccuracies in the quick-actions language rule

Both from review on #6345, non-blocking there and deferred to keep that
merge unblocked.

"these instructions" was self-referential: the LANGUAGE RULE is itself an
instruction, so a model reading "ignore these instructions when choosing
the language" could read the rule as disowning itself. It means the
system prompt, so it now says so.

The system prompt claimed the user message "ends with" the LANGUAGE RULE.
It does not — the task line follows it. Reworded to "contains a LANGUAGE
RULE line near the end", which is what the renderer actually produces.
The rule's position is unchanged: still after the conversation and the
replayed labels, which is the part that matters.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 14:21:25 +08:00
Bao Ming
738217c275 MUL-5516: fix(issues): re-trigger blocked issues when resumed (#6155)
* fix(issues): re-trigger blocked issues when resumed

* fix(issues): enforce agent permission on resume

---------

Co-authored-by: baoming.lv <baoming.lv@ly.com>
2026-08-04 14:16:57 +08:00
Xichang(Seacen) Zhao
671a1a0c21 fix(server): read the chat's channel from its binding instead of guessing (#6330)
A claim reported chat_channel_type by trying the binding lookup once per
channel it knew the name of. The list was Slack alone, then {Slack, Feishu}
after MUL-4899. Any channel added later — WeCom is the first — fell off the
end and claimed as a web chat, so the runtime brief told the agent to deliver
files with `multica attachment upload` into a conversation that cannot carry
an attachment. That is the same failure MUL-4899 fixed, reintroduced by the
shape of the fix.

Every channel writes the same channel_chat_session_binding row and differs
only in channel_type, and UNIQUE (chat_session_id) allows at most one, so the
row already holds the answer. Read it by session id and take channel_type off
what comes back. chat_in_thread stays Slack-only, now keyed on the row's own
channel_type: the two commands it picks between are hardwired to the Slack
history reader, and no other channel has one.

The channel_type-scoped query stays for the outbound senders, which are
per-platform by construction and must not deliver into a foreign channel.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 14:04:01 +08:00
Naiyuan Qing
a5c1d44701 MUL-5642: fix(agents): stop the creation studio polling, and stop it losing work (#6307)
* feat(agents): make AI agent creation resumable (#6246)

Leaving the Agent Creation Studio destroyed the conversation. The unmount
cleanup called deleteChatSession, so a sidebar click, a tab close or a route
change deleted the builder session and every message in it — the bug external
users reported. Archiving instead (PR #6247) would have stopped the deletion
without giving anyone a way back in: builder sessions hang off a hidden
`kind = 'system'` carrier agent, which the `kind = 'user'` filter keeps out of
every chat list, so an archived one is unreachable rather than recoverable.

A creation conversation is now a durable object with its own address.

Server:
  - GET /api/agent-builder/sessions lists the caller's unfinished creations.
    Creator-scoped like every other chat read. It reports the CARRIER's
    runtime, not chat_session.runtime_id — the latter is the daemon's resume
    pointer and is deliberately left stale after a switch, so resuming from it
    would put the picker on a runtime that executes nothing (MUL-5163).
  - PUT /api/agent-builder/sessions/{id}/draft stores the configuration,
    including the edits the user typed but never sent. Migration 251 adds
    agent_builder_draft (no FK per repo rule; pruned explicitly by
    DeleteChatSession, the runtime teardown and the workspace teardown, and
    registered in the workspace-deletion manifest).
  - The payload is opaque to the server: its shape is the studio's AgentDraft,
    validated client-side. Teaching Postgres and the handler about it would
    create a second definition to keep in sync for no gain.

Client:
  - The session id lives in `?session=`, so a refresh, a back/forward and a
    reopened tab land back in the same conversation.
  - Leaving no longer deletes anything. The only destructive path is an
    explicit "discard", confirmed in a dialog, next to the create button.
  - Creating the agent archives the conversation instead of deleting it: it is
    the record of how that agent was designed, and an idle carrier costs
    nothing since usage is booked per task.
  - The configuration autosaves (debounced) and restores on arrival, with the
    applied-assistant-message marker stored alongside it so a restore cannot
    re-apply the last reply over edits made after it.
  - The 1.5s polling of messages and pending-task is gone. The global realtime
    sync already invalidates both per session id, exactly as it does for the
    main chat window, which has never polled.
  - The `<agent_draft>` block collapses to one "configuration updated" line.
    The regex now also swallows an unterminated block, which is what streaming
    produces — the raw payload used to scroll past on every turn.

The 2185-line agent-creation-studio.tsx is split into three routes
(`/agents/new`, `/agents/new/manual`, `/agents/new/ai`), its pure logic moves
to packages/core/agents/ with its tests, and the unreachable template flow —
`setMode("templates")` had no caller — is removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(agents): let the builder panes resize

The conversation / configuration split was not draggable. Two structural
reasons, both fixed by giving the group the same shape the chat page uses:

  - The panels reached the group through BuilderWorkspace's fragment, so they
    were not children the group could measure.
  - The group's children alternated between one panel (runtime setup) and two
    (conversation), under one persisted layout id.

The group now lives inside BuilderWorkspace with its two panels as its only
children, and the setup screen renders no group at all — it has nothing to
split.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ui): give the resize handle a cursor on hover

The separator had no cursor of its own, so the only signal that a split was
draggable arrived after the drag started — the library writes a global
`cursor: ... !important` while dragging, and nothing before it.

Fixed on the shared handle rather than at one call site: every split surface
(chat, inbox, issue detail, project detail, the agent builder) was missing the
same affordance. The library's drag-time rule still outranks this one, so the
cursor keeps narrowing to `e-resize` / `w-resize` once a panel hits its bound.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(agents): render the builder's draft block as an inspectable row

Every builder reply ends in an <agent_draft> block that rewrites the form on
the right. Flattening it to a line of prose said that something changed but not
what, and the payload — the only record of what the builder actually claimed —
was unreachable.

A settled reply now carries a full-width row saying the configuration was
updated, which opens the exact payload. A streaming one keeps a text line
instead: the block is still being written, so there is nothing complete to
open, and without the line the half-finished JSON scrolls past.

ChatMessageList gains an optional `renderAssistantAddon`. It is opt-in per
surface and undefined everywhere but this one, because no other chat speaks
this protocol — the alternative was to keep pushing an embedded protocol
through `transformContent`, which can only ever produce prose.

`extractBuilderDraftBlock` returns an unparseable payload verbatim rather than
withholding it: a malformed block is exactly when someone wants to read it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(agents): address review blockers on the resumable builder

1. Migration prefix collision. `251_agent_runtime_unbind` landed on main after
   this branch cut, so the backend's prefix-uniqueness guard failed. Renumbered
   to 252.

2. #6287 — the manual form still lost everything. The route split moved where
   you land, not what survives: the draft was `useState`, so a tab switch (the
   desktop shell mounts only the active tab) remounted it empty, and the
   beforeunload guard covered a hard reload and nothing else. It now persists
   through the repo's draft-store factory, scoped by what is being created — a
   blank agent and a copy of agent X are different work, and a copy of X is not
   a copy of Y — cleared once the agent is committed, and registered for logout
   / workspace-delete cleanup.

3. A saved draft with no messages was unreachable. The configuration form is
   editable from the moment a builder session exists and autosaves, so someone
   could open it, type a name and leave before the first turn; the list keyed
   "is this a draft" on messages alone, so that row existed and nothing could
   reach it. A session now qualifies on a message OR a stored draft, and sorts
   by whichever it has.

4. The debounce dropped the last edits. Its timer died with the component, so
   navigating away inside the 800ms window lost exactly the keystrokes the user
   had just made. The pending payload is now flushed on unmount.

`useUnsavedDraftWarning` is gone with its last caller: both routes persist, so
the browser prompt would have been warning about work that is already saved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ui): stop the resize cursor flipping mid-drag

The library narrows its cursor the moment a panel hits a bound — col-resize
while both directions are open, a one-way arrow once only one is. Truthful, but
it reads as a glitch: the icon changes under your hand halfway through a drag
you never stopped making.

`disableCursor` turns that global rule off; the handle's own `cursor-col-resize`
is now the only source. A drag captures the pointer and walks it across the
panels, away from the 8px handle, so the group carries the same cursor for as
long as a separator is active — otherwise it would fall back to a text caret the
instant the pointer left the handle.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(agents): key manual drafts by owner, drop the draft-block rendering

Two changes.

One slot destroyed the other flow's work. The manual draft was stored under a
single key: opening a blank form, or a copy of a different agent, refused to
adopt the stored draft and then immediately wrote its own empty form over it —
so a half-finished copy of agent A died the moment the user opened anything
else, before typing a character. Drafts are now keyed by what is being created,
the same shape the chat composer uses for its per-session drafts, and a slot is
dropped when its content is gone rather than parked blank (which also stops the
map growing a dead key per agent ever opened for duplication). Committing an
agent clears that flow's slot only.

The `<agent_draft>` block goes back to being hidden outright. Labelling it and
opening its payload dressed up machinery as content: the block drives the
configuration form, and the form is where its effect is already visible.
`renderAssistantAddon` goes with it — ChatMessageList is back to what it was,
since no surface needs the slot. The two-pattern strip stays: an unterminated
block is what streaming produces, and without matching it the raw JSON scrolled
past the reader on every turn.

Also removed six barrel exports nothing imported through, and unexported five
types only their own file used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(agents): keep a manual draft whose only edit is a picker

The "is this worth storing" predicate listed six fields by name, and the draft
serializes eleven. A form whose only change was the model, the thinking level,
the service tier, the access scope or a team grant read as untouched, so the
next save deleted its slot — picking a model before typing a name and switching
tabs lost the model.

Enumerating was the mistake, not the specific omissions: the predicate stops
covering every field added after it is written, and the failure is invisible
because each field saves correctly as long as some *other* field is also set.
It now compares the whole draft against a fresh one. The runtime stays outside
that comparison, on the entry rather than in the draft, because the form seeds
it on every visit and counting it would store a draft for a form nobody
touched.

Covered field by field, one edit at a time, so a future field cannot quietly
fall out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 19:00:00 +08:00
Multica Eve
4fe94a6d40 revert: "MUL-5637 fix(mcp): agent mcp_config is an authoritative allowlist (#6292)" (#6314)
This reverts commit aa349fed02.

Preflight flagged the server/daemon mixed-version gate as a release blocker for
today's v0.4.17 window (MUL-5655): a managed, non-inheriting mcp_config claimed
by a daemon that does not advertise authoritative-mcp-v1 fails the task with
mcp_config_daemon_outdated, and that failure is not auto-retryable. Reverting to
unblock the release; the fix should return once daemon capability coverage in
production is confirmed.

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 17:43:03 +08:00
Multica Eve
aa349fed02 MUL-5637 fix(mcp): agent mcp_config is an authoritative allowlist (#6292)
* fix(mcp): treat agent mcp_config as an authoritative allowlist

An agent's saved mcp_config was silently widened with the runtime host's
own user-level MCP servers, so an explicitly empty `{"mcpServers":{}}`
resolved to the COMPLETE host set instead of no servers at all — the
opposite of what the operator configured (GitHub #6283).

`--strict-mcp-config` was being passed correctly; the merge happened
before it, in the daemon, so strict mode constrained an already-widened
set. Introduced by #5277 and present in v0.4.16 through main.

Restore the three-state contract in resolveEffectiveMcpConfig:

  null / unset        -> inherit the provider's native MCP configuration
  {"mcpServers":{}}   -> strict empty, no host servers
  non-empty object    -> strict allowlist, exactly those servers

Two explicit inherit paths keep the additive behaviour reachable without
weakening the default:

- runtime_config.mcp.inherit_runtime = true opts an agent back in.
- The claim response now carries mcp_config_overlay_only so the daemon
  can tell an agent-authored config from a per-task Composio overlay.
  Without it, enabling an integration on an agent that never configured
  MCP would have stripped the host servers it was already inheriting.

Both decode paths fail closed: malformed runtime_config never enables
inheritance, and a failed runtime merge falls back to the agent's own
config.

The web MCP tab and the `agent create/update --mcp-config` help text
described the old additive behaviour, which is how a tightened config
could look correct while exposing every host server; both now state
which mode is in effect.

Note for rollout: the fix lives in the daemon, so self-hosted users must
upgrade the daemon — a server/UI upgrade alone does not apply it.

Co-authored-by: multica-agent <github@multica.ai>

* fix(mcp): close review gaps in the authoritative mcp_config change

Addresses the four must-fix findings from review of #6292.

1. Deleting the last managed server no longer widens access.
   removeManagedMcpServer cleared the config to null, which now means
   "inherit the host's MCP servers" — so a delete took the agent from one
   allowed server to every server on the host. It now leaves an explicit
   `{"mcpServers":{}}`. Restoring inheritance moved to a separate
   clearManagedMcpConfig action behind its own confirmation that states the
   widening. The delete dialog no longer claims "Runtime servers are not
   affected", which was the opposite of the truth.

2. The UI no longer promises a boundary an old daemon does not enforce.
   The strict semantics live in the daemon, so a config saved against an
   older daemon is not yet in effect. Adds the authoritative-mcp-v1 daemon
   capability:

   - The daemon advertises it and reports authoritative_mcp on the
     runtime-capabilities response.
   - The claim path fails closed: a managed, non-inheriting mcp_config
     claimed by a daemon without the capability cancels the task and
     returns 412 with an actionable message, instead of letting that daemon
     merge the host's servers in. runtime_config.mcp.inherit_runtime is the
     documented escape hatch, and it is honest — it declares that the
     operator accepts the host's servers.
   - The MCP tab shows "needs upgrade" rather than "Not exposed" while the
     bound runtime lacks the capability.

3. Saving OpenClaw settings no longer drops the inherit opt-in.
   parseOpenclawRuntimeConfig discarded unknown keys and the tab persisted
   the result as the whole runtime_config, so one unrelated routing save
   silently deleted mcp.inherit_runtime. Unknown keys now round-trip
   through OpenclawRuntimeConfig.passthrough, excluded from the dirty check
   so they cannot make the form look edited.

4. Documents the new semantics in the built-in creating-agents skill and
   its source map: the three states, the persisted
   runtime_config.mcp.inherit_runtime field, and the claim-time capability
   gate.

Also corrects the PR's rollout claim: there is no database migration, but
this does add a persisted JSON field and change the meaning of an existing
one.

Co-authored-by: multica-agent <github@multica.ai>

* fix(mcp): stop the authoritative-daemon gate from blocking valid claims

CI's backend job failed three handler claim tests with the new 412. Two
distinct problems, both real:

1. The gate fired for a non-object mcp_config. 66 handler fixtures seed
   `[]`, which is not a valid MCP config and cannot carry `mcpServers`, so
   it expresses no boundary to protect. An old daemon does not widen it
   either: mergeRuntimeAndAgentMcpConfig fails to unmarshal a non-object
   and falls back to the agent config alone (verified directly). Gating
   these blocked tasks with no security benefit, so the gate now requires a
   JSON object.

2. The shared daemon test-request helper advertised no capabilities, so
   every claim test was accidentally simulating a pre-#6283 daemon. It now
   defaults authoritative-mcp-v1 on, matching what every current daemon
   sends. Only that capability — skill-bundles / coalesced-comments / rpc
   are feature negotiations whose absence tests real legacy behaviour, so
   they stay opt-in per test.

Adds claim-level coverage for the gate itself, which is what the unit tests
alone could not catch: an outdated daemon gets 412 with an actionable
message and the task is cancelled; a capability-advertising daemon gets
200; the inherit_runtime opt-in lets an outdated daemon through; and an
unmanaged or non-object config is never gated.

Verified against a real migrated schema this time (throwaway Postgres),
which is how the three failures were reproduced locally and confirmed
fixed: `go test ./internal/handler ./internal/daemon` both ok.

Co-authored-by: multica-agent <github@multica.ai>

* fix(mcp): surface the daemon-upgrade refusal and stop gating safe providers

Addresses the second review round on #6292.

1. The refusal is now visible wherever the operator looks. The default
   claim path is the machine-level BATCH endpoint, which skips build
   failures and still answers 200 {"tasks":[]}, so the previous bare
   CancelTask showed a task that vanished with no stated reason — turning an
   explicit upgrade requirement into an unexplained failure. The claim path
   now fails the task with a new classified reason,
   mcp_config_daemon_outdated, plus the actionable message. That reaches the
   user on all three claim paths and on any daemon version, which a new
   response field could not: the audience is by definition a daemon too old
   to read one. The per-runtime path keeps its 412.

   The reason is deliberately not auto-retryable — the same outdated daemon
   would claim the retry and fail it again.

2. The gate no longer cancels safe tasks. It applied to every provider, but
   only claude / codebuddy / codex / cursor / opencode / openclaw were ever
   merged with host MCP by an old daemon (loadRuntimeMcpServerConfigs).
   Qwen was never merged and already had strict semantics, so its tasks were
   being failed for a risk that does not exist. Scoped via
   providersOldDaemonsMergedRuntimeMcp; an unknown provider does not gate,
   because the gate should only fire where the old behaviour is concrete.

3. The new authoritative_mcp flag now goes through the API schema layer.
   Both local-skills responses were returning raw network JSON, so the flag
   that decides whether the UI may assert an MCP boundary rested on an
   unchecked type assertion. Adds RuntimeLocalSkillListRequestSchema with
   authoritative_mcp and mcp_supported defaulting to FALSE — the fail-closed
   direction — and a MALFORMED_ fallback that cannot express a guarantee.

Claim-level tests now cover all three paths, which is what the previous
helper-only tests missed: per-runtime 412, batch recording the refusal on
the task while still delivering the healthy tasks in the same batch, WS RPC
refusing and accepting, the qwen negative case, the inherit_runtime escape
hatch, and unmanaged / non-object configs.

Verified against a real migrated schema (throwaway Postgres):
go test ./internal/handler ./internal/daemon ./pkg/agent ./pkg/taskfailure
./internal/service all ok.

Co-authored-by: multica-agent <github@multica.ai>

* fix(mcp): register the new failure reason and wire its copy into the UI

Addresses the third review round on #6292.

1. mcp_config_daemon_outdated was declared but never registered in
   taskfailure.allReasons, so metrics.NormalizeFailureReason missed the
   known-value map and fell through to free-text Classify() — relabelling a
   platform-side refusal as `agent_error.unknown` (verified directly) and
   leaving the Prometheus series un-pre-warmed. Registered it, canonical
   count 22 → 23 (platform 8 → 9), with the wire value and IsAgentError split
   pinned. New test pins the WHOLE canonical set through
   NormalizeFailureReason so forgetting the next reason fails a test instead
   of quietly mislabelling a metric; NormalizeFailureReason had no coverage
   at all before.

2. The upgrade copy was dead. The locale strings landed last round but
   neither consumer mapped the reason: chatFailureCopy fell back to generic
   failure text with the actionable detail buried in the collapsed raw
   error, and task-failure.ts rendered the bare wire value
   `mcp_config_daemon_outdated` in the agent activity list and issue
   execution log. Both are mapped now, with regression tests, plus the
   runtime class pinned in failure-class.test.ts. This directly contradicted
   the claim in the claim-path comment that every path reaches the user, so
   that is now actually true.

3. providersOldDaemonsMergedRuntimeMcp is documented as what it is: a FROZEN
   record of what pre-capability daemons merged, not a mirror of the daemon's
   current provider switch. The old "keep the two lists in lockstep" note was
   actively harmful advice — runtime MCP discovery for a new provider can only
   ship in a daemon that already advertises the capability (never gated), so
   adding it here would fail tasks on old daemons that never merged for it,
   re-creating the qwen false-positive. Pinned with a test.

Also corrects a stale count in task-failure.ts (7 → 9 platform reasons).

Verified against a real migrated schema (throwaway Postgres): full backend
suite green apart from the pre-existing environmental cmd/multica guard; all
9 TestMcpGate_* integration tests pass.

Co-authored-by: multica-agent <github@multica.ai>

* docs(taskfailure): correct taxonomy counts and finish the reason registration

Non-blocking nits from the fourth review round on #6292.

- Taxonomy counts now say 23 reasons / 9 platform-side. Registering
  mcp_config_daemon_outdated last round updated the assertions but not the
  prose. Swept the whole repo rather than only the flagged lines, which
  turned up four more that were already stale at 21 and drifted further:
  handler/dashboard.go, daemon/poisoned.go, core/types/agent.ts, and the
  db/queries/task_usage.sql comment sqlc copies into the generated file.

  The generated file's comment was updated by hand to match its source.
  Running `sqlc generate` churned 58 lines across 47 unrelated files — the
  local sqlc version differs from the one that produced the checked-in
  output — so that churn was reverted and only the one intended line kept.

- failure_test.go's `required` list now includes
  ReasonMcpConfigDaemonOutdated. Length and label assertions already covered
  the reason, but the list is documented as the complete canonical set, so
  the omission contradicted its own comment.

- Restored the line break in chat-message-list.test.tsx that a previous edit
  of mine collapsed.

Comment, test-fixture and formatting only; no behaviour change.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 16:06:07 +08:00
Wangjue Yao
69d3208586 fix(channels): handle /new in shared router (#6251)
* fix(channels): handle /new in shared router

* test(channels): cover /new provider reset end to end

* fix(channels): preserve command source across rewrites

---------

Co-authored-by: 无瑕 <yaowangjue.ywj@antgroup.com>
2026-08-03 15:40:33 +08:00
Xichang(Seacen) Zhao
cb04226fd4 MUL-5632: fix(engine): broadcast the full issue payload for chat-created issues (#6278)
* fix(engine): broadcast the full issue payload for chat-created issues

An issue opened with /issue from a chat channel broadcast
{"issue_id": <uuid>} — the minimal fallback IssueService.Create emits
whenever a caller leaves IssueCreateOpts.BroadcastPayload nil, which the
engine did. Every issue:created consumer reads payload["issue"], so the
missing key meant extractIssueFields (cmd/server/subscriber_listeners.go)
failed its map assertion and returned false before it ever reached the
id / creator_id check. The listener returned early and no auto-subscribe
rule ran, so the person who typed /issue was never subscribed to the
issue they had just filed and got no notifications for it. Both channels
that register with the engine — Feishu/Lark and Slack — were affected.

The HTTP handler already supplies a BroadcastPayload and autopilot
publishes its own event through issueToMap; only the engine path was
short. Export issueToMap as IssueToMap and have the engine use it, so a
single builder is the one source of truth for that shape instead of
three descriptions drifting apart. The workspace issue prefix comes from
the GetWorkspace call the /issue path already made for the chat reply's
identifier, hoisted so it is read once and used for both.

No regression — the payload only gains keys, and consumers read named
fields. The activity and notification listeners type-assert
payload["issue"] to handler.IssueResponse and still skip a map, exactly
as they already do for autopilot-created issues. The subscriber listener
accepts either shape. The workspace WS fanout marshals the payload
as-is, so open clients now render a chat-created issue live instead of
waiting for a refetch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(service): complete the issue:created payload contract

IssueToMap omitted project_id, stage, metadata and properties, all of
which handler.IssueResponse — the other rendering of the same event —
always emits. Clients type both as a complete Issue and insert the
object straight into the list cache without runtime validation, so an
issue created by autopilot, quick-create or the chat /issue command
appeared in open clients missing its project and custom properties
until the next refetch. Broadcasting the full issue from the chat path
would otherwise have spread that existing defect to a third entry
point.

Fill in the missing keys and pin the contract with a test that fails if
the two renderings ever drift apart again, so adding a field to
IssueResponse without adding it here is caught in CI rather than in the
UI. metadata and properties render as {} when unset, matching the
"always present" part of the contract.

Also route both identifier renderings through service.IssueIdentifier,
so a workspace lookup that degrades the prefix cannot show "#42" in the
chat reply while the realtime list shows "-42".

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* docs(service): name the actual IssueToMap call sites

The doc comment and the shape test listed quick-create among the
non-HTTP publishers of the issue payload. It is not one: the three call
sites are autopilot and the channel engine on issue:created, and the
background stuck-issue status reset on issue:updated.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 14:29:47 +08:00
Bohan Jiang
d38da27ed4 MUL-5619: fix(cli): surface the server's 409 message instead of the generic conflict template (#6267)
* fix(cli): show the server's conflict message instead of the generic 409 template

Every 409 this API returns is a deterministic refusal that names its own fix
("a skill with this name already exists", "set parent_id (--parent) to <id>").
The CLI replaced all of them with a template that says the opposite — that the
state changed underneath you and you should re-fetch and retry. Agents took the
retry hint literally: GH #6264 reports 15+ identical retries over 10 minutes
followed by hours spent chasing an optimistic-concurrency theory that never
existed, and GH #5948 is a second user misdiagnosing the same way. MUL-4417 had
already written the useful message server-side; it just never reached anyone.

Route 409 through the same server-message extraction 400/422 already uses, so
roughly forty hand-written conflict messages across skills, agents, runtimes,
labels, projects and comments become visible by default. A body we cannot
recognize still falls back to the template, so this never dumps a raw response.

extractServerMessage now prefers prose over a bare identifier, because a few
endpoints put a stable code in "error" and the sentence in "message".

MUL-5619

Co-authored-by: multica-agent <github@multica.ai>

* fix(comments): stop telling a wrong --parent that it posted a top-level comment

The reply guard returns one message for two different mistakes. A resumed
session that carries a previous turn's --parent forward (GH #6264) did not ask
for a top-level comment, but is told it did — which sends it looking for a
new-thread opt-in (GH #5383) instead of correcting the parent it already passed.

Split the copy: name the rejected parent when one was supplied, and keep the
existing top-level wording for the parentless case. Both still point at the
trigger comment to use.

MUL-5619

Co-authored-by: multica-agent <github@multica.ai>

* fix(runtime): return 500, not a 409 echo, when the update store fails

InitiateUpdate answered every UpdateStore.Create failure with a 409 carrying
err.Error(). The in-memory store only ever returns errUpdateInProgress, so this
looked safe — but the Redis store also wraps infrastructure failures as
"reserve active update: <dial error>" and "persist update request: <error>".

Surfacing 409 bodies in the CLI turns that into a user-visible leak of internal
addresses, and labels an outage as a conflict the caller could fix by retrying.
Classify instead: errUpdateInProgress keeps its 409 and its actionable message,
everything else is logged and answered with a 500 and fixed copy.

Also pins the prose-over-machine-code preference for validation bodies, which
the shared extractor applies to 400/422 as well as 409. Only the issue-table
endpoints are shaped that way and none is reachable from the CLI today, but the
change is intentional and should fail loudly if reverted.

MUL-5619

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 13:35:04 +08:00
Multica Eve
b06af2ae17 feat(runtime): unbind agents on runtime delete instead of destroying them (#6220)
* feat(runtime): unbind agents on runtime delete instead of destroying them

Deleting a runtime archived its agents and then hard-deleted the rows, so the
agents and every conversation with them disappeared — while the confirmation
dialog said "archive", which a user reasonably reads as recoverable. Retiring a
laptop is an ordinary action; losing the agents configured on it is not an
ordinary consequence.

An agent is now a persistent business object and a runtime is replaceable
execution capacity: deleting a runtime unbinds its agents. `runtime_id IS NULL`
means unbound — orthogonal to archived — and the agent keeps its instructions,
skills, chats, labels, channel installations, autopilots and task history.
service.AgentReadiness already refused an agent with no runtime, so the
scheduling safety gate needed no change.

Two columns become nullable, not one. Without `agent_task_queue.runtime_id`,
deleting the runtime still cascades the task history away (and task_message /
task_usage / task_token with it), so the agents would survive with no record of
anything they did — the same class of loss. A NOT VALID CHECK keeps NULL confined
to history: an active task must always have a runtime, so claim / dispatch /
delivery-CAS paths can never observe one without. It is written against
completed_at rather than a status list so a future non-terminal status fails
closed instead of slipping through.

Two prerequisites this depends on:

- 'deferred' (migration 128) was missing from CancelAgentTasksByRuntimeOrAgent.
  It went unnoticed because the delete used to cascade those rows away; with the
  new CHECK it would abort the delete and make the runtime undeletable.
- The channel-installation / label / chat-pin / invocation-target / draft-restore
  cleanups were scoped to "archived agents on this runtime". Archived user agents
  now survive, so that scope is narrowed to kind='system' — otherwise the fix
  would produce a subtler loss: agent alive, configuration wiped.

Also removes the squad guard that refused (409) when an active squad's leader was
an archived agent on the runtime, plus the archived-squad delete that existed
only to get past squad.leader_id's RESTRICT FK. The leader is no longer deleted,
so nothing needs to be given up to retire a machine. Autopilots are no longer
paused either: their assignee survives, and a rebind restores them without the
owner having to remember to re-enable.

Reason codes: an unbound agent reports agent_runtime_required, not
runtime_offline. The copy for runtime_offline tells users to reconnect a machine;
an unbound agent has no machine to reconnect, and the fix is to bind a runtime.
Chat's bare 409 string gains the same code so the composer can offer that action.

API: agents gain runtime_bound. runtime_id stays a string (empty when unbound) so
installed clients keep parsing and no gated two-release rollout is needed. The
confirmed-delete endpoint is /unbind-agents-and-delete; /archive-agents-and-delete
still routes to it, and the compared expected_active_agent_ids set is unchanged —
widening it would 409 every older client forever.

Co-authored-by: multica-agent <github@multica.ai>

* fix: make runtime unbinding recoverable

Co-authored-by: multica-agent <github@multica.ai>

* fix: address runtime unbind review nits

Co-authored-by: multica-agent <github@multica.ai>

* fix: resolve runtime unbind review blockers

Co-authored-by: multica-agent <github@multica.ai>

* fix(migrations): renumber runtime unbind after main merge

Co-authored-by: multica-agent <github@multica.ai>

* test(daemon): avoid late-request lease flake

Co-authored-by: multica-agent <github@multica.ai>

* test(autopilots): bind validation fixture runtime

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 12:39:27 +08:00
Jiayuan Zhang
28b6105edc fix(subscribers): notify the human an agent files sub-issues for (MUL-5483) (#6209)
When an agent created a sub-issue while working on a human's behalf, that human
received no notifications for it at all. issue_subscriber modelled ACTOR
identity, so an agent-created, agent-assigned issue had a full subscriber list
and zero members to deliver to. The platform already knew who the work was for
(agent_task_queue.originator_user_id, MUL-4302); notification never asked.

- attribution.DelegatedSubscriber: one shared rule over the same origin
  waterfall ClassifyDirect uses. agent_create subscribes the originator as
  'delegated'; quick_create keeps the direct 'creator' tier; autopilot and
  degraded attribution subscribe nobody.
- Delegated is a reduced delivery tier: in_review/done/cancelled/blocked plus
  failures and mentions. Routine churn is suppressed, and the parent bubble
  cannot re-deliver what the tier dropped.
- Unsubscribe becomes stateful: an unsubscribed_at tombstone survives later
  rule passes, and opt_out_scope distinguishes "this issue" from "this subtree"
  so a narrow opt-out no longer silently suppresses future children.
- Subtree unsubscribe is its own endpoint. A body flag cannot fail loudly
  against an older backend (Go drops unknown fields); an unknown route 404s,
  which the UI now surfaces with a distinct message.
- Eligibility and the write share one statement under a (workspace, user)
  advisory lock that subtree unsubscribe and member revoke also take, closing
  the check-then-insert races. Revoke additionally clears the departing
  member's subscriptions in the same tx.
- UI explains a delegated subscription and offers both unsubscribe scopes.

Migrations 249/250 add the delegated reason, the opt-out tombstone, and the
opt-out scope, using NOT VALID + VALIDATE CONSTRAINT so the widened CHECK does
not scan issue_subscriber under an exclusive lock.

Reviewed across eight rounds; an earlier write-time subtree roll-up was built
and then removed in full once it proved unfixable without serializing every
topology mutation. The parent's own status transition already carries that
signal.

Closes MUL-5483.
2026-07-31 16:52:17 +08:00
Multica Eve
f48bd655bc MUL-5562: optimize application-owned workspace deletion (#6230)
* fix(workspace): optimize application-owned deletion

Co-authored-by: multica-agent <github@multica.ai>

* fix(workspace): address deletion review findings

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 16:09:19 +08:00
Jiayuan Zhang
f13969b996 refactor(chat): generate quick actions server-side via the LLM layer (MUL-5573) (#6214)
* refactor(chat): generate quick actions server-side via the LLM layer (MUL-5573)

Follow-up suggestions were produced by a second, full provider CLI invocation
per chat turn: the daemon resumed the just-finished session and ran a
suggestion-only pass. That pass inherited the main turn's exec options, so its
20s budget had to cover process spawn, every MCP handshake, session replay, and
model reasoning at the agent's own thinking level — typically 8-15s of visible
skeleton, and every turn paid two provider cold starts.

Generate them here instead, through the same pkg/llm layer that backs chat
auto-titling. Suggestions need no tools, workdir, or agent identity — only the
tail of the conversation — so a bounded 8s call on the deployment's small model
replaces the whole resumed turn.

Quality changes that came with the move:

  - The prompt now states the frame explicitly ("you write FOR THE USER"). The
    old pass ran inside the agent's session and inherited the runtime brief's
    identity, which drifted suggestions toward agent-operations actions.
  - Previously-offered labels are replayed as ALREADY SUGGESTED. The old
    architecture had the opposite effect: on providers that append on resume,
    each pass saw its predecessor's JSON and anchored on it.
  - A failed generation broadcasts failed=true. Before, a timeout delivered an
    empty array — indistinguishable from "nothing worth suggesting", so every
    slow pass read as a quality problem.
  - The in-band footer is still stripped from replies but its actions are now
    discarded, so a pre-upgrade session is not pinned to the retired
    suggestions with the replacing pass suppressed.

The refresh path no longer enqueues an agent task: it validates the target and
calls the same generator, which also drops the not-resumable refusal — a session
whose runtime was rebound can now be refreshed. Client contract is unchanged
(chat:done pending flag, chat:quick_actions supplement); the only frontend
change is the pending window, resized from 30s to 12s to match the new budget.

Also removes the daemon's TMPDIR-after-cleanup hazard by construction: the old
pass started after runTask's defers had already deleted the temp dir it was
still pointed at.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(chat): drop the quick-actions opt-out setting (MUL-5573)

Suggestions are always on. The Settings → Chat toggle is removed along with
the whole per-turn opt-out path it fed: the persisted client preference, the
quick_actions_enabled send field, the quick_actions_disabled task stamp, and
the eligibility gate that read it.

The toggle predates server-side generation, when it could only hide chips a
provider pass had already paid for. Now that generation is a bounded call the
server decides on, an off switch buys nothing a user would miss, and it was
the last piece of UI implying the feature might be unavailable.

agent_task_queue.quick_actions_disabled is no longer written (dropped from
CreateChatTask's INSERT; the column keeps its false default). Left in place
alongside regenerate_quick_actions_for for a later drop migration — removing
columns an already-running binary still inserts would break mid-deploy.

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): address quick-actions review findings (MUL-5573)

Four defects from review of the server-side generation change.

1. Automatic failures were reported as refresh failures. The generator
   broadcast failed=true on any LLM error, but the client turns every
   failed=true into a "couldn't refresh" toast — so an automatic timeout
   popped a toast for an action the user never took. This also contradicted
   ChatQuickActionsPayload.Failed, which documents false for the automatic
   pass. The caller now passes its origin; only an explicit refresh reports.

2. Generation context was not bound to the target turn. The pass re-read the
   session's newest messages while always writing to the task it was handed,
   so a turn landing between the completion callback and the detached read
   supplied the context for a reply it did not belong to. Worse, a user
   typing a follow-up in the second after a reply left the window ending on
   a user row, which the old code treated as "nothing to build on" — that
   turn silently never got pills. The window is now anchored on the target
   assistant message and queried strictly before it.

3. No concurrency or idempotency bound on generation. Refresh stopped
   creating a task, so the busy check could not see a pass already running:
   two refreshes both returned 202, spent two upstream calls, and raced to
   write one row. Nothing bounded generation process-wide either. Adds a
   per-session in-flight guard (refresh now 409s on a duplicate) and a
   process-wide ceiling; a shed pass still resolves the client placeholder
   so no skeleton hangs on work that never started.

4. A new daemon could not safely talk to an older server. The refresh task
   discriminator was deleted, so a regenerate task from such a server fell
   through to the ordinary chat path: no user message, but the agent would
   answer anyway and the server would persist it as a real reply. The field
   is restored as a refusal marker only — the task completes empty, which is
   the shape the retired pass produced and which that server writes no row
   for. Not a restored execution path.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 15:59:50 +08:00
Multica Eve
51f44873cc fix(labels): always enable resource labels (MUL-5563) (#6225)
* fix(labels): always enable resource labels (MUL-5563)

Co-authored-by: multica-agent <github@multica.ai>

* docs(labels): clarify resource label rollback safety (MUL-5563)

Co-authored-by: multica-agent <github@multica.ai>

* docs(labels): correct compat client range (MUL-5563)

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 14:41:05 +08:00
Multica Eve
c3f5df8bf4 MUL-5492: fix timeline cap dropping newest entries + stop double-broadcasting descriptions (#6175)
* fix(timeline): cap the issue timeline at the newest end and report the clamp

The per-issue timeline cap was applied with ORDER BY created_at ASC LIMIT
2000, so once an issue accumulated more than 2000 comments or activities
the cap discarded the NEWEST rows. The timeline appeared to stop at some
point in the past and every later event was invisible, with nothing in
the response indicating anything was missing.

Activity is machine-paced — description autosave, every agent run, status
and assignee changes all write rows — so this was reachable in normal use,
not only on pathological issues.

Take the window with the keyset ordering (created_at DESC, id DESC) in a
subquery and re-sort ascending in the outer query. This keeps the
chronological contract for every existing caller, including the comment
list endpoint that shares ListCommentsForIssue, and is served as an
index-only scan by the idx_*_keyset indexes already added in migration
068 — no new migration, no call-site changes.

Two things beyond the ordering flip:

- Clamp both lists to a shared window floor. The two caps are applied
  independently, so each list has its own floor. Merging windows with
  different floors produces a timeline that looks continuous but, below
  the higher floor, contains only one of the two kinds — e.g. comments
  with no interleaved activity. That is worse than a timeline that
  visibly stops, because nothing about it looks wrong. Both lists are
  now clamped to the newest floor, so the result is a contiguous,
  correctly interleaved slice.

- Stop truncating silently. The unpaginated response is a bare JSON
  array with nowhere to put a flag, so the clamp is reported via
  X-Timeline-Truncated and X-Timeline-Window-From, added to
  ExposedHeaders because a custom response header is otherwise
  unreadable from browser JS. The legacy wrapped shape's has_more_before
  is now truthful instead of hardcoded false.

Queries read one row past the cap so "hit the cap" is distinguishable
from "holds exactly 2000 rows", which would otherwise report a complete
timeline as truncated and drag the other list's window down with it.

Regression tests cover all four properties and were confirmed to fail
against both the original query and a floor-less DESC flip.

Co-authored-by: multica-agent <github@multica.ai>

* perf(realtime): stop broadcasting two full descriptions on issue:updated

issue:updated carried prev_description alongside the new description in
the issue object, and the WS forwarder reuses the producer's payload map
verbatim. Every debounced description autosave therefore pushed two full
copies of the description to every connection in the workspace, including
users who did not have the issue open. The DB write is O(1); the fanout
was O(connections x description size), and it repeats on every pause in
an editing session.

prev_description and prev_title exist only for in-process listeners —
subscriber_listeners adds newly @mentioned users, notification_listeners
builds mention notifications, activity_listeners records the title
change. No client reads them: IssueUpdatedPayload in
packages/core/types/events.ts does not declare either field.

Project the payload on the way out. The bus dispatches bus.Subscribe
handlers before the SubscribeAll forwarder, so the in-process consumers
are unaffected, and projecting at the forwarder covers both the single-
node Hub and the Redis relays since that is where the frame is
serialized. The producer's map is copied rather than mutated.

The removed keys are listed in a table rather than an if on one event
type. The bug was structural, not a typo: the next large field added to
a published payload inherits the same cost silently, and a declarative
list puts the internal/external payload boundary in one reviewable
place.

issue.description itself is deliberately kept — clients apply it to
their cache, so stripping it would trade fanout bytes for N refetches.
Cutting the remaining fanout needs the per-issue scope routing already
scaffolded server-side for MUL-1138, which is blocked on the client
sending subscribe frames.

Tests assert both halves: the keys are absent from the serialized frame,
and the in-process listener still receives them.

Co-authored-by: multica-agent <github@multica.ai>

* fix(timeline): keep comment threads whole under the newest-N cap

Review found that the newest-N window can orphan a reply, and an orphaned
reply is invisible rather than merely mis-nested: the timeline builds its
top level from "activities + comments with no parent_id" and renders
replies by looking them up under their parent, so an orphan sits in the
map with no card to render it. MUL-1847 / #2263 was exactly this shape —
1 root + 29 replies, root dropped, all 29 vanished from the UI while the
API returned them.

Root cause of the regression: capping with the OLDEST n could never
orphan anything, because a reply is always newer than its parent, so a
prefix of the timeline is closed under "parent of". A newest-n window is
a suffix and has no such property. Flipping which end the cap bites
silently invalidated a structural property the comment tree relies on.

Two changes.

Drop the cross-kind clamp. The previous revision trimmed both lists to a
shared floor so the window was provably contiguous. That was the wrong
trade and it was also the dominant source of orphans. Comments are
human-paced (p99 ~30, max ever observed ~1.1k) and essentially never
reach the cap, while activity is machine-paced and reaches it routinely
— so the shared floor was almost always the activity floor deleting
comments that had been fetched successfully and would have rendered
fine. On an issue with thirty comments it was pure loss. Each list now
reports its own truncation and X-Timeline-Truncated names which kinds
were affected. Not clamping costs only activity density in the older
part of the range, which is metadata rather than content, and it is
reported rather than hidden.

Complete parent chains for the case that remains — comments themselves
exceeding the cap. ListMissingAncestorComments walks parent_id upward via
a recursive CTE and returns the ancestors not already held; the handler
merges them and restores the ascending order. This only ever ADDS rows,
so unlike clamping it cannot hide anything the caller would have seen,
and it is bounded by the number of distinct missing ancestors. Whole-
thread windowing was considered and rejected: a single thread can exceed
any row budget, so its degradation is not definable.

Applied to the shared query's default list path too, not just the
timeline. foldResolvedThreads documents a COMPLETE-thread set as its
precondition and comment.go asserts the default list mode satisfies it;
a half thread made that assertion false and a resolved thread whose root
was cut stopped folding correctly.

Also drops X-Timeline-Window-From. It was second-precision RFC3339 while
the real ordering key is (created_at, id) at full precision, so it could
not resume a read without skipping or repeating rows inside a shared
second. A resumable cursor should be opaque and carry both halves; worth
designing when there is a consumer rather than shipping as a lossy
approximation.

Tests: the reviewer's exact scenario, plus a no-orphaned-replies
invariant on both endpoints, the fold-still-works case, and a guard that
activity truncation does not delete comments. Each was confirmed to fail
with the fix disabled. TestListTimeline_JointWindowHasNoOneSidedRegion
was rewritten rather than deleted — it pinned the clamp behaviour being
abandoned here, so leaving it would lock in the wrong contract and
deleting it would drop the coverage.

Co-authored-by: multica-agent <github@multica.ai>

* fix(comments): bound parent-chain completion and stop folding partial threads

Second review round on MUL-5492. Four must-fixes, all stemming from one
conflation: parent-chain closure, a newest-N window, and a complete
thread are three different things. Closure makes a reply renderable; it
does not license thread-level derivations.

Do not fold a truncated read. fetchCommentsForList closes parent chains,
but older siblings and descendants of a retained reply stay outside the
window, so the set holds partial threads. Folding them produced wrong
answers rather than incomplete ones: a resolution reply outside the
window made a resolved thread look unresolved, and folded_count reported
a total derived only from retained replies. The previous revision claimed
closure restored foldResolvedThreads' COMPLETE-thread precondition; it
did not, and that claim is removed. --recent and untailed --thread still
return whole threads and still fold.

Bound the walk. The recursive CTE climbed to the root with no depth
limit, so a deep chain could drag its entire ancestry back and defeat the
row cap it was meant to preserve. Depth is genuinely unbounded in stored
data: the general write path stores the exact comment being replied to
(only the agent path collapses to the thread root), so chains can run far
deeper than the two levels the UI renders. Replaced with a layered walk
under explicit budgets — 2000 extra rows, 64 levels — making a response
provably bounded by 4000 comments, or 6000 timeline entries with
activities.

Scope every level to the tenant. The CTE's recursive branch matched on
parent_id alone. parent_id carries a foreign key to comment(id) but not
to a matching issue, so a stray cross-issue parent reference is
representable, and the walk would have followed it into another issue's
comments. The replacement filters issue_id and workspace_id on every
level. A negative test confirms the leak: with the filter removed it
reports "a comment from another issue leaked into this issue's response".

Degrade by pruning, not by orphaning. When a budget is exhausted, a
parent row is missing, or a parent is out of scope, keepRootConnected
drops the affected comments instead of returning replies the UI cannot
render. Dropping a node also drops its descendants, since their chains
run through it. Returning fewer new replies is conservative and already
signalled as a truncated read; leaking another tenant's data, returning
an unbounded response, or emitting invisible orphans are all worse.

Also: probe read on the comment list so exactly-2000 is not misreported
as truncated, which would needlessly suppress the fold; CommentsTruncated
is carried on fetchCommentsResult rather than inferred from the result
length, which is meaningless once completion adds rows. The new query
returns db.Comment directly instead of a hand-copied row, which is how
quick_action_id came to be dropped after the rebase — a backfilled
quick-action root would have rendered as a raw prompt.

Corrected three inaccurate comments: the "index-only scan" claim (the
index avoids the sort but does not cover SELECT *), the "write path
collapses replies to root" claim, and a test header still describing the
abandoned contiguous-window behaviour.

Tests cover exactly-at-cap still folding, truncated reads not folding
(both reply-resolved and root-resolved), depth beyond budget pruned not
orphaned, shared ancestors fetched once, cross-issue parents never
crossing the boundary, and quick_action_id surviving backfill. Each was
confirmed to fail with its specific fix disabled; the reply-resolved fold
test was reshaped after the first version passed for the wrong reason.

Co-authored-by: multica-agent <github@multica.ai>

* fix(timeline): preserve complete threads under comment cap

Co-authored-by: multica-agent <github@multica.ai>

* fix(comments): preserve newest bounded views

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 14:40:16 +08:00
Naiyuan Qing
13b06f038e fix(issues): count agents working in the surface, not the workspace (MUL-5525) (#6191)
* fix(issues): count agents working in the surface, not the workspace (MUL-5525)

The "N agents working" chip ran its own workspace-wide
`/api/working-agents` read while the list it filters came from the
surface's own compiled query. Two definitions of the same question, so on
a project page the chip could advertise agents working nowhere near that
project and then open an empty list. Every other narrowing the list knows
about — status, priority, assignee, creator, label, custom property,
date, sub-issue display, the /issues Members/Agents tabs — was invisible
to the count for the same reason. Only /my-issues (relation) and the
issue-detail sub-issue chip (parent) were narrowed, because those were
the two cases the endpoint had grown parameters for.

Rather than add a `project_id` parameter and leave the next dimension to
be discovered the same way, the count now comes from a `working_agents`
facet on the existing issue-table facets endpoint: same scope, same
filters, same compiled WHERE clause the rows come from, joined to running
issue tasks and grouped by agent. Correct-by-construction instead of
correct-by-keeping-two-lists-in-sync.

- Facet is disjunctive like every other one: it drops `working_issue_ids`
  / `working_only`, so the answer is identical whether the filter is on or
  off and the number does not move when you click the chip.
- Facet keys are agent ids, so they pass the same visibility gate as the
  other workspace-wide agent aggregations — a private or non-allow-listed
  agent is not disclosed by id, count, or presence.
- Gantt keeps a client-side count: its canvas projection (scheduled +
  dated + showCompleted) cannot be expressed in the Table query spec, so
  it counts the agents holding canvas rows instead.
- The chip is now presentational; `undefined` renders the existing
  indeterminate label rather than a zero it cannot stand behind.
- Removes the MUL-4884 `workingScopeIssues` plumbing, dead since the count
  moved to the endpoint in MUL-5200, keeping only the Gantt branch that
  still has a real consumer.

Also fixes the empty state that bug dropped you into: a filtered-empty
surface claimed "No issues linked — create one" while 41 issues sat behind
the filter. Shared filtered-empty state now precedes each surface's own
copy and offers to clear exactly the filters it blames.

Verified: pnpm typecheck, pnpm test (469 files), pnpm lint (0 errors),
go test ./internal/handler (new facet tests cover project scope, status
and sub-issue narrowing, filter-independence, and the access gate).

Co-authored-by: multica-agent <github@multica.ai>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(issues): keep the working-agents unknown state unknown (MUL-5525)

The chip correctly refused to print a number for an unresolved projection,
then handed the hover card `agents ?? []` — so hovering an indeterminate
chip read "No agents working right now". That is the same unearned claim
this issue is about, made by the one surface with room to spell it out:
the label said "—" while the body next to it asserted zero.

- `WorkingAgentsHoverContent` takes `readonly WorkingAgentSummary[] |
  undefined` and distinguishes all three states: `undefined` renders new
  `agent_activity.unknown_hover` copy, `[]` keeps the empty sentence, a
  non-empty list keeps the roster. The chip passes its projection through
  untouched.
- The colour tier had the same collapse: unknown wore the neutral tier
  WITH muted text, which is exactly the "nothing is happening here" tier a
  known zero wears. `chipAppearance` now takes a `ChipActivity`
  ("unknown" | "none" | "some") instead of a boolean, so the three cases
  cannot be written as two, and unknown stays neutral but undimmed.
- The sub-issues chip is unaffected: it passes a resolved array and
  renders nothing at zero, so it never claimed anything either way.

Regression tests cover the hover path specifically — reverting either
downgrade fails "does not let the hover body downgrade an unresolved
projection to zero", "does not dim the chip while the projection is
unresolved", and the chipAppearance unknown case (verified by reverting).
`WorkingAgentsHoverContent` also gets direct unknown / empty / roster
tests, and `chipActivity` one for the three-way split.

Verified: pnpm typecheck, pnpm test (469 files), pnpm lint (0 errors).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 13:04:42 +08:00
Multica Eve
e6610c0831 fix(usage): close the per-agent rollup windows so the leaderboard cannot exceed the totals (MUL-5551) (#6194)
The Usage page showed a single agent with 1021.0M tokens under a workspace
Tokens KPI of 805.9M for the same 1D window.

Both halves read the same rows and disagreed only on the window.
parseSinceParamInTZ deliberately returns N+1 calendar days of headroom, and
the date-bucketed series (usage/daily, runtime/daily) get trimmed back to
-(days-1) client-side before the KPIs and the chart are computed. The two
per-agent rollups behind the leaderboard carry no date column, so nothing
trimmed them and they kept the full N+1 span: at days=1 that is today PLUS
yesterday. One busy agent's two-day total then trivially exceeded the
workspace's one-day total.

Same defect and same fix already applied to failures/by-agent: switch
usage/by-agent and agent-runtime to parseExactSinceParamInTZ. This also
realigns the Run time / Tasks KPI tiles, which are sourced from
agent-runtime and were therefore a day wider than the Cost / Tokens tiles
beside them.

Co-authored-by: Eve <eve@multica-ai.local>
2026-07-31 12:54:25 +08:00
yushen
0a54485ab6 chore(llm): use gpt-5.6-luna by default 2026-07-31 12:04:01 +08:00
Jiayuan Zhang
0fdc38704e MUL-5149: add agent-generated Chat quick actions (#5766)
* feat(chat): add agent-generated quick actions

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): preserve mid-response quick-action fences

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): drop quick actions on empty reply to keep no_response fallback

An actions-only completion — a quick-actions footer with no visible text —
wrote an empty-content assistant message (message_kind=message). Older
Desktop/mobile clients ignore the quick_actions field and render that as an
empty bubble, breaking the MUL-4351 contract that an empty turn always gives
old clients a visible no_response fallback.

Drop the quick actions when the visible body is empty so an actions-only turn
falls through to the visible no_response outcome, and revert the completion
switch to gate the message row on visible text only. Update the completion
test to pin the corrected behavior.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* feat(chat): generate quick actions via daemon suggestion pass

Replace the in-band runtime-brief instruction with a dedicated post-completion
provider turn: after a direct chat reply finishes, the daemon resumes the same
session with a JSON-only suggest prompt and forwards the raw output on the
complete callback. The server parses it leniently and reuses the existing
sanitize/redact/store/broadcast pipeline; the stripped in-band footer stays as
a fallback for older daemons and pre-upgrade sessions. The footer strip now
covers every chat completion, fixing the intro-turn protocol leak. Adds a
Settings → Chat toggle (client-persisted, default on) that hides the chips.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(chat): deliver quick actions async with skeleton placeholders

Decouple suggestion generation from the turn: the daemon reports completion
immediately (chat:done carries quick_actions_pending as a per-turn capability
signal) and runs the suggestion pass in the background, delivering results
through a new supplement endpoint + chat:quick_actions broadcast. A new turn
on the same session cancels the stale pass. Clients render pill skeletons
under the finished reply until the supplement resolves them (entrance
animation on arrival, 30s safety timeout); older daemons never raise the flag
so no skeleton dangles. Suggest usage re-reports merged totals because
task_usage upserts replace per (task, provider, model). Prompt now asks for
exactly 3 actions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(chat): make the quick-actions toggle stop generation, not hide pills

The Settings → Chat toggle previously only hid rendered pills while the
daemon kept burning a suggestion call every turn. It now travels with each
send (quick_actions_enabled, absent = enabled for older clients), is stamped
on the chat task (migration 213), forwarded on the claim, and gates the
daemon's suggestion pass at the source — no call, no pending flag, no
skeleton. Existing suggestions stay visible; settings copy now says
'generate' instead of 'show'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(migrations): renumber quick-action migrations onto current main

Merging current origin/main brought the vcs migrations to their canonical
216-221 prefixes, which collided with the quick-action migrations that were
sitting at 219/220 (backend CI red in
TestMigrationNumericPrefixesStayUniqueAfterLegacySet). Renumber them to the
next unused prefixes:

- 219_chat_message_quick_actions        -> 222_chat_message_quick_actions
- 220_agent_task_quick_actions_disabled -> 223_agent_task_quick_actions_disabled

Contents are unchanged; sqlc regeneration produces no drift since the added
columns are independent of the vcs tables.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* fix(mobile): render async chat quick actions via chat:quick_actions

The daemon generates quick actions in a background pass after the turn
finishes, delivering them on a separate chat:quick_actions event. Mobile
only handled chat:done (which invalidates + refetches an actions-less
message list) and keeps the messages query at staleTime: Infinity, so an
active mobile session never rendered async-generated quick actions until a
manual pull-to-refresh or refocus.

Add applyChatQuickActionsToCache — mirroring web's patcher — which patches
the supplement onto the targeted assistant message in the flat messages
cache, and subscribe to chat:quick_actions in use-chat-session-realtime.
Patch-only (no invalidate), matching web and mobile's cellular
patch-over-invalidate rule; an empty supplement is a terminal no-op. Covered
by chat-ws-updaters.test.ts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): cancel in-flight messages refetch before quick-actions patch

The chat:done invalidate can leave a messages refetch in flight that read the
assistant row before the daemon persisted the quick actions. If that refetch
resolves after the chat:quick_actions setQueryData patch, it overwrites the
freshly-patched actions with an actions-less row. Both message caches are
staleTime: Infinity, so the overwrite never self-heals and the actions vanish
permanently (MUL-5149, Howard review).

applyChatQuickActionsToCache now awaits cancelQueries for the affected caches
(web: flat messages + messagesPage, mobile: flat messages) before patching, so
a stale in-flight refetch is cancelled and cannot land after the patch. Cancel
must precede setQueryData because cancelQueries reverts to the pre-fetch state.
WS handlers call it via `void` (fire-and-forget).

Adds an active-query race regression test on both web and mobile that holds a
refetch open across the supplement and asserts the patched actions survive;
verified to fail without the cancel.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* feat(chat): quick-actions refresh/regenerate + review hardening (MUL-5149)

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): address quick-actions re-review (MUL-5149)

- Ack alignment: refresh request carries the target message_id; server
  atomically confirms it is still the session's latest turn (409 stale
  otherwise), so the client marker always matches the resolving
  chat:quick_actions — no response reconciliation. Adds a regression test.
- Converge the pending marker on every terminal path: HandleFailedTasks
  (sweeper/orphan) now resolves it, and the daemon reports a failed supplement
  so FailTask resolves it instead of leaving a completed-but-unresolved task.
- Timeout fallback now clears the real query state (useQuickActionsPendingTimeout)
  instead of a component-local flag that only masked the UI; drop the skeleton's
  and pill row's local timers.
- frontend-test type-scale: text-xs -> text-caption. Strip EOF blank line.

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): close quick-actions refresh races and failure feedback (MUL-5149)

Third-round review of the refresh button surfaced three issues; all three
are addressed here.

§1/§2 Session-busy race + concurrent-refresh double-spend: a newer reply
that is queued/running but whose assistant row hasn't landed leaves the old
turn as latest-persisted, so the stale check passes and the regen resumes
the newer provider state — attaching suggestions to the wrong turn. And two
concurrent refreshes each enqueue a quota-spending pass. Add
HasActiveChatTaskForSession and refuse a refresh (ErrChatQuickActionsBusy →
409) whenever the session has any task in flight, checked under the same
session lock as the enqueue so no sibling insert slips past.

§3a Timeout re-arm on surface switch: the pending marker now carries an
absolute expires_at deadline instead of a per-mount timer, so switching
between the floating window and the chat tab resumes the same deadline
rather than restarting a fresh 30s window each remount.

§3b Generation failure masked as success: runChatSuggestPass now returns ok
so an explicit refresh distinguishes a failed pass (didn't start / didn't
complete / timed out) from a completed-but-empty one. On failure the regen
task reports failure, resolveFailedRegenerateQuickActions broadcasts a
FAILED chat:quick_actions, and the client resolves the spinner AND toasts
"couldn't refresh" instead of silently stopping on unchanged pills.

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): count deferred tasks in refresh busy check; solid refresh icon tone (MUL-5149)

Two re-review blockers on 2dc9404d.

§1 (deferred window): HasActiveChatTaskForSession only treated
queued/dispatched/running/waiting_local_directory as in-flight, so a chat
auto-retry armed with a backoff fire_at — inserted 'deferred' by
CreateRetryTask, as provider_network's ~5s final attempt is — slipped past
the busy check. In that window the failed turn has no assistant row yet, so
the old turn is still latest-persisted and refreshable; the regen would then
resume a session the retry is about to advance and pin the new turn's
suggestions onto the old one. Add 'deferred' so the set matches the
canonical in-flight status list the rest of the queries already use
(agent.sql has-active-task checks). New regression test covers a deferred
active turn.

CI (text-contrast gate): the refresh icon button used
text-muted-foreground/70 (transparency standing in for a text tone), which
the frontend-test contrast gate rejects. Switch to the solid
text-faint-foreground token — the tone the gate recommends for icons/glyphs,
already used repo-wide and clearing WCAG 1.4.11.

Co-authored-by: multica-agent <github@multica.ai>

* test(chat): assert quick-actions pending marker carries expires_at (MUL-5149)

The chat:done supplement-flow test still expected the 2-field marker from
before the absolute-deadline change; applyChatDoneToCache now stamps
expires_at, so the deep-equal failed on frontend-test. Assert the deadline is
present (expect.any(Number)) rather than a wall-clock-dependent value — its
timing semantics are covered by the pending-timeout hook.

Co-authored-by: multica-agent <github@multica.ai>

* fix(db): renumber regenerate-quick-actions migration 237 -> 240 (MUL-5149)

main merged Issue Quick Actions (MUL-5465) taking migrations 237/238/239
(quick_action, quick_action_workspace_index, comment_quick_action). This
branch independently took 237 for agent_task_queue.regenerate_quick_actions_for.
The two 237s do not textually conflict (different filenames) so the PR reads
mergeable, but the merged tree would carry two migration 237s. Renumber this
one to 240 so it applies after main's chain. The migration is a standalone
`ALTER TABLE agent_task_queue ADD COLUMN IF NOT EXISTS` — order-independent,
touches a column none of main's migrations reference.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Walt <walt@multica.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Naiyuan Qing <145280634+NevilleQingNY@users.noreply.github.com>
Co-authored-by: NevilleQingNY <nevilleqing@gmail.com>
2026-07-30 21:25:03 +08:00
Jiayuan Zhang
5e3b7a8c37 feat(issues): Issue Quick Actions — preset agent + prompt, one-click from the sidebar (MUL-5465) (#6132)
* feat(issues): Issue Quick Actions — preset agent + prompt, one-click from the sidebar (MUL-5465)

Preset "who to call and what to say" once in Settings, then trigger it from
any issue's sidebar with a single click.

Running one is NOT a new dispatch path. The server renders the prompt, posts
a `quick_action` comment carrying the target's mention markup, and hands off
to the existing comment -> mention -> task trigger. Permission
(canInvokeAgent), attribution, squad-leader routing, the execution log, and
pending-task coalescing are inherited rather than reimplemented — the
MUL-3375 lesson about four drifting copies of one trigger decision.

Three things the UI has to be honest about, because the backend already
decided them:

- One pending task per (issue, agent) is a DB invariant
  (idx_one_pending_task_per_issue_agent). A second click against a busy agent
  starts no new run; the comment merges into the pending task. The toast says
  "Added to Lambda's current run", not "Lambda started working".
- An offline target defers rather than fails; the run reuses the existing
  dispatch.ReasonCode vocabulary instead of inventing one.
- Private agents are deny-by-default with no admin bypass. The sidebar filters
  by the caller's own invoke verdict, so a dead button is never rendered, and
  a direct API call still 403s with `invocation_not_allowed`.

Visibility is DERIVED from the bound agent's permission_mode on every request,
never stored — so it cannot drift after someone flips an agent between private
and public_to. Binding a workspace action to a private agent is allowed (the
alternative pressures people into making agents public just to satisfy a
config constraint) but the settings form says so at bind time, and the
catalog badges it. The target's name is withheld from callers who cannot see
it, so the response never discloses a private agent's existence.

Prompt templating is flat substitution over a closed whitelist. No
conditionals, loops, or filters — the agent already reads the whole issue, so
natural language is the control flow. One optional runtime input ({{input}})
keeps a single action from splitting into five near-identical variants; both
directions of the input/{{input}} agreement are rejected at write time so a
typo can never land silently.

Surfaces: sidebar (top 5, rest behind More), the `/` menu in the comment
composer (inserts the server-rendered body to edit before sending), and
Alt-click for the same hand-off from the sidebar.

Migrations 234-236: quick_action table, its listing index (CONCURRENTLY, own
file), and comment.type + comment.quick_action_id.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(issues): simplify quick action permissions to a stored public/private intent (MUL-5465)

Replaces the derived four-value visibility model with a two-value choice made
at creation, and collapses permission handling to a single check.

The old model computed visibility per request from the bound agent's
permission_mode and used it to filter the sidebar. That filtering was the
problem: two people on one issue saw different sidebars with nothing to
explain the difference, which is harder to debug than a button that tells you
why it refused. It also required the list endpoint to run an invocation-target
query per action per request.

Now:
  - `visibility` is stored INTENT — 'public' or 'private' — chosen up front.
  - A public action must bind a target every workspace member can invoke
    (public_to carrying a workspace target), enforced at write time. So a
    public action is runnable by construction and dead buttons are eliminated
    at the source rather than filtered out later.
  - A private action allows any target and is returned only to its creator.
    That scoping is what the field MEANS, not a permission check.
  - Permission is checked in exactly one place: RunQuickAction. A refusal is a
    structured 403 the client renders as one dialog. The dialog does not
    distinguish "no permission" from "the binding drifted" — the person
    reading it takes the same next step either way, and the person who can fix
    it looks at settings.

Removed: can_run, position + manual ordering (settings sorted by usage while
the sidebar sorted by position — one list, two orders), the derived
visibility_broken flag, the runnable_only projection and its second cache
entry, target_name redaction, the alt-click composer hand-off (the `/` menu
covers insert-then-edit and is discoverable), and the sidebar_limit response
field (now a shared constant).

Ordering is use_count DESC everywhere. Settings shows the target's current
reachability as plain metadata ("Nova · private"), so a public action pointing
at a now-private agent reads as visibly wrong without a bespoke error state.
The tradeoff — no active signal when that drift happens — was accepted
deliberately: drift is rare and the failure is loud at click time.

Migration 234 is edited in place rather than layered, since the PR is
unmerged and the table has never been deployed.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(issues): drop quick action variables and runtime input (MUL-5465)

V1 ships a preset prompt sent verbatim, triggered from the sidebar or the `/`
slash command. Two features are removed and one guard is kept.

Runtime input goes because `/` already covers it. Typing `/code review` drops
the rendered body into the composer, where any part of it can be edited before
sending — strictly more flexible than one fixed field, and the field was
specified before `/` was in V1. Two UIs for one need.

Variables go because none of them passed their own test. The rule was that a
variable earns its place only if it changes what the agent ATTENDS TO, not what
it KNOWS. Checked one by one — {{issue.title}}, {{issue.identifier}},
{{issue.url}}, {{user.name}}, {{date}} — the agent already has every one from
the issue context and from the fact that the comment is authored by the person
who triggered it. They were inherited from autopilot's title template rather
than justified.

The REJECTION survives the feature: any `{{...}}` is refused at write time,
naming the offending token. Someone carrying the habit over would otherwise
have `{{issue.title}}` rendered literally into an agent's instructions and
never notice — the exact silent-typo failure the whitelist existed to prevent.
The check is a fraction of the interpolation engine it replaces and keeps the
door open to enabling variables later without touching stored data.

Removed: 4 columns (input_enabled/label/placeholder/required),
renderQuickActionPrompt + the variable whitelist + quickActionIssueURL, the
two-way {{input}} agreement logic, the run/render `input` parameter, the
variable insert chips, the entire "Ask for input on click" block, and the
sidebar's Popover branch — every row is now a plain button. The settings
dialog drops from six field groups to four.

Migration 234 is edited in place rather than layered, since the PR is unmerged
and the table has never been deployed.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(settings): align Quick Actions with the Labels/Properties list, then fix what the UI review found (MUL-5465)

The tab used a bespoke card list while its two siblings — Labels and
Properties — share one table layout. These three are the workspace's catalog
of small named things and should read as one surface, so Quick Actions now
uses the same structure: search + primary action row, bordered card, responsive
column grid that collapses to stacked rows under `md`, and an overflow menu
instead of a row of icon buttons. Columns are Name / Runs as / Who / Used /
Updated. The tab joins the max-w-5xl group for the same reason.

A UI review pass over the result found five things, four of which are fixed
here:

- The visibility chooser communicated selection through border and background
  only, so a screen reader announced both options identically. Added
  aria-pressed.
- The editor dialog was max-w-xl while both siblings use sm:max-w-lg, and the
  unprefixed cap applied at every breakpoint.
- The empty-state hint diverged from the Properties tab it was copied from
  (text-sm and no max width vs mx-auto max-w-sm text-xs).
- Two hardcoded `text-amber-600 dark:text-amber-400` usages replaced with the
  `text-warning` semantic token, per the repo's design-token rule.

Also fixed a signal-quality bug the review surfaced: the usage column
highlighted anything with use_count 0, so an action was flagged the instant it
was created. Staleness now means "has had time to be used and wasn't" — 90
days since last use, or 90 days since creation for one never used.

Not fixed here: the overflow trigger is size-7 (28px), under the 44px touch
floor. Labels and Properties use the identical size, so changing only this tab
would break the consistency this commit exists to create; it needs one pass
across all three.

Co-authored-by: multica-agent <github@multica.ai>

* fix(issues): drop the quick_action comment type, widen the mention guard, harden the slash race (MUL-5465)

Second review round on PR #6132. All four remaining findings.

**Comment type removed entirely (#2 blocker + #3).** Adding a `quick_action`
type meant dropping and re-adding comment_type_check, and re-adding a CHECK
holds ACCESS EXCLUSIVE on `comment` for a full table scan — a read/write stall
on one of the hottest tables in the product, every deploy. It was also
forgeable: `type` is client-supplied on POST /comments, so any member could
post type='quick_action' and have an ordinary comment render as an action
audit record with its body collapsed out of view.

Both go away by not having the type. A quick action now posts an ORDINARY
comment marked with `quick_action_id`, and the collapsed card keys off that id.
There is no request field for it, so the marker cannot be forged, and the
migration is a bare nullable ADD COLUMN — metadata-only and instant. Verified
against a fresh database: comment_type_check is untouched.

The generic comment endpoint now also validates `type` instead of letting the
DB CHECK reject it. An unknown type surfaced as a 500 on a constraint
violation, which reads as a server fault for plainly bad input; it is a 400
now. `status_change` and `system` are excluded from what a client may author —
claiming those would be forging system narration.

**Member mentions rejected too (#1).** The first pass allowed
`mention://member/...` in prompts on the reasoning that it "only renders a
link". That was wrong: notification_listeners.go adds member mentions to the
recipient set and creates an inbox item, so a saved prompt pinged that person
on every single click. Only `mention://issue/...` reaches nobody and stays
allowed.

**Slash race, properly this time (#4).** The previous fix checked only that the
range still started with "/". Rewriting `/review` into `/fix` while the request
was open passed that check, and the stale response overwrote the new command.
The exact original text is now captured and compared; if the command was
edited, moved, or removed, the pick is abandoned rather than inserted
somewhere wrong. Adds the three regression tests the review asked for:
delayed resolve, rejection, and edit-during-flight.

Co-authored-by: multica-agent <github@multica.ai>

* fix(issues): stop the quick action card repeating its own prompt, and insert the `/` body as markdown (MUL-5465)

Two fixes, one reported and one found while verifying it.

**The card printed the prompt twice.** The collapsed header previewed the
prompt's first line, and expanding showed the mention line plus that same
prompt again. The header now identifies WHICH action ran — "Code Review via
Lambda" — which is both non-redundant and something the body never told you:
the prompt text alone does not say which action produced it. This is what the
original design called for; previewing the prompt was the implementation
drifting from it.

When the action cannot be resolved — deleted, or another member's private one
and so absent from this viewer's catalog — the header falls back to the
prompt's opening line, which is the previous behaviour.

**The `/` menu inserted its body as literal text.** insertContentAt was called
with a plain string, so Tiptap treated the server-rendered markdown as text
rather than parsing it. The mention never became a node; it serialised back out
with escaped brackets (`\[@Lambda\](mention://agent/…)`) and rendered as raw
markup in the thread. Passing `contentType: "markdown"` — the same option the
description editor already uses — parses it properly. Found by reading the
comment rows while checking the first fix: one had escaped brackets and no
quick_action_id, which is what a slash-inserted comment looked like.

The existing async test now asserts the contentType, so the option cannot be
dropped again without failing.

Co-authored-by: multica-agent <github@multica.ai>

* docs(issues): correct the stale quick actions sidebar comment (MUL-5465)

The comment still claimed the section renders nothing when no action is
runnable by the member. Permission filtering was removed several rounds
ago -- the list is deliberately unfiltered and a refusal is explained at
run time -- so the comment described behavior that no longer exists.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(settings): cut the quick action dialog's helper copy in half (MUL-5465)

The dialog had five blocks of explanatory prose around four fields, and
three of them wrapped to two lines, so the form read as a paragraph with
inputs in it.

Each helper now earns its line or loses it:

- The header explained the implementation ("keeps the same history,
  permissions, and execution log as an @mention") -- an architecture note
  the person creating an action does not need. Reduced to the one fact
  they do: it posts a comment.
- "Who can use it" is a question, so the hints answer it as noun phrases
  ("Everyone in the workspace" / "Only you") instead of restating the
  verb. Both now fit one line, which also makes the two cards the same
  height -- the shorter one used to sit in dead space.
- The target and prompt hints front-load the constraint rather than
  burying it mid-sentence.

70 words to 32 across the dialog, with no fact dropped. Field spacing
goes 4 -> 5 so the gap between groups beats the gap inside one.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(issues): render a quick action comment as an ordinary comment (MUL-5465)

The card had a collapsed one-line header that expanded to reveal the
prompt, on the theory that repeated runs of the same action would bury
the discussion. That was solving a problem the feature does not have:
prompts are a sentence or two, the header restated what the body already
said, and the disclosure only put a click between the reader and the
text.

A quick action posts a real comment through the real mention path, so
the honest rendering is the one every other comment gets. Drops
QuickActionCommentBody, its query for the action catalog, and the
now-orphaned quick_action_ran_via string in all four locales.

quick_action_id stays on the comment: it is provenance, and it was never
the reason the card looked different -- keying the special rendering off
it is what is going away, not the record itself.

Co-authored-by: multica-agent <github@multica.ai>

* fix(settings): use the faint tone token for the empty-state icon (MUL-5465)

main added apps/web/app/text-contrast.test.ts, a guard that rejects
transparency standing in for a text tone. The empty-state Zap used
text-muted-foreground/60, which is exactly the pattern it forbids: an
alpha-dimmed tone lands at a different contrast on every surface it is
composited over, so it cannot be reasoned about the way a token can.

text-faint-foreground is the token the guard names for icons and glyphs.

The rule arrived on main after this branch's last merge, so local runs
never saw it -- CI tests the merge commit, which is why only CI caught
it. Merged main first so the branch is checked against the same rules.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 21:01:48 +08:00
Bohan Jiang
44ce16d9b8 MUL-5549: fix(agent): stop reporting a failed model discovery as a real catalog (#6196)
* fix(agent): stop reporting a failed model discovery as a real catalog (MUL-5549)

Selecting the CodeBuddy runtime showed a model list that shares no IDs with
what the CLI actually supports, so every pick was an ID codebuddy rejects
(GH #6180). The list in the report is codebuddyStaticModels() verbatim: the
daemon had fallen back, but nothing downstream could tell.

discoverCodebuddyModels returned (staticModels, nil) on all three failure
paths, and copilot/cursor/grok do the same. A failed discovery therefore
arrived as a successful one, which defeated every guard built to catch it:
the daemon reported status "completed", the picker's discovery_failed hint
only renders on isError, and cacheableModelCatalog — whose own comment says
an empty list means transient failure — waves through a non-empty stand-in
and stores it as last-known-good for the full 24h serve window. One blip got
pinned as the answer for a day.

Discovery now returns a Catalog carrying a Fallback marker, which the daemon
forwards as an additive `fallback` field (older servers ignore it; an older
daemon omitting it keeps the previous behaviour). A fallback catalog is still
rendered — the picker stays populated and manual entry still works — but it
is kept out of both the daemon's 60s discovery cache and the server's catalog
cache. On the server it maps to Keep rather than Drop: a stand-in is no
grounds to evict a real catalog, matching how a `failed` report is treated.

Also stop codebuddyHelpOutput swallowing the exec error. CombinedOutput folds
in stderr, so a codebuddy whose `#!/usr/bin/env node` interpreter is missing
from a GUI-launched daemon's PATH had `env: node: No such file or directory`
parsed as help text — and cached as such for 60s.

Verified against CodeBuddy CLI v2.130.0: the parser itself is fine (16 models
from real --help), so this fixes the reporting of the failure, not the parse.

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): run codebuddy --help at most once per model-list request (MUL-5549)

Review catch on the previous commit. Model discovery and effort discovery both
read `codebuddy --help`, and the effort pass called it independently. That was
free while a failed --help was (wrongly) memoised, but once failures correctly
stopped being cached, the failure path ran the 35s command twice in a single
request — past the server's 60s running timeout, so the request timed out and
the late report was then discarded as stale. The user got nothing, not even the
fallback list the previous commit exists to preserve.

discoverCodebuddyModels now owns the thinking annotation, so the one help
capture feeds both catalogs, and the failure path uses codebuddyFallbackCatalog
to apply the static effort levels without exec'ing at all: whatever broke
--help for the model catalog breaks it for the effort catalog too.

Also strengthen the handler tests. They decoded into a struct declared in the
test rather than calling ReportModelListResult, so a wrong JSON tag or a
mis-wired cache branch would have passed. They now drive the real endpoint with
daemon auth and chi params, covering: a fallback report leaving a previously
discovered catalog intact, an older daemon omitting the field still warming the
cache, and an authoritative empty catalog still dropping the snapshot.

Both fixes are mutation-tested — reverting either makes the new tests fail.

Co-authored-by: multica-agent <github@multica.ai>

* docs(agent): correct codebuddy --help comments after the single-capture refactor (MUL-5549)

Review nit. The comments still described the pre-refactor call graph, where
both discoverCodebuddyModels and codebuddyEffortSuperset called
codebuddyHelpOutput and the cache was what stopped the duplicate run. The
effort parser now takes an already-captured string, and the single-invocation
guarantee is structural rather than cache-dependent — which matters, because a
failed --help is deliberately not cached, so a second caller would re-run the
full 35s timeout.

Also note on codebuddyHelpOutput that it has exactly one caller and why a new
one would reintroduce the bug.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 20:34:01 +08:00
Jiayuan Zhang
f0110da555 feat(inbox): mark a notification unread from the row context menu (MUL-5496) (#6137)
The inbox auto-marks a notification read the moment it is selected, so
"opened" and "handled" were the same signal — a row you glanced at and
meant to come back to was gone from the unread count with no way back.

Right-click any inbox row for a shared context menu: Mark as read /
Mark as unread, plus Archive (Unarchive in the archived view).

- POST /api/inbox/{id}/unread + MarkInboxUnread query, publishing
  inbox:unread. Item-scoped, mirroring mark-read: the list renders one
  row per issue carrying that group's newest item, so flipping the whole
  group would resurrect siblings the user already dealt with.
- useMarkInboxUnread patches both lists optimistically and re-pulls the
  cross-workspace unread summary on settle.
- One shared menu per list rather than a Base UI root per row (the same
  shape IssueContextMenuProvider uses): only one is ever open, and a
  per-row root would unmount with its menu when the row scrolls out of
  the virtualized viewport.
- The read toggle is main-view only — archived rows deliberately render
  as read and the unread count excludes them, so a toggle there would
  report success and change nothing on screen.
- Parking the row that is currently open holds the auto-read effect off
  that one item while it stays selected; re-opening it later marks it
  read again.
- Mobile subscribes to inbox:unread so the unread dots agree across
  clients.

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 19:18:36 +08:00
Bohan Jiang
f73250654f fix(issues): stop blaming permission for an unresolved mention target (MUL-5548) (#6190)
* fix(issues): stop blaming permission for an unresolved mention target (MUL-5548)

A well-formed but wrong agent mention UUID comes back as
`invocation_not_allowed`, and the UI rendered that as "You don't have
permission to use this target". The server never claimed a permission
cause: `invocation_not_allowed` is deliberately ambiguous so a blocked
reason cannot confirm that a private agent in another workspace exists
(dispatch/reason.go). The copy turned a typo into an access-control
investigation — GH #6181 hit exactly this on a squad handoff.

- Reword the blocked-trigger labels in all four locales to name both
  possibilities instead of asserting permission, and record in
  blocked-trigger-copy.ts why a label must not narrow the wire code.
- Report a mention id that is not a valid UUID at all
  (`mention://agent/-`) as `target_unavailable`, matching the squad
  branch beside it and the autopilot admission path. A non-UUID names no
  entity in any workspace, so it conceals nothing and must not be blamed
  on invoke permission.

The well-formed-but-unresolved case is unchanged and still shares
`invocation_not_allowed` with a private agent — the enumeration boundary
this issue asked us to move stays exactly where it is.

Also refresh the multica-mentioning skill: the mention path gates on
`canInvokeAgent`, not `canAccessPrivateAgent` (split in MUL-3963), and
the skill now tells agents to check a mention UUID against the roster
before touching any visibility setting.

Co-authored-by: multica-agent <github@multica.ai>

* docs(skills): correct the mentioning skill's "silent no-op" framing (MUL-5548)

Review nit on #6190: the source map still titled its table "Guards that
make a valid mention a silent no-op", but a parsed mention is never
silently dropped — it is either blocked with a reason_code or folded into
a running task. Several rows in the same table also pointed at
comment.go:14xx line numbers that had drifted.

- Retitle to "Guards and outcomes for a parsed mention" and split the
  outcome into its own column, so each row states the reason_code it
  produces.
- Replace the drifted line numbers with stable search anchors, matching
  the convention the newer rows in this file already use.
- Correct two rows that were wrong, not just stale: archived / no-runtime
  targets are blocked (target_unavailable, runtime_offline), and an
  already-pending target is a coalesce/defer fold, not a skip.
- Apply the same correction to SKILL.md, where the frontmatter and the
  "What does NOT happen" section told agents an already-pending mention
  was dropped. It is folded into the running task and still gets read —
  worth being exact about, since believing otherwise invites a re-post.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 18:42:14 +08:00
Multica Eve
9072cef12c Revert "MUL-5493: feat(chat): add a visible follow-up queue (#6133)" (#6171)
This reverts commit b13657be71.

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 15:37:13 +08:00
TeAmo
b13657be71 MUL-5493: feat(chat): add a visible follow-up queue (#6133)
* feat(chat): add a visible follow-up queue

Add a visible, manageable FIFO follow-up queue for Web and Desktop chat while preserving the existing per-session scheduler and backward-compatible pending-task response.

* fix(chat): preserve queue after deferred cancellation

---------

Co-authored-by: TeAmo <liu.junhui3@iwhalecloud.com>
2026-07-30 15:14:23 +08:00
Bohan Jiang
577018649e feat(issues): support human-readable issue URLs using issue keys (MUL-5354) (#6117)
* feat(issues): support human-readable issue URLs using issue keys (MUL-5354)

Closes #5987. `/{ws}/issues/MUL-123` now opens the issue, the copy-link
action shares that form, and a UUID URL rewrites itself to it. Existing
UUID links keep working.

Backend already resolved identifiers on `GET /api/issues/{id}`, but it
compared the number only — every prefix with the right number opened the
same issue, so no identifier URL could be canonical. Resolution now
validates the prefix against the workspace's own (case-insensitively,
matching `lookupIssueByIdentifier`), and the number parser bails on
int32 overflow instead of truncating a digits-only UUID group into a
plausible issue number.

On the client the identifier stays a presentation concern: the route
resolves it to the UUID before rendering, because the realtime updaters
patch `issueKeys.detail(wsId, issue.id)` with the UUID from the
websocket payload. A view keyed on the identifier would sit on a cache
entry no realtime event can reach and silently stop updating. Resolution
reuses the request the detail view would have made anyway and seeds the
UUID-keyed entry, so an identifier URL costs no extra round trip. The
desktop tab title/status glyph hops through the same resolution for the
same reason.

The URL rewrite lives in the new route wrapper rather than IssueDetail:
the inbox renders IssueDetail in a side panel, where replacing the URL
would navigate the user out of the inbox.

No migration — `issue (workspace_id, number)` is already unique/indexed.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(issues): make the single-request guarantee for identifier URLs explicit

Review flagged that opening `/{ws}/issues/MUL-123` fires two detail
requests. It does not, under the app's own QueryClient — but the
guarantee was resting on something implicit, so make it structural.

The old shape seeded the UUID-keyed entry from a `useEffect` after
resolution, while the route enabled the UUID query in the same render.
That held only because the seed effect happened to be declared before
the UUID query's own effect, and because `createQueryClient` sets
`staleTime: Infinity` so a seeded entry is never refetched. Neither is
obvious from the code, and a diagnostic run under a bare `new
QueryClient()` (staleTime 0) does show two calls — the second being a
staleness refetch of an already-seeded entry, i.e. a harness artifact.

`useCanonicalIssueId` becomes `useCanonicalIssue`, which owns both the
resolution query and the canonical detail query and hands the resolution
response to the latter as `initialData`. That is applied while the
observer is created, so the canonical query never observes an empty
cache and never starts a fetch of its own — no dependency on effect
ordering, and no cache write that could race a realtime patch
(`initialData` is ignored once the entry holds data).

Callers collapse to one hook each: the route no longer runs its own
detail query, and the desktop page drops its duplicate.

Tests now build the client with `createQueryClient()` rather than a bare
`new QueryClient()`, so request-count assertions measure production
behavior instead of the harness, plus a direct assertion that an
identifier URL costs exactly one request.

Co-authored-by: multica-agent <github@multica.ai>

* fix(issues): stop the request loop when an identifier names no issue

Opening `/{ws}/issues/ZZZ-134` never reached "not found". It spun an
unbounded request loop and left the UI on the loading skeleton forever.

The route treated a failed resolution as "nothing resolved" and handed
the raw identifier down to IssueDetail. IssueDetail mounted a second
observer on the query that had just failed; `retryOnMount` refetched it,
which flipped the resolve hook back to pending, which unmounted
IssueDetail, which remounted it when the refetch failed — and around
again. Measured with retry disabled to isolate it: 8,192 requests at
300ms, 32,768 at 600ms. Under the app's `retry: 1` the backoff only
paces the loop, it still never converges.

`useCanonicalIssue` now reports a terminal `notFound` read from the
resolution query's own error state, rather than leaving callers to infer
failure from "not resolving and no id" — an inference that cannot
distinguish failed from in-flight. `IssueDetailRoute` renders the
not-found UI itself and never hands an unresolved segment to a view that
would query it again, so no second observer exists to restart the cycle.
Same measurement after the fix: 1 request, settled, "not found" on
screen.

The not-found UI moves out of IssueDetail into a shared `IssueNotFound`
so both render the identical state.

Regression tests at both levels, with retry off so any count above 1 can
only be a remount refetch: the hook settles a failed resolution without
looping, and the real IssueDetailRoute holds at one request across
waits and rerenders. Both fail against the previous code.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 14:22:13 +08:00
Bohan Jiang
2187b58292 feat(attachments): let callers opt out of pre-signed URLs in bulk responses (MUL-5372) (#6119)
* feat(attachments): let callers opt out of pre-signed URLs in bulk responses

A CloudFront-signed download_url is ~800 chars, ~630 of which are a
Policy+Signature pair re-minted on every request (the policy embeds
now+TTL at second granularity). Emitting one per attachment on every
list response costs three ways: raw payload, a fresh RSA sign per
attachment per request, and — because the bytes differ on every read —
it defeats any cache keyed on response content.

Agents pay all three and use none of it. `multica attachment download
<id>` fetches a fresh signature from the single-attachment endpoint, so
the id is the only part the CLI needs; measured on a 15-attachment
issue, the URL fields were 28.3% of the payload and ~10k chars of that
rotated on every read.

Callers now advertise `stable_attachment_urls` in X-Client-Capabilities
to receive the stable /api/attachments/{id}/download path instead. That
endpoint re-signs and 302s on every hit, so the value stays correct
forever at ~95 chars. The CLI advertises it unconditionally — it is a
protocol detail, not a user-visible flag.

The single-attachment endpoint keeps signing regardless of the
capability: it is the one source of fresh, natively-loadable URLs, and
stable-mode callers exchange a stable path for a signature there. That
carve-out is what makes the rest safe.

The server default never moves. A caller that advertises nothing — every
installed mobile build, every third-party script — gets byte-identical
responses, so this ships ahead of any client and each surface migrates
on its own release cadence. Web/desktop (Phase 2) and mobile (Phase 3)
are deliberately not touched here.

MUL-5372

Co-authored-by: multica-agent <github@multica.ai>

* test(attachments): pin the CLI capability header and the rotation premise

Two review nits from #6119.

The CLI's setHeaders call is the only place Phase 1 is switched on, and
nothing observed it: the server tests prove a request carrying
`X-Client-Capabilities: stable_attachment_urls` gets stable paths, but a
typo in the CLI token — or dropping the header — would have left every
test green while the CLI silently went back to ~800-char signed URLs.
Assert the header verbatim across GET, GET-with-headers, POST, and the
no-token client. Verified the assertion is load-bearing by mutating the
constant to `stable_attachment_url`: the test fails.

TestAttachmentToResponse_SignedModeRotatesButStableDoesNot only checked
the "stable does not" half, so its name overstated coverage. Add the
rotation half. It drives the signer with two explicit expiries rather
than calling attachmentToResponse twice: that function mints its expiry
from time.Now() at second granularity, so consecutive calls usually land
in the same second and asserting on them directly would be a flaky test
of a real property. Also pins that only the query rotates, and that a
fixed expiry re-signs deterministically — without which the rotation
assertion would prove nothing about the clock.

MUL-5372

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 18:25:14 +08:00
Multica Eve
5f820a1d22 perf(agents): raise the model catalog serve window to 24h (MUL-5444) (#6115)
The serve window shipped at 15 minutes, which was calibrated for a
constraint that no longer exists. In the first cut the browser held a
flat 5-minute staleTime decoupled from the server, so the two windows
compounded and a short server window was the only bound on observable
staleness. Since the review round, a server-cached response is stale on
arrival client-side, so nothing compounds.

What 15 minutes actually cost: no background job keeps a snapshot warm —
only a picker open reads or refreshes it — and the browser's react-query
cache dies with the tab. So opening the agent form once today and again
tomorrow was guaranteed to be a cold miss both times, paying the full
1-15s daemon round trip. The cache only helped inside a single session.
Guarding a change that happens on a scale of days (a CLI upgrade) with a
minutes-scale window bought no freshness for that price.

Freshness is not traded away, because the two constants do different
jobs and only the other one governs it: any served snapshot older than
modelCatalogRevalidateAfter (60s, unchanged) queues a background refresh,
and the client treats a cached answer as immediately revalidatable. The
observable behaviour stays "this open shows the old list, the next open
shows the new one" at any window length.

The window stays bounded rather than infinite because a *failed*
discovery report deliberately leaves the snapshot in place — a transient
failure must not empty the picker — so for a runtime whose CLI was
uninstalled or logged out, expiry is the only thing that eventually
retires the stale catalog. 24h covers the dominant daily-usage pattern
while keeping that backstop inside a day.

Both backends derive from the one constant (the in-memory retainFor and
the Redis TTL), so there is no second place to keep in sync. Storage is
not a factor: a measured snapshot is 206 bytes for a Claude catalog and
6.4KB for a codex catalog with full thinking levels and service tiers.

Adds a behavioural regression test instead of restating the constant: an
hours-old and a just-inside-the-window snapshot are still served, a
past-the-window one is a miss, and serving an aged snapshot still
enqueues the refresh plus its wakeup hint. It also fails loudly if the
window is ever dropped back to minutes or if the revalidate threshold
stops being well below it.

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 17:25:20 +08:00
Multica Eve
c25a82eee0 perf(agents): fast model discovery on runtime switch (MUL-5444) (#6098)
* perf(agents): make runtime model discovery fast on runtime switch (MUL-5444)

Switching runtime in the agent creation form left the model picker
spinning for ~8-20s. Two costs stacked up:

- the list-models request sat in the store until the daemon's next
  scheduled heartbeat (0-15s, avg 7.5s of pure dead wait), and
- the daemon then enumerated the catalog locally (static for claude,
  but a CLI/ACP round trip up to ~15s for everyone else).

Both are addressed with the two standard techniques for a slow,
low-frequency, read-only operation: push instead of poll, and
stale-while-revalidate.

Push (removes the heartbeat wait):
- new additive `daemon:pending_work` hint, runtime-scoped, delivered
  through the existing daemon WS hub and the Redis relay so the API node
  holding the socket does the delivery.
- the daemon answers a hint with ONE immediate heartbeat and dispatches
  what it claimed. The hint deliberately carries no work, so nothing has
  to be un-claimed when delivery fails and a duplicate hint cannot
  duplicate work - PopPending stays the atomic claim.
- per-runtime coalescing plus a 1s floor keeps a caller-triggered hint
  from becoming a heartbeat amplifier.

Cache (removes the discovery wait on repeat opens):
- server-side per-runtime catalog cache (in-memory single-node, Redis
  multi-node) written on every successful report.
- a snapshot younger than 15min answers the POST immediately as an
  already-completed request; older than 60s it also enqueues a
  background refresh that only warms the cache.
- only supported, non-empty catalogs are cached; a completed-but-empty
  report invalidates instead, while a failed report keeps serving the
  last known good list.

Frontend: staleTime 60s -> 5min and gcTime 30min, so a runtime revisited
in the same session renders from cache and revalidates in the background
instead of showing the spinner again.

Compatibility: every wire change is additive. Old daemons ignore the
unknown hint type and keep using the scheduled heartbeat; new daemons
against an old server simply never receive one. The cached response is
shaped exactly like a completed live discovery apart from the optional
`cached` / `cached_at` markers.

Co-authored-by: multica-agent <github@multica.ai>

* fix(agents): address review on model discovery SWR (MUL-5444)

Sol-Boy's review on #6098 found the client cache could outlive the
server's own staleness promise, and that the two changed endpoints were
still cast rather than validated.

Must-fix 1 — client freshness now derives from the served answer.
`staleTime` was a flat 5min, so a 14-minute-old snapshot (which the
server returns while queueing its own refresh) was held as fresh for
another 5min: observable staleness became server window + client window,
and the refreshed catalog never reached the tab that triggered the
refresh. `staleTime` is now a function of the query data: a `cached`
answer is stale on arrival (bound stays the server's window alone, and
the next mount/focus picks up the refreshed snapshot), while a live
discovery — which just measured the truth — is trusted for the full 5min
so a cold runtime is never re-enumerated inside one form session.
`gcTime` stays 30min, so a revisited runtime still renders from cache and
revalidates in the background; the pickers gate their spinner on
`isLoading`, which stays false throughout.

Must-fix 2 — both model-discovery responses go through a zod schema.
`POST /api/runtimes/{id}/models` and its poll companion were casting
network JSON to `RuntimeModelListRequest`, which the root CLAUDE.md
API-compatibility rules forbid. Added a lenient schema (`status` stays
`z.string()`, `supported` defaults to true, `.loose()` keeps unknown
fields) plus a fallback record whose `status` is `failed`: a malformed
body now surfaces "discovery failed" with manual entry still usable
instead of a fabricated empty catalog or an endless spinner.
`resolveRuntimeModels` was tightened to match — only an explicit
`completed` is a catalog, so an unrecognised status is an error rather
than a silent empty list, and `supported` can no longer be `undefined`.

Nit — the in-memory catalog cache now deep-copies each entry's
`Thinking` (and its level slice) and `ServiceTiers`, so it delivers the
independent value its comment promises and matches the Redis backend's
JSON round-trip semantics.

Tests: staleTime policy for cached/live/no-data; a QueryObserver test
proving the refreshed catalog reaches the same client with no blank
loading state; unknown-status and omitted-`supported` handling; schema
tests for live, cached, old-backend and nine malformed shapes; client
tests that both endpoints degrade to an explicit failure; nested-field
mutation isolation for the cache.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 16:03:26 +08:00
Bohan Jiang
30318b79bc MUL-5426: fix(daemon): retire sessions whose history the provider refuses to replay (#6083)
* fix(daemon): retire sessions whose history the provider refuses to replay

A run killed mid-reply (machine shutdown, force-quit, SIGKILL) can leave an
empty assistant message in the agent CLI's transcript. Every later resume
replays it, the provider rejects the request, and the (agent, issue) pair is
bricked with no self-healing and no user-facing recovery.

Multica already has the mechanism for this — poisoned-session classification —
but its detector paired "400" with "invalid_request_error", which is the
Anthropic wire shape. The same defect reported by any other provider carried
neither token, so it classified as agent_error.unknown: resume-safe by
omission. GetLastTaskSession kept handing back the dead session on every
follow-up, manual Rerun resolved it through the same predicate, and the
in-turn fresh-session retry never fired because ResumeRejected is false here
(nothing rejected the resume — the transcript loaded and the provider refused
to replay it).

Add taskfailure.UnresumableHistory, which recognises the defect by what the
provider says is wrong — some content is empty, and here is which message in
the history — rather than by status code or provider name. Both signals are
required, so a tool reporting "field must not be empty" does not match.

Wire it into the four places that decide whether a session survives:

- classifyPoisonedError, so the task is written as api_invalid_request
- shouldRetryWithFreshSession, so the turn recovers on all 17 backends
  instead of the subset whose adapter learned to detect it; the tools == 0
  gate is unchanged, so a run that already used a tool is never re-run
- ResumeUnsafeFailure, covering the manual-Rerun path
- both resume queries, as defense-in-depth for hosts whose daemon predates
  this (self-host daemons upgrade on their own cadence)

Fixes #6066. Also covers the daemon half of #5760.

Co-authored-by: multica-agent <github@multica.ai>

* fix(session): close the Chat and fresh-retry paths that resurrect a poisoned session

Review found the previous commit stopped short in two places, both of which
put the dead transcript back in play.

Chat never consulted the guarded query. The claim handler reads
chat_session.session_id first and only falls back to GetLastChatTaskSession
when it is empty, so a poisoned pointer there bypasses every filter that query
applies. The fail path merely declined to OVERWRITE the pointer, leaving it in
place. It now clears it in the same transaction, matched on session and
runtime so a concurrent turn's newer pointer survives. The promote guard moves
to ResumeUnsafeFailure as well — the reason-only check passed an un-upgraded
daemon's agent_error.unknown row and re-pinned what the clear had just removed.

GetLastChatTaskSession also kept the row-level filter the issue query dropped
in GH #5975: it discarded the newest poisoned row and fell back to an older
completed row carrying the same dead session. It now judges each session by
its latest terminal state, matching GetLastTaskSession.

A recovered turn could not retire anything. A terminal report carried one
session_id, and an empty one meant both "nothing to report" and "forget the
old session", so a fresh-session retry that SUCCEEDED left the id it retried
away from selectable — through an older completed row on the issue, or through
the chat pointer. agent_task_queue.retired_session_id records the abandonment
itself, reported on every terminal path including completed, and both resume
lookups exclude it. This is the contract gap the previous PR deferred; the
fresh-retry path now runs on all backends, so deferring it is not safe.

Also narrows what the cross-backend test claims: it pins the shared decision,
not that all 17 adapters surface the error into Result.Error (#5760 is the
counter-example), and says so.

Co-authored-by: multica-agent <github@multica.ai>

* test(session): require pgx.ErrNoRows in the resume-exclusion assertions

The `if err == nil && prior.SessionID.Valid` form these tests shared is
false-green: any real fault — undefined column, syntax error, dead connection
— makes err non-nil, so the condition is false and the test passes. Run
against a database missing this branch's new column, the exclusion tests
reported PASS on a SQLSTATE 42703, meaning they could not have caught a broken
query.

requireSessionExcluded demands pgx.ErrNoRows specifically and fails loudly on
anything else, so a green run now means the filter worked rather than the
query never ran.

Applied to all nine sites, not just the four this branch added: the other five
guard the same GetLastTaskSession exclusion behaviour that this branch
changes, so leaving them false-green would leave the change under-tested. All
nine pass on a correctly migrated database.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 15:54:51 +08:00
Bohan Jiang
607209de7c fix(avatar): serve avatars through a signed endpoint on private buckets (MUL-5393) (#6088)
* fix(avatar): serve avatars through a signed endpoint on private buckets (MUL-5393)

Avatar uploads persisted the raw storage object URL into `avatar_url`. On a
deployment whose bucket is private and has no public CDN domain (S3 with Block
Public Access, R2, MinIO) that URL is a guaranteed 403 in the browser:
ATTACHMENT_DOWNLOAD_MODE only ever applied to the attachment download
endpoint, so every user / agent / squad / workspace avatar rendered broken
even though the upload itself succeeded.

Resolve at read time instead of at upload time. What is persisted stays the
durable object reference, so nothing with a TTL is ever written to the
database and avatars saved by an older build are fixed without a backfill.
What is served is `/api/avatars/<sig>/<key>`, a stable URL the server resolves
per request through the deployment's existing storage download policy
(presigned redirect, CloudFront-signed redirect, or proxied body).

The endpoint is unauthenticated and the HMAC signature is the credential: the
session cookie is SameSite=Strict, so an auth-gated URL cannot be a native
<img src> from Desktop, a mobile webview, or a split-origin self-hosted web
app. The signature covers the storage key and only image extensions resolve,
so an avatar_url pointed at a private document cannot launder it into a
publicly fetchable URL.

Deployments that already work are untouched: a public CDN domain without
per-request signing, and the local-disk backend whose /uploads/* route is
public, both keep returning the raw URL.

Fixes #6024

Co-authored-by: multica-agent <github@multica.ai>

* fix(avatar): only publish avatar-class objects through the signed endpoint (MUL-5393)

Review found that being able to name a storage object was treated as
permission to publish it. `ownedStorageKey` proved only that a URL came from
this deployment's storage, and every image-shaped key was then signed — while
the avatar update endpoints accepted any raw storage URL. A caller who had
seen a private image attachment's URL could submit it as their own avatar, and
the unauthenticated endpoint would keep re-signing it indefinitely. A user
avatar propagates to every workspace that user belongs to, so the leak crossed
workspace boundaries.

Add the missing authorization rule: an object is serveable as an avatar only
when it is avatar-class — a standalone image upload not attached to an issue,
comment, chat session, chat message, or task. The check resolves the backing
attachment row from the id UploadFile embeds in the object filename, so it
needs no lookup by URL and no new index.

It is enforced on both sides. The write side rejects such a value with 403
before anything is stored; the read side re-checks per request, which is what
makes the guarantee hold for rows written before this existed and revokes the
URL if an object is later bound to a comment or chat.

Scope is the `workspaces/` namespace — the only place that can hold content
belonging to someone other than whoever is setting the avatar, covering both
uploads and channel media ingest. Keys elsewhere (the per-user standalone
namespace, or objects an operator placed in the bucket) stay usable, which
keeps the documented "an explicit avatar_url is preserved" contract intact.

Uploader identity is deliberately not part of the rule: duplicating an agent
legitimately reuses the source agent's avatar object, which a different admin
may have uploaded. Publishing someone else's unbound image would require
knowing its URL, and unbound rows appear in no listing endpoint.

Also clamp the 302's cache lifetime to half the signed URL's own TTL (0 ->
no-store). ATTACHMENT_DOWNLOAD_URL_TTL takes any positive duration, so the
fixed 60s could outlive the target it pointed at on a short-TTL deployment.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 15:39:18 +08:00
Bohan Jiang
bdae0d2a03 fix(attachments): serve proxy-mode downloads with a scoped capability URL (MUL-5292) (#6092)
Desktop users on self-hosted deployments saw a save dialog but never got a
file. Electron's native download is a browser-level request: it carries
neither the desktop client's Authorization header nor a session cookie, so
GET /api/attachments/{id}/download answered 401.

The authenticated GET /api/attachments/{id} already exists to hand native
loaders a URL they can fetch without our credentials, and it already does so
in two of three modes -- a CloudFront-signed URL, or an S3 presigned URL.
Proxy mode (local disk, private object host) had no equivalent and kept
returning the auth-gated API path, which is the whole of the bug: one
unfinished branch of an otherwise correct design.

Finish that branch. In proxy mode the already-authenticated metadata endpoint
now mints a capability -- an HMAC-SHA256 signature over (version, attachment
id, expiry) with a key domain-separated from the JWT secret, valid for 60
seconds and scoped to exactly one attachment -- and a separate public route
redeems it. Membership is checked when the capability is minted, never at
redemption; the signature is the proof that check happened.

Nothing moves out of middleware.Auth: the existing authenticated download
route is untouched, so clients that predate this keep working and there is no
second copy of the header/cookie/PAT/task-token resolution. The capability
route always proxy-streams, so it emits no cross-origin redirect and the
signed query cannot leak to a CDN in a Referer.

The capability is site-relative and minted only by GetAttachmentByID. Both
matter: an absolute URL would be picked up by the inline-media re-sign path
and pinned into an <img> far longer than the TTL, and a capability in a list
response would expire before anything used it.

Verification: go test ./internal/handler/ ./cmd/server/ and the
@multica/views editor tests pass; gofmt, go vet, tsc --noEmit clean.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 15:20:13 +08:00
Multica Eve
d4dac0e77c perf(agents): index-backed latest-terminal lookup in task snapshot (MUL-5436) (#6085)
ListWorkspaceAgentTaskSnapshot took each agent's latest completed/failed
task with a workspace-wide DISTINCT ON, so every presence load read and
sorted the workspace's whole terminal history. Neither existing index
matches that shape: (agent_id, status) has no completed_at, and migration
231's (completed_at) partial index is completed_at-first for the Usage
rollups.

Replace the outcome half with a per-agent JOIN LATERAL Top-1 and add a
partial index on (agent_id, completed_at DESC NULLS LAST, created_at DESC,
id DESC) WHERE status IN ('completed','failed'). On a 40-agent workspace
with 200k terminal rows this goes from 6631 shared buffers / 48.3 ms to
162 buffers / 0.1 ms, with an identical row set.

The (created_at, id) tie-break also makes the pick deterministic when
completed_at ties or is NULL — completed_at DESC alone left the winner up
to the plan.

Report #6075 asked to delete the outcome half as dead code, but PR #2608
made the Squad hover card (AgentLivePeekCard) read those rows for its
"last activity" line, so removing them would be a product regression for
shipped desktop builds. The response contract is unchanged here; splitting
the outcome into a lazy endpoint stays follow-up work.

Also tighten pickLatestTerminal to completed/failed only, matching the
snapshot's filter — it accepted cancelled, which the endpoint never
returns and which would have masked an agent's last real outcome.

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 14:28:44 +08:00
Bohan Jiang
cc5ee169f0 fix(agents): allow agent owners to manage their own agent's env (MUL-5438) (#6080)
* fix(agents): let agent owners manage their own agent's env (MUL-5438)

GET/PUT /api/agents/{id}/env admitted only workspace owner/admin, which
made env the one endpoint in the agent permission model that ignored
agent ownership: canManageAgent already lets the owner update/archive,
and canViewAgentSecrets already lets the owner read mcp_config. The
asymmetry was worst on create — POST /api/agents accepts custom_env from
any member and stores that member as owner_id — so a member could write
secrets into their own agent and then never read or rotate them, with
UpdateAgent rejecting custom_env outright and no workaround left.

authorizeAgentEnv now admits a workspace owner/admin OR the agent's own
human owner. The agent-actor rejection still runs first and unchanged:
an agent process is denied even when its backing human owns the target
agent. The owner comparison uses member.UserID rather than the raw
X-User-ID header because agent.owner_id is nullable and uuidToString
renders NULL as "", and canManageAgentEnv rejects an empty owner from
the other side too.

The web Environment tab is gated on the same rule via the existing
canEdit decision, so it stops offering a "Reveal & edit" action that is
a guaranteed 403. The server remains the boundary.

Closes #6076

Fixes: https://github.com/multica-ai/multica/issues/6076
Co-authored-by: multica-agent <github@multica.ai>

* docs(agents): correct stale "owner/admin only" env permission wording

Follow-up to the MUL-5438 permission change: several comments and the
published docs still described the env endpoints as workspace
owner/admin only, which now contradicts the code.

- router.go / agent.go / types/agent.ts: the three sites flagged in
  review.
- agents-create.mdx (en/zh/ja/ko): the user-facing callout said reading
  values requires a workspace owner or admin. It now names the agent's
  own owner first, and spells out that the agent-actor denial holds even
  for an agent the same human owns.
- daemon.go / middleware/auth.go: these called the env endpoints
  "owner-only" as shorthand for "reject agent actors". That property is
  unchanged, but "human-only" is what they actually mean now.

Comments and docs only — no behavior change.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 14:01:09 +08:00
Multica Eve
45e24e3565 fix(comment): @all no longer swallows an explicit @agent mention (MUL-5411) (#6048)
* fix(comment): @all no longer swallows an explicit @agent mention (MUL-5411)

computeCommentAgentTriggers short-circuited on `util.HasMentionAll` before it
looked for explicit mentions, so a comment carrying both `@all` and an
`@agent` / `@squad` mention enqueued nothing at all — the named target never
ran. Evaluate the explicit-mention branch first; `@all` now only suppresses the
implicit routes (assignee / thread parent / conversation), which was its intent.

`all` is neither "agent" nor "squad", so it is still skipped inside
resolveMentionedAgentCommentTriggers and never enqueues a run of its own.

Tests: preview + create coverage for @all + @agent (mentioned agent only, no
assignee fallback), @all + @squad (leader wakes), @all + @member (still
suppressed), plus an end-to-end subtest in the @all suppression integration
test. Refreshed the builtin multica-mentioning skill and its source map, which
documented the old short-circuit and a function name that no longer exists.

Co-authored-by: multica-agent <github@multica.ai>

* fix(comment): malformed mention id no longer panics the trigger resolver

Review finding on PR #6048. MentionRe accepts any `[0-9a-fA-F-]+` id, so
`mention://agent/-` parses as a real mention, and resolveMentionedAgentCommentTriggers
handed it to parseUUID (util.MustParseUUID) — a panic on attacker-controlled
comment text. Preview returned a 500; on the create path the comment row was
already committed before the panic fired.

Parse both the agent and the squad mention id with the error-returning
util.ParseUUID (the convention routeFirstExplicitRootMentionOwner already
follows) and record a blocked outcome instead: agent → invocation_not_allowed,
squad → target_unavailable. Those are the same enumeration-safe codes a
well-formed id that owns no entity already produces, so a malformed id reveals
nothing new and never enqueues.

The panic predated the @all reordering on plain explicit mentions; the reorder
made it reachable for @all + malformed id too, so it is closed here.

Tests: table-driven preview + create regression for a bare `-` agent id, a
short hex agent id, a malformed squad id, and both @all combinations —
asserting no panic, no trigger, and the exact blocked outcome. Verified the
tests fail with the panic before the fix.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 13:15:10 +08:00
Bohan Jiang
2e9a3d0119 fix(dashboard): stop leaking private agents from the per-agent rollups (MUL-5409) (#6051)
* fix(dashboard): stop leaking private agents from the per-agent rollups (MUL-5409)

Three per-agent dashboard endpoints authorized on workspace membership alone
and returned a bare agent_id for every agent in the workspace:

  GET /api/dashboard/usage/by-agent
  GET /api/dashboard/agent-runtime
  GET /api/dashboard/failures/by-agent

That told a plain member which private agents exist, how much they spend, how
long they run and what they fail on. The client already collapsed those rows,
but client-side filtering is decoration — one curl bypasses it.

Server: rows for agents the caller may not view are now folded onto a
`__restricted_agents__` sentinel before serialization, via one shared helper.
Folded, not dropped: each of these responses is the per-agent half of a pair
whose other half (usage/daily, runtime/daily, failures/daily) is workspace-
scoped and unfiltered, so dropping rows would make the per-agent breakdown stop
adding up to the KPIs rendered beside it. The bucket keeps its provider/model
and failure_reason dimensions — both are derivable by subtraction from the
workspace-level series anyway, and the client needs them to price the bucket and
compute its failure rate.

Owner/admin and agent actors short-circuit before any extra query, so the
governance view is unchanged. Hard-deleted agents are deliberately excluded from
the fold — they have no visibility left to protect and keep their own bucket.

Client: fixes the mislabelling that shipped with this. A live private agent was
folded into a row labelled "Deleted agents" with a bin icon, and counted into
the card's "· N deleted" caption — telling the user N agents were deleted when
they are alive and still running. The restricted bucket is now its own row with
neutral copy, keeps its real Time / Tasks values, and counts as neither an agent
nor a deletion in the caption.

Tests: handler regression coverage proving a plain member's response contains no
private agent UUID while every aggregate still sums to the privileged view's
total, plus view coverage for the label and caption.

Co-authored-by: multica-agent <github@multica.ai>

* fix(dashboard): fold hidden system agent carriers into the restricted bucket (MUL-5409)

Review follow-up. The first pass built the restricted set from ListAllAgents,
which filters `kind = 'user'` — so it missed the hidden `kind = 'system'`
execution carriers behind agent-builder sessions.

Those carriers run real tasks and book real usage, and all three rollups
aggregate over agent_task_queue / task_usage with no kind filter of their own.
No list endpoint returns them either (ListAgents / ListAllAgents both filter on
kind), so no client can resolve one to a name. Net effect: the exact two bugs
this PR exists to fix, still live — a bare UUID exposing one member's builder
session (with its spend and failure profile) to every other member, and, once
the agent list loads, a running agent folded into the client's "Deleted agents"
row and counted as a deletion.

restrictedAgentIDs now reads a new ListAllAgentsAnyKind and restricts every
non-user-kind agent for EVERYONE, workspace owner included — nobody can name
one, so a bare UUID row is wrong for every viewer, not just plain members. User
agents keep the per-viewer visibility rule. The invocation-target lookup is
skipped for actors that rule can never restrict (agent actors, owner/admin), so
the added cost is one indexed list query.

Because the bucket now also carries carriers that are nobody's "restricted"
agents, its copy drops to the neutral "Other agents" — the same wording the
Errors card already uses for its equivalent row, in all four locales.

Adds a regression test seeding a kind=system private carrier with tasks and
usage: no endpoint may return its UUID to either the plain member OR the
workspace owner who owns it, a bucket must be present to carry its rows, and
every metric delta (tokens, seconds, tasks, failures, runs) must equal its exact
contribution. Verified to fail on all three endpoints for both viewers with the
kind-filtered query restored.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 13:08:48 +08:00
Multica Eve
cf4114cd5d MUL-5396: validate agent concurrency limits (#6034)
* fix(agent): validate concurrency limits

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): harden concurrency duplication

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 12:47:14 +08:00