Commit Graph

116 Commits

Author SHA1 Message Date
Bohan Jiang
62a2d70afd refactor(daemon): stage-2 — Mentions and Comment Formatting judgment rewrite (MUL-5442) (#6453)
* refactor(daemon): rewrite Mentions and Comment Formatting to their judgment form (MUL-5442)

Stage 2, final two sections, bundled per the small-sections rule.

Mentions: the four-link side-effect table stays verbatim (platform
facts); the two H3 subsections merge into one policy paragraph keeping
every anti-loop anchor — the no-mention default with its cost mechanism,
the no-sign-off-mention ban, the end-with-no-mention rule, the three
mention-warranted cases, and the silence closer. The retired headings'
pins re-anchor to the policy phrases.

Comment Formatting: both variants keep the full operational contract
(file-first sentence verbatim, both bans with incident ids, workdir
scope, --parent continuity, cleanup, newline rule); what goes is the
mechanism narration (what the shell rewrites, how flags get swallowed,
PowerShell version/encoding detail — the consequence stays).

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): correct the PowerShell version claim, pin the mention scope qualifiers, section-scope the formatting assertions (MUL-5442)

Review catches by Elon on #6453, all three accepted:

1. The compressed Windows rationale over-generalized a version-specific
   fact: '$OutputEncoding drops non-ASCII' is true of Windows PowerShell
   5.1 (ASCII default), false of PowerShell 6+ (utf8NoBOM). Now reads
   'Windows PowerShell 5.1 ... may replace non-ASCII characters with ?';
   the Go comment documents the version split and why file-first stays
   version-agnostic (agents cannot rely on which shell services the pipe).
2. The merged Mentions paragraph was pinned only at the list head — the
   scope qualifiers ARE the anti-repeat-notify boundary: 'not yet
   involved', 'for the first time', 'explicitly asks to loop someone in',
   and the loop-cost mechanism are each pinned individually now.
3. The Comment Formatting assertions ran against the whole file, where
   '#4182' also appears in Available Commands — the HEREDOC ban could
   vanish with green tests. The assertions now slice the section (matched
   at the line-start heading, since Available Commands references the
   heading inline) and cover all seven contract elements within it.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-06 12:44:09 +08:00
Bohan Jiang
f44df04f96 refactor(daemon): rewrite the Output section to its judgment form (MUL-5442) (#6425)
Stage 2 section 2. The delivery contract keeps every platform fact and
boundary: comment-add as the only delivery channel, the invisible
terminal, exactly one comment per run before turn exit, --attachment as
the only file path, the runtime-local-path ban with the code-location
form and the say-so-in-words fallback. What goes is derivation: the
invisible-task consequence restatement, the plans-in-your-reasoning
elaboration, the good/bad style examples, the exists-right-now clause,
and two quick-create explanation tails.

One pin re-anchor: 'Do not assume any workspace issue prefix' follows the
rewrite to 'never assume a workspace issue prefix'. Every other Output
and delivery-invariant pin passes unchanged, including the MUL-4899 trio
and the anti-dangling pointer target the Attachments section names.

Issue-kind Output: 1,493 -> 978. Brief (real-UUID fixture):
13,382 -> 12,907 (-475); quick-create variant sheds ~160 more on its own
kind.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 21:15:13 +08:00
Bohan Jiang
2bb13667c6 refactor(daemon): cross-channel command dedup — brief states the loop, per-turn carries the commands (MUL-5442) (#6421)
* refactor(daemon): move the workflow's command templates to the per-turn channel (MUL-5442)

Cross-channel dedup, brief side. MUL-5721 measured that every issue-turn
variant of the per-turn message already carries the ready-to-run commands
with real ids (issue-get line, reading hints, reply cookbook), fresher
than the brief's static copies. The six workflow steps now state the loop
shape — command names and flag mnemonics only; full templates and every
embedded issue UUID leave the brief. The Ownership status commands and the
squad-activity call switch to <issue-id> placeholder form.

Doctrine pins stay on the brief side (mandatory catch-up, scan-first
order, both motivation anchors); command-template pins re-anchor to names
and mnemonics. Two design-consequence test updates: the static-catch-up
assertion repoints at the doctrine (the full command now lives in every
per-turn variant), and the byte-identity non-vacuity guard varies agent
identity instead of issue id — because the brief is now deliberately
issue-id-independent, which opens the cross-issue prefix-cache door.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): point the reply cookbook at Comment Formatting instead of restating it (MUL-5442)

Cross-channel dedup, per-turn side. The cookbook's prose restated the
brief's Comment Formatting rules (file-first rationale, inline-content and
HEREDOC hazards, the full Windows $OutputEncoding mechanics) around the
command that already demonstrates the shape. It now keeps the file-first
order, the ready-to-run command, the literal-newline rule, and the
pointer; the hazard mechanics live once, in the brief section the pointer
names. Default cookbook: 639 -> 494 bytes (-145 per reply turn,
uncacheable channel); the Windows variant sheds ~200 more.

Pins re-anchor from the deleted prose to the surviving anchors (file-first
order, the pointer, the command form); the MUL-2904/#4182 banned-shape
negative guards are untouched.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): narrow the per-turn coverage claim, test the cross-issue invariant, strip the command shape from the guest-visible prohibition (MUL-5442)

Review catches by Elon on #6421 plus the CI failure, all three related:

1. The steps header overstated what the per-turn message carries — it
   ships the issue id and ready-to-run context-read commands; other calls
   are assembled from Available Commands. Reworded to say exactly that (a
   factual claim in a prompt cannot rely on the reader discovering it is
   wrong).
2. The cross-issue byte-equality this PR claims as a design benefit is now
   asserted directly per provider (same stable inputs, different 36-char
   issue UUIDs, byte-identical brief) — a Contains-based negative cannot
   catch truncated/transformed/conditional id use.
3. CI: the squad guest-leader contract test bans any runnable
   'issue status <issue-id> in_review' shape in guest-visible text; the
   placeholder rewrite made the leader dispatch rule's NEGATIVE sentence
   match that ban. The prohibition now states itself without a command
   form ('do NOT move it to in_review or done on this turn') — negative
   sentences should not carry copy-pasteable command shapes at all.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 14:29:01 +08:00
Multica Eve
b98f760f1f MUL-5604: 支持 Reasonix runtime (#6370)
* feat(agent): add Reasonix runtime

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): make Reasonix sandbox host-adaptive

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 14:19:28 +08:00
Bohan Jiang
c3aab05a38 refactor(daemon): rewrite Background Task Safety to its judgment form (MUL-5442) (#6381)
* refactor(daemon): rewrite Background Task Safety to its judgment form (MUL-5442)

Owner-authorized judgment rewrite per the Claude 5 context-engineering
guidance: state the platform fact and the boundaries an agent cannot
infer, drop the enforcement details a frontier model derives.

The section is now three paragraphs: (1) the fact everything derives from
— turn exit is task-terminal, no wakeup exists, never background-and-yield
— with the foreground-collect, synchronous-fallback and no-standing-by
consequences stated once; (2) the external-systems/CI boundary: not
run-owned, the named --watch ban (kept because MUL-5223 proved the
principle alone did not stop CI-watching), merge-gate is not acceptance
criteria, the hand-off phrasing, and the single explicit-ask exception;
(3) the persistent-service handoff contract, review-locked verbatim.

Deliberately dropped as derivable: the run-owned work enumeration, the
tool-promise enumeration, the wait/collect split rule, the
persistent-service scope bullet, auto-merge and snapshot elaborations.

Both pin lists rewritten to the surviving semantic anchors (21 each);
retired pins are listed in the test comments with the rationale.

2,980 -> 1,931 (-1,049). Incident history stays in the Go comment where
it costs no context.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): delete the duplicated CI exception, pin the full ban clause, sync the comment ledger (MUL-5442)

Review catches by Elon on the judgment rewrite, all three accepted:

1. The old exception bullet survived below the handoff paragraph — the CI
   exception appeared twice, the second copy scope-widened by its
   position, and every substring pin stayed green. Deleted; both BTS
   tests now count-guard exactly one 'The one exception' occurrence.
2. The compound watch/poll ban was pinned only by its first command; the
   full clause is the MUL-5223 boundary, so the pin now covers all three
   members.
3. The Go comment ledger still described the pre-rewrite bullet model
   (retired pins as kept, a forward reference that no longer exists).
   Rewritten to the three-paragraph model so the safety-decision history
   reads true.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 13:30:50 +08:00
Bohan Jiang
18f8f78af5 refactor(daemon): per-turn message slimming — OPT-1 UUID dedup + OPT-2 initiator prose (MUL-5721) (#6382)
* refactor(daemon): drop duplicate UUID commands from per-turn comment hints (MUL-5721)

OPT-1 from the MUL-5721 measurement: within one hint, a second full
`multica issue comment list <uuid> ...` command restated the issue UUID
(and, in the resumed hint, both anchors) for no routing value. The warm
hint's issue-wide catch-up and the cold hint's roots scan become flag
swaps on the thread command they follow; the resumed hint loses its
anchor-restating sentence (the read command carries the thread anchor,
the reply cookbook carries the trigger id as --parent).

Fixture-measured: warm hint 555->481, resumed 551->417, cold 429->396;
reply-turn totals -74/-134/-33 B. Ownership prompt and the reply
cookbook (with its concrete --parent command) are untouched.

Every dropped duplicate gets a deny assertion so it cannot silently
return; the demoted fallbacks get semantic anchors in its place.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): compress Task Initiator attribution prose (MUL-5721)

OPT-2: the block's 400 B attribution paragraph carried two derivable
restatements ("in a workspace many people can reach", "so this
attribution does not widen what you can read or write"). The compressed
paragraph keeps every rule: per-person privacy/access rules apply, the
initiator (not the runtime owner) is who you are answering, credentials
stay scoped to the runtime owner, and the initiator may not see
everything you can. Both MUL-2645 test-pinned phrases survive verbatim;
the fact line and heading are untouched.

Fixture-measured: block 513->370 B (member+email form). The two
previously unpinned surviving rules gain anchors ("is who you are
answering", "do not assume the initiator can see everything you can").

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): restore initiator read/write authorization boundary (MUL-5721)

Review must-fix (#6382): the compression treated "attribution does not
widen what you can read or write" as derivable from credential scoping
and dropped it. It is an independent boundary: credential scoping names
whose credentials run, the visibility sentence constrains disclosure —
neither stops the model from treating an initiator's request as extra
authorization. Restored as a short explicit negation with a semantic
anchor.

Initiator block 370 -> 440 B (still -73 vs the original 513); typical
reply turn 2,191 B (net -147 for the PR).

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-05 13:02:51 +08:00
Bohan Jiang
1e661bb04b refactor(daemon): stage-1 judgment rewrite — mode router and small guardrail sections (MUL-5442) (#6365)
* refactor(daemon): compress the mode router and Reply-block preamble to their facts (MUL-5442)

Stage 1 batch A of the judgment rewrite. The mode-router paragraph spent
~800 bytes explaining WHY the file cannot know the turn mode; the fact an
agent needs is only WHERE the mode is named and that exactly one mode block
applies. Compressed to a single Turn-mode lead carrying every routing fact:
both mode lines, the shared-steps note, the one-block rule, and the
no-mode-line default. The Reply block's first bullet drops its derivations
(session history, confusion warnings) and keeps the id-source rule; the
reply-warranted bullet keeps both pinned anchors and every behavioural
outcome, dropping only the worked examples.

Re-pins: the 'Mode router' heading anchor becomes '**Turn mode.**'.
Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): trim the mechanism prose from the small guardrail sections (MUL-5442)

Stage 1 batch B. Comment Formatting (both variants), Issue Body
Formatting, and Always Use the CLI keep every ban, every incident id, and
every pinned phrase verbatim; what goes is the mechanism narration (what
the shell rewrites, why the CLI rejects outside paths, the resource-type
enumeration). Zero test changes — every pin survives as-is.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): restore two platform boundaries and pin the routing rules (MUL-5442)

Stage-1 review catches by Elon on #6365:

1. Two behavioural boundaries had been deleted as explanation. The
   issue-body H1 rule narrowed 'body or description' to 'body' — but
   'description' is the CLI/API field name, so the alias is a
   cross-surface mapping. And step 2 lost 'CLI failures are normal' —
   platform failure semantics (a failed metadata read must not block the
   task), not tool mechanics. The second deletion also went undisclosed
   in the PR description. Both restored; both now pinned.

2. The compressed Turn-mode lead was guarded only by its markers: every
   routing rule (shared steps, apply-exactly-one, status difference,
   no-mode-line fallback, no status change) could have been dropped with
   green tests. All five now pinned as semantic anchors, and both Reply
   outcomes (work-produced arm, silent-exit arm) pinned individually
   alongside the existing heading anchors.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): pin the full no-mode-line mapping and gofmt the touched tests (MUL-5442)

Final-review catches on #6365: runtime_config_kind_test.go:99 was not
gofmt-clean (a scripted insertion), and the no-mode-line fallback was
pinned as two separable halves — 'No mode line' plus a 'Reply mode' that
other file content satisfies — so the text could mutate to route the
fallback at Ownership mode with green tests. The anchor is now the exact
mapping clause 'No mode line → Reply mode'; the status anchor stays.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 18:01:46 +08:00
Bohan Jiang
dc60429366 refactor(daemon): slim Issue Metadata to its judgment core, defer discipline to the skill (MUL-5442) (#6351)
* refactor(daemon): fold the metadata ban list into the write rule (MUL-5442)

The Issue Metadata section spent 1,024 bytes teaching a KV bag. The
What-NOT-to-pin bullet was a separate heading restating the write rule's
negative space; it folds into Write on exit as one sentence with every ban
kept explicitly (secrets/tokens/API keys, logs or comment summaries,
runtime bookkeeping, single-run details). The parenthetical examples and
connective prose go; every rule stays.

Pins updated in kind: the merged ban list is pinned as one sentence plus
the result-comment redirect, so the next compression pass cannot drop a
ban without failing CI. -202 bytes (16,099 -> 15,897 on the standard
fixture).

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): restore the runtime-bookkeeping examples that define the ban (MUL-5442)

Review catch by Elon on #6351: 'runtime bookkeeping' has no other
definition anywhere in the brief or skills, so its examples (attempts, run
timestamps, agent IDs) are the category's boundary, not connective prose —
and all three are high-probability miswrites that can look like re-readable
diagnostics or durable facts. Restored inside the merged bullet.

Also adopts the review's pin advice, which is our own anchor-pin doctrine
applied properly: the whole-sentence pin from the previous commit is
replaced with separate semantic anchors (each ban plus the bookkeeping
examples plus the result-comment redirect), so the sentence can be reworded
later without CI churn while dropping any single ban still fails.

+52 bytes; the section lands at 874 (from 1,024), the PR nets -150.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): defer the metadata write discipline to the working-on-issues skill (MUL-5442)

Owner decision on MUL-5442: metadata is deliberately free-form custom
key-value state — the recommended-keys block never matched the feature's
intent and is removed outright, not relocated. The full ban list defers to
the multica-working-on-issues skill, which already carried a near-complete
copy; the skill gains the two bans it lacked (secrets/tokens/API keys,
agent ids) so nothing loses its home.

The brief keeps only what the interface cannot express: the read-as-hints
stance, the will-a-future-run-re-read-it bar, and the two write-time
boundaries (never secrets, never long content). Section: 874 -> 505 bytes
(1,024 at the round's start).

Pins move with the content: the brief side slims to the surviving
semantics plus the skill pointer; the skill-side contract test gains the
relocated ban anchors so the pointer cannot dangle.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): remove the curated key list from the skill and restore the ban categories (MUL-5442)

Round-3 review catches by Elon on #6351, both accepted:

1. The recommended-keys concept survived in the skill — which loads exactly
   when an agent is about to write metadata, so it still steered free-form
   KV into a platform vocabulary. The owner's ruling was that the concept
   should not exist, not that it should move. The key list, the
   'high-signal keys only' heading, and the pr_url-specific example are
   gone; the section now describes free-form durable custom state and a
   generic set example. A mustNotContain guard keeps the curation from
   creeping back.

2. The relocated ban list had dropped its two defining categories (runtime
   bookkeeping, other single-run details), leaving unlisted values looking
   writable. Restored with the reviewer's structure: each category names
   its examples, and the test pins every category AND every example as
   separate line-safe anchors — no item can be silently dropped again.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): remove the last key-vocabulary residue from the property-vs-metadata note (MUL-5442)

Round-4 review catch by Elon on #6351: the property-vs-metadata bullet
still read 'free-form scratchpad for run state (pr_url, waiting_on, ...)'
— recommending the ruled-out fields AND calling metadata a home for run
state, in direct conflict with the runtime-bookkeeping ban restored two
sections above. Now reads 'free-form bag for durable custom issue state',
consistent with both the owner ruling and the ban list.

Test side per the review: the broad pr_url anchor narrows to the full
stale-warning phrase (the one sanctioned pr_url reference — a negative
compatibility note, not a write recommendation), and the curation guard
gains the removed residue phrases so neither the vocabulary nor the
run-state framing can creep back.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 16:09:05 +08:00
Bohan Jiang
766c7a7fa2 refactor(daemon): retreat deep read semantics to --help, trim BTS prose (MUL-5442) (#6347)
* feat(cli): carry the comment-read contract in comment list --help (MUL-5442)

The --recent flag help gains the MUL-5372 saturation semantics (N caps
threads, not comments; every thread returns uncapped; small issues return
the whole history) and a pointer at the bounded alternative. The --before
help names the stderr cursor labels alongside the response header.

This is the relocation target for the brief's deep read semantics: the
contract follows the flag, and TestIssueCommentListHelpCarriesReadContract
pins it here so the brief-side pointer cannot dangle.

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): retreat deep read semantics to --help and trim BTS prose (MUL-5442)

Unblocked by #6309: the platform no longer recommends --recent anywhere, so
the flag's deep semantics no longer need to ride in every brief.

- Available Commands comment-list entry: keep the signature, the two bounded
  read shapes, and the one-clause saturation warning; per-thread cap detail,
  folding rules, and cursor labels retreat to --help (pinned there by
  TestIssueCommentListHelpCarriesReadContract).
- Workflow step 3: drop the worked examples and the redundant trailing
  pointer; the mandatory two-step read and both motivation pins stay.
- Background Task Safety: tighten connective prose only — every pinned
  phrase and every behavioural constraint is untouched.
- multica-squads SKILL.md quick-start: replace the last remaining --recent
  recommendation on the platform with the bounded scan — same class of fix
  as #6309, found while relocating the warning.

-676 bytes on the standard fixture (16,766 -> 16,090).

Co-authored-by: multica-agent <github@multica.ai>

* fix(cli,daemon): repair the help rendering, restore the BTS orphan scope, complete the squads read (MUL-5442)

Three review catches by Elon on #6347:

1. The --before usage string wrapped the cursor labels in backticks, and
   pflag's UnquoteUsage hijacked the first pair as the flag's value
   placeholder — rendered help showed '--before Next thread cursor' instead
   of '--before string' (same regression class as
   TestLoginTokenHelpOutputRendersCleanly). Labels now use double quotes,
   and the --recent help no longer suggests composing mutually exclusive
   flags ('--roots-only + --thread --tail' -> an explicit two-step read).
   TestIssueCommentListHelpCarriesReadContract now asserts on the RENDERED
   FlagUsages output, pinning '--before string' and the two-step order.

2. The compressed BTS opening said 'anything still running is orphaned',
   which swept externally-owned work (GitHub Actions) into the orphan rule
   the same section later scopes out. Restored: 'any run-owned work still
   active is orphaned', pinned.

3. The squads quick-start's roots-only scan never returns reply bodies,
   where mention triggers and failure reasons usually live. Added the
   bounded drill-down step and the sequence rationale; both reads pinned in
   TestSquadsSkillCoversLeaderRoutingContract, which also guards against a
   regression back to --recent.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-04 14:22:51 +08:00
Bohan Jiang
e92d828f0c refactor(daemon): compress brief prose and demote the sub-issue playbook (MUL-5442) (#6310)
* refactor(daemon): compress brief prose and demote the sub-issue playbook (MUL-5442)

Six static edits to the issue brief, -1,249 bytes on the standard fixture
(17,977 -> 16,728). All edits are identical for every issue run; prompt-cache
byte stability (#6008) is untouched.

- BTS persistent-service bullet: keep the handoff contract (deliverable-only,
  lifecycle detached, readiness verified, best-effort survival), drop the
  operational walkthrough. Pins re-anchored to concepts (durable logs,
  recorded PID, verify readiness).
- BTS CI-ban bullet: keep the command blacklist and the complete-hand-off
  rule, drop the auto-merge/snapshot elaboration.
- Ownership mode: state the identity-forbids clause once on the header
  instead of once per status bullet.
- Delivery invariant: merge the three sub-bullets into the lead paragraph;
  also fixes the stale '(below)' — the per-surface delivery line renders
  above the invariant, not below it.
- Sub-issue Creation: demote the todo/backlog/stage playbook to the
  multica-working-on-issues skill; the brief keeps a one-line flag map plus
  the skill pointer. Skill-side anchors added to
  TestWorkingOnIssuesSkillCoversIssueLoopContracts so the pointer cannot
  dangle.
- Attachments: collapse the two-sentence CLI-fetch intro into one line.

Every pinned behavioral phrase is either carried verbatim or re-pinned to an
equivalent semantic anchor in the same assertion; no assertion is deleted
without a replacement.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): restore the full persistent-service handoff contract (MUL-5442)

Review catch by Elon on #6310: the compressed bullet weakened the contract
two ways — the reply requirement dropped 'logs' from the URL/logs/stop
triple (durable logs alone are unobservable if the user is never told where
they are), and the general ownership/cleanup handle narrowed to a bare PID
(a supervisor/profile-managed service has no single stable PID). Both test
pin sets had been updated to the weakened phrases, which would have made the
regression look legitimate.

Keeps the compressed sentence shape; restores both halves of the contract
and pins them ("cleanup handle such as PID/profile", "URL, logs, and stop
instructions") so they cannot be compressed away again. +38 bytes; the PR
still nets -1,211 against main.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 20:05:07 +08:00
Bohan Jiang
3c6176cbcd refactor(daemon): de-duplicate cross-section rules in the runtime brief (MUL-5442) (#6302)
* refactor(daemon): fold four restatements of the end-of-turn rule into one (MUL-5442)

Background Task Safety opened with four bullets that were four views of the
same rule -- do not end the turn with run-owned work outstanding: the general
ban, the "tool says it will notify you" case, the unobservable-result case,
and the "standing by" sign-off. Separating them cost bytes without adding a
distinct behaviour.

Fold them into the leading bullet. Every phrase the behaviour tests pin is
carried over verbatim, including "Do NOT end your turn while background
tasks", "Never background-and-yield", "wait for a future
notification/reminder", "running in the background so you can keep working",
"run the work synchronously instead" and "standing by".

MUL-5442

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): demote workflow-step restatements of delivery, mention and metadata policy to pointers (MUL-5442)

The six workflow steps and the Reply mode block restated policy that already
has a dedicated section. Each restatement is a second place to edit when the
policy changes, which is how they drifted apart in the first place.

Give each rule one canonical home and leave a pointer at the call site:

- delivery ("only a comment reaches the user") -> ## Output; steps 5 and the
  Reply block point at it.
- mention discipline -> ## Mentions. The reply-time phrasing the loop-hardening
  test pins moves into that section rather than being duplicated in the Reply
  block, so the anti-loop signal is unchanged.
- metadata read/write bar -> ## Issue Metadata; steps 2 and 6 point at it
  instead of paraphrasing the bar.

Tests: the metadata scope test pinned the old pointer wording, and the mention
test's comment claimed the sign-off rule lived in the workflow steps. Both are
updated to the new placement; every behavioural phrase they guard is still
asserted, file-wide.

MUL-5442

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): give the comment-read surface and the file-safety rules one home each (MUL-5442)

Three cross-section duplications, each resolved toward the section that owns
the rule:

- comment reads: the workflow step and the Available Commands entry both
  explained what --roots-only / --summary / --thread --tail do. The step keeps
  the two reads it mandates and its anti-stale motive; flag semantics stay in
  Available Commands, which is the single discovery point. The saturation trap
  ("caps THREADS, not comments") and the pagination cursor labels stay put --
  they are load-bearing after MUL-5372 and remain asserted.
- --content-file: the comment add entry restated the guardrail Comment
  Formatting owns; it now names the rule and points there for the rationale.
- workdir path rule (MUL-4252): issue create carried its own copy of the
  stale-file rationale; it keeps the rule and defers the why.
- inbound attachments: trimmed to the pinned rule plus a pointer.

MUL-5442

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): merge the Agent Identity action list, gate squad maintenance, trim the attachment restatement (MUL-5442)

Three follow-ups from review, each re-examined against what the text guards
today rather than against the fact that a test pinned it.

1. Agent Identity enumeration. Instruction Precedence and workflow step 4 were
   added in the same commit (#3802) and each carried its own list of actions
   Agent Identity can forbid -- and the lists disagreed: one named status
   changes, the other named issue create/update and delegation, neither
   contained the other. Merge them into Instruction Precedence, which owns the
   rule. Step 4 keeps only what that section cannot express: a delegation-only
   role stops once its delegation is delivered.

2. Squad maintenance. `multica squad member set-role` shipped to every run,
   including every agent that leads no squad and therefore has no squad whose
   roles it could change. Gate it on IsSquadLeader -- agent configuration, not
   per-run state, so the brief stays byte-stable across runs of one session
   (MUL-5377), the same predicate the workflow already branches on.

3. Inbound attachments. The section restated Output's no-clickable-local-path
   rule verbatim. Keep the framing Output cannot express -- a downloaded
   attachment feels shared but landed in a private workdir -- and point at
   Output for the rule. The delivery test now also asserts the pointed-at rule
   is present, so the pointer cannot dangle.

Brief size, plain issue task: 19,231 -> 18,009 bytes for an ordinary agent
(-6.4%), 20,718 -> 19,759 for a squad leader (-4.6%).

MUL-5442

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 19:00:34 +08:00
Multica Eve
a54662f30e fix(daemon): provision config-referenced Codex instruction files into task home (#6291)
* fix(daemon): provision config-referenced Codex instruction files into task home

A per-task CODEX_HOME copies the user's config.toml verbatim, so a
model_instructions_file reference survived the move while the file it
names did not. Codex resolves the relative value against CODEX_HOME —
now the task home — and failed loading its configuration before the task
prompt was ever delivered (#6271).

Generalize the existing model_catalog_json provisioning into a keyed
table of path-valued config keys and add model_instructions_file plus its
deprecated experimental_instructions_file alias. Semantics are unchanged
per key: absolute/~ values are left for Codex to read directly, relative
values must stay inside the task home, the copy refreshes on reuse, and a
missing source fails during environment preparation with a diagnostic
naming the key instead of an opaque os error 2 at thread/start.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): root-scope task-home writes for config-referenced Codex files

Review must-fix: filepath.IsLocal only proves the config string has no
lexical '..'. A task home is reused, and the task that ran in it can
replace an intermediate directory of the copy destination with a symlink,
which made the daemon's next MkdirAll/Remove/copy follow that link and
delete or overwrite a file outside the task home.

Do the mkdir, stale-copy removal, and write through os.OpenRoot(codexHome)
so links leaving the task home are rejected (links staying inside it are
harmless), and refuse outright when codexHome itself is a symlink, since
OpenRoot would resolve it before confining anything below. The pre-existing
model_catalog_json path is covered by the same helper.

Also stop reporting every source stat failure as a missing file — a
permission or IO error now says so — and document the source-side symlink
policy plus the stale-copy contract.

Co-authored-by: multica-agent <github@multica.ai>

* test(daemon): pin stale-copy contract when a referenced Codex file is repointed

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): bind codex-home root check to the opened handle

Review must-fix: checking the path with os.Lstat and then opening it with
os.OpenRoot leaves a window where the directory is swapped for a link to
somewhere else. OpenRoot resolves the path it is given, so the resulting
root is confined to the wrong tree and every root-scoped write lands
outside the task home without ever escaping its root. Task homes are
reused and Windows cannot confirm descendant cleanup, so a leftover
process that knows its old CODEX_HOME can create that window.

Open first, then prove identity against the handle: compare root.Stat(".")
with a no-follow os.Lstat of codexHome via os.SameFile, and reject a
symlink outright. A swap before the open now fails the check; a swap after
it cannot matter because all writes go through the verified handle.

verifyCodexHomeRoot is split out so the swap is tested deterministically
instead of raced, plus an end-to-end test for a task home that is already
a link to an outside directory.

Co-authored-by: multica-agent <github@multica.ai>

* docs(daemon): scope the codex-home symlink test name and root-handle claim

Review nit: the end-to-end test asserted only that the referenced-file copy
refuses a symlinked task home, but its name read as though the whole
prepare were safe. Rename it accordingly and state in both the test and
openVerifiedCodexHomeRoot that the earlier path-addressed steps of
prepareCodexHomeWithOpts are out of scope, tracked in MUL-5647.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 16:23:43 +08:00
Multica Eve
e4b6f7a31b MUL-5581: add Qoder CN CLI runtime (#6232)
* feat(agent): add Qoder CN CLI runtime

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): address Qoder CN review nits

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): defer Qoder CN version gate

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 16:01:53 +08:00
Bohan Jiang
98766cc06a feat(daemon): remove the Linux Codex per-task HOME; default Linux to danger-full-access (MUL-5578) (#6233)
* feat(daemon): default Linux Codex to danger-full-access on the real HOME (MUL-5578)

Linux Codex tasks ran under the `workspace-write` Landlock sandbox with a
generated per-task HOME: the daemon rewrote HOME/XDG_*/npm_config_cache into
`<envRoot>/home`, symlinked a hand-maintained allowlist of seven credential
paths back into it, and granted that directory as a `writable_roots` entry.

That mechanism could not converge. Any host CLI outside the allowlist (aws,
kubectl, gcloud, glab, rclone, cargo, …) started every task as if unconfigured
even though it works in the daemon user's shell, and three open issues asked to
extend it in three incompatible directions (#5636, #5573, #3867). The
containment it bought was also narrower than it looked: workspace-write
restricts writes only — reads and network were already unrestricted, and `.ssh`
was seeded into the task home — so credentials under the real HOME were
readable and exfiltratable regardless.

Linux now runs `danger-full-access` on the daemon user's real HOME and inherited
XDG environment, matching macOS and Windows. The task filesystem boundary is the
boundary the daemon itself runs inside (VM, container, or dedicated Unix user),
which is what the other providers already assumed — Claude Code runs with
`--permission-mode bypassPermissions` today.

This also removes the split-brain the mechanism could enter: the task-HOME
decision read the platform default only, so a `-c sandbox_mode=danger-full-access`
override (which passes arg filtering and wins over config.toml) left a task
running unsandboxed while the daemon still redirected HOME and emitted
writable_roots for a sandbox that was not in effect. With no HOME rewrite there
is only one environment contract left to disagree about.

Removed: prepareTaskHome, prepareCodexSandboxHome, TaskHomeEnv, Env.TaskHome,
both seed allowlists, and the now-unfed WritableRoots plumbing through
codexSandboxPolicy / CodexHomeOptions. Env roots created by older daemons keep
working; their leftover `home/` directory is simply ignored and reclaimed with
the env root.

Task-scoped CODEX_HOME is untouched — it is managed Codex state, not the Unix
HOME. macOS, Windows, and non-Codex providers are unchanged.

Co-authored-by: multica-agent <github@multica.ai>

* docs: document the agent execution security model (MUL-5578)

The docs had no page describing what a task can reach on the machine that runs
it — the sandbox posture was only discoverable from source. Adds one, stating
plainly that tasks run with the full permissions of the daemon user and that
isolation must come from a dedicated Unix user, container, or VM.

Also separates what Multica genuinely isolates (per-task workdir, task-scoped
CODEX_HOME, agent+task-bound API tokens) from what is not a boundary (the coding
tool's own sandbox and approval settings), and links it from step 5 of the
self-host quickstart, where the daemon is first installed.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): address review on the Linux full-access default (MUL-5578)

Docs, both must-fix:

- The Chinese page rendered the product concept as `Agent`; the repo glossary
  in developers/conventions.zh.mdx mandates 智能体. Fixed all nine occurrences.
  ja/ko already used エージェント / 에이전트 and needed no change.
- The security page asserted without qualification that every task runs with
  the daemon user's full permissions and that the Codex filesystem sandbox is
  always off. codexSandboxPolicyForWindows still keeps workspace-write when a
  user explicitly opts into a native windows.sandbox, so an authoritative
  security page contradicted the code. It now states that Multica makes no
  filesystem-sandbox guarantee, names the Windows opt-in as the one current
  exception, and says which combinations sandbox anything is a compatibility
  detail that moves with tool versions. All four locales updated, plus the
  historical callout which now says Linux matches the macOS/Windows *default*.

Both stale comments from the non-blocking notes:

- prepareCodexHome no longer claims to assume workspace-write + network_access;
  it pins GOOS=linux, which now resolves to danger-full-access.
- ensureCodexSandboxConfig's warn-level logging is no longer described as a
  macOS-only fallback; it fires for every danger-full-access resolution.

Adds TestCodexTaskShellEnvInheritsRealHome at the daemon env-assembly layer:
HOME and the XDG base dirs must reach a Codex task's shell tools from the
inherited daemon environment. Verified it fails when that pass-through breaks.
It guards the pass-through, not runTask's decision not to inject a HOME of its
own — that decision is inline in runTask and has no unit seam.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 15:42:06 +08:00
Bohan Jiang
abdfd3e28c refactor(skills): make the brief's skill list a names-only index (MUL-5529) (#6207)
* refactor(skills): make the brief's skill list a names-only index (MUL-5529)

Step 3 of MUL-5529. Every runtime CLI discovers the SKILL.md files the daemon
writes and builds its own listing from their frontmatter — verified against 11
locally installed CLIs plus official docs for 5 more. The brief's copy of those
descriptions was therefore the same routing signal paid for twice: measured on
a real task, `## Skills` was 13,295 chars, 40% of the entire brief, against a
16,304-char CLI listing of the same 28 skills.

Now 850 chars for that same set — roughly 3,100 tokens back per brief.

The index itself stays. It is the one skill listing Multica controls; each
CLI's own listing is theirs, and its format — or its existence — can change
with any release.

Three changes:

  - Descriptions dropped from the `## Skills` entries.

  - The per-provider branch is gone. Its fallback told providers outside a
    hardcoded list to read `.agent_context/skills/`, but the only providers
    that ever reached it were grok and traecli, whose files are written to
    `.grok/skills` and `.traecli/skills` and which discover natively. The
    pointer was wrong for everyone it addressed, so removing the branch
    deletes the bug rather than relocating it. This closes MUL-5537.

  - issue_context.md and its quick-create / autopilot variants no longer render
    `## Agent Skills`. That copy duplicated the brief once both were
    names-only, and nothing ever read it: no prompt references the path, and
    grepping the server finds only the writer. `.agent_context/skills/` had the
    same fate for hermes (issue #5242). Quick-create, previously skipped in the
    brief and served only by that unread copy, now gets the brief section like
    every other kind — one index, one place.

Not included: skills carrying `disable-model-invocation` are still written to
disk for every provider. The plan assumed that key needed provider-specific
handling for everything except claude; probing the installed CLIs shows 9 of 11
honor it, and only opencode and hermes do not. The remaining question is
narrow and a genuine product tradeoff — withholding the file honors the
author's intent but also removes explicit invocation — so it is left to a
separate decision rather than folded in here.

Co-authored-by: multica-agent <github@multica.ai>

* docs(skills): align stale comments with the names-only brief contract (MUL-5529)

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Steve Jobs (Multica Agent) <agent-steve-jobs@multica.ai>
2026-07-31 13:27:50 +08:00
Bohan Jiang
5199278780 fix(skills): give every skill one name across brief, directory, and frontmatter (MUL-5529) (#6189)
* fix(skills): give every skill one name across brief, directory, and frontmatter (MUL-5529)

A skill could answer to three different names at once. The runtime brief listed
`AgentSkillData.Name` verbatim — a workspace skill's human display name ("PR
review") — while the invocable identity on disk is the sanitized slug
(`pr-review`). Separately, ensureSkillFrontmatter returned valid upstream
frontmatter untouched, so a SKILL.md could declare `name: multica-dev-workflow`
inside a directory called `multica-git-workflow`.

That last divergence is the sharp one: runtimes disagree on which field
identifies a skill. Claude routes on the directory name, OpenCode on the
frontmatter `name`. So the same skill is invocable under different names
depending on where it runs, and the brief's instruction to use "only names from
the listing" pointed at names that resolve nowhere.

The slug is authoritative: it is what lands on disk, it derives from the name
users see in the product, and it is the only value with a uniqueness guarantee
(allocateCollisionFreeSkillDir). A frontmatter `name` is author-supplied and two
imported skills may both claim the same one.

- modelVisibleSkills normalizes Name to the slug. All four model-visible
  listings (runtime brief + the three issue_context renderers) already route
  through it, so they cannot drift apart.
- ensureSkillFrontmatter rewrites `name` to the allocated slug and keeps every
  other key byte-identical, so deliberately shaped upstream frontmatter still
  survives. The rewrite follows the collision fallback slug too.
- Name matching is now top-level only. An indented `name:` belongs to a nested
  mapping; treating it as the skill's identity both missed that the block had no
  top-level name and would have spliced a top-level key into the nested one.

Known gap, tracked separately: the listings render the natural slug, so a
collision fallback to `<slug>-multica` still leaves the brief naming the user's
skill. Closing it needs the allocated slug threaded back from Prepare, which the
renderers cannot reach without giving up the byte-identical-brief guarantee.

Co-authored-by: multica-agent <github@multica.ai>

* fix(skills): cover multi-line name values and in-batch slug collisions (MUL-5529)

Two holes in the name-unification change, both found in review, both
reproduced before fixing.

1. setFrontmatterName replaced only the `name:` line, not the rest of the YAML
   value. A value may continue onto indented lines — `name: >-\n  upstream`,
   multi-line plain scalars, wrapped quoted scalars — so the continuation
   survived and YAML folded it into the new value: the block parsed as
   "my-slug upstream-name", not "my-slug". The directory == frontmatter-name
   invariant this change set exists to establish was still broken, just less
   visibly. frontmatterNameSpan now covers the whole value.

   As a side effect this also fixes `name:` with the value entirely on the
   following line, which previously read as "no name" and got a second
   top-level `name` injected above it — two `name` keys, which strict loaders
   reject outright.

2. The listings derived slugs with sanitizeSkillName alone, which is not
   injective: "A B" and "A-B" both reduce to "a-b". writeSkillFiles resolved
   that at write time, so the second skill landed in `a-b-multica` while both
   were listed as `a-b` — the second skill had no invocable name and the model
   was pointed at the first. This needed no user-installed skill and no
   local_directory; two such skills bound to one agent reproduce it in a clean
   workdir. resolveSkillSlugs now deduplicates the batch up front and both the
   listings and the writer derive from it, with skillSlugCandidate shared so
   the in-memory and filesystem allocators cannot disagree on the suffix
   sequence.

   Slugs are allocated over the unfiltered batch: hidden
   (disable-model-invocation) skills are still written to disk and still
   consume a slug, so filtering first would shift every later suffix.

Filesystem-dependent collisions against user-installed directories remain out
of scope and are still tracked in MUL-5550.

Regression tests assert the parsed YAML value rather than the output text —
the first line looked correct in every one of these cases — and set-equality
between listed names and the directories actually written. Both were confirmed
to fail against the previous implementation.

Co-authored-by: multica-agent <github@multica.ai>

* fix(skills): bound the frontmatter name value with the YAML parser (MUL-5529)

Review round two found a valid multi-line name the indentation rule still
mangled. A quoted scalar may wrap onto a line at the *same* indentation as its
key:

    name: "upstream
    continued"

The rule stopped at the key's line, stranding `continued"` and producing
invalid YAML. Probing the parser showed the same holds for flow collections
(`name: [a,\nb]`, `name: {a: 1,\nb: 2}`), so this was not a quoting special
case: indentation simply does not bound a YAML value, and no amount of
patching the heuristic would have made it one.

Value extent now comes from yaml.v3's own line numbers. frontmatterNameValueSpan
locates the `name` key node and ends the span where the next top-level key
begins, stepping back over blank lines and unindented comments so they survive
the rewrite. Unindented is the operative word: block scalar content is always
indented, so an unindented `#` can only be a comment, while `  # text` inside a
block scalar is value and stays in the span.

Detection stays lexical, in lexicalFrontmatterNameSpan. It runs on malformed
blocks too, where there is no parse to consult, and only decides which branch
to take.

setFrontmatterName now re-parses its own output and returns verified=false
unless `name` really is the slug; ensureSkillFrontmatter then routes to the
existing re-synthesis path rather than emitting a block it cannot vouch for.
This adds no new fallback — it feeds an already-present one. The invariant is
the whole point of the change, so an unprovable rewrite is worth less than a
reformatted block: with the span deliberately broken, output stays valid YAML
carrying the right name and loses only upstream formatting.

Tests extend the table to same-indent single/double quoted scalars, flow
sequences and mappings, and name-as-last-key, and now assert the *input* parses
so a case cannot pass by silently taking the re-synthesis path. Two more cover
comment survival and a `#` line inside a block scalar name. All four
same-indent cases fail against the previous implementation.

Co-authored-by: multica-agent <github@multica.ai>

* fix(skills): preserve policy keys when the surgical name rewrite fails (MUL-5529)

Review round three: the post-condition check added last round was routing valid
YAML into bare re-synthesis, and re-synthesis emits only name and description.

An anchor on the name value is one way to get there. Replacing the line drops
`&skill_name`, so `description: *skill_name` no longer resolves, the check
correctly rejects the rewrite — and the block was then rebuilt as just:

    name: my-slug

`disable-model-invocation: true` went with it. That key is the author's
instruction that a runtime must not surface the skill on its own, and the
SKILL.md we write is what native discovery reads, so losing it advertises a
skill that was deliberately hidden. Confirmed on the written output:
skillDisablesModelInvocation went from true to false. That is a semantic
regression, not the formatting loss the fallback was justified by.

The two failure modes are now separate:

  - invalid YAML → re-synthesize, unchanged; there is nothing to preserve.
  - valid YAML, rewrite unprovable → renameFrontmatterNameViaNode rebuilds the
    block from the parsed node with `name` set to the slug, keeping every other
    key. Formatting normalizes; semantics survive.

The anchor is deliberately kept on the name node. Dropping it is what
invalidates a document that aliases it; keeping it means the alias resolves to
the slug — that value changes, but the document still loads and every policy
key is intact. The rebuilt block is re-parsed and checked like the surgical
path, so an unprovable result still falls through to re-synthesis.

Regression test asserts name, disable-model-invocation, and a custom key all
survive, and that the written file still reads as hidden to
skillDisablesModelInvocation. It fails without the new path.

Co-authored-by: multica-agent <github@multica.ai>

* fix(skills): materialize aliases before renaming an anchored name (MUL-5529)

Round three kept the anchor on the name node so aliases would not dangle. That
avoided an invalid document but created a worse one: every alias resolves
through the anchor, so renaming the anchored value silently rewrote whatever
those fields meant.

    name: &shared "true"
    disable-model-invocation: *shared

skillDisablesModelInvocation went true -> false across the rewrite. Same
skill-exposing regression as round three, reached by keeping the key instead of
dropping it — which is the lesson: preserving a key is not the invariant,
preserving each key's resolved value is.

Aliases pointing at the name node are now materialized to the value they
resolved to *before* the rename, and the anchor is dropped afterwards as
unreferenced. Ordering matters: the clones are taken first, so they capture the
original value rather than the slug. Copies are per-alias, since sharing one
node would make the encoder re-emit an anchor/alias pair.

Nested aliases are covered by walking the whole document, not just the top
mapping.

Tests: three shapes (alias carrying the policy value, alias nested in another
mapping, two aliases of one anchor) each assert the fixture starts hidden and
stays hidden, that no anchor or alias survives, and that name is the slug. The
round-three test now also pins its aliased `description` to the pre-rename
value instead of merely tolerating the slug. All four fail without the change.

Co-authored-by: multica-agent <github@multica.ai>

* fix(skills): let the parser decide whether a name key exists (MUL-5529)

The branch that chooses between "rewrite the name" and "inject a name" was
still gated on a lexical scan, which only recognizes a bare `name:` carrying a
value on the same line. Two valid spellings therefore read as nameless:

    "name": upstream        # quoting is syntax, not identity
    name:                   # a key with no value is still the key

Both got a second `name` injected above the existing one, and a duplicate
mapping key is rejected outright — `mapping key "name" already defined`. The
skill does not end up misnamed, it fails to load. Single-quoted keys have the
same problem.

For valid YAML the parsed top-level mapping now answers the question, so any
spelling of the key routes to the rewrite. The lexical scan is confined to the
invalid-YAML branch, where there is no parse to consult and it is only choosing
between re-synthesis and injection. Nesting still reads as absent: a `name`
under another mapping is not the skill's name, so one is added.

Once past the gate the existing paths handle both shapes unchanged, since
frontmatterNameValueSpan already bounds the entry by node line numbers rather
than by how the key is written.

Also tightens the alias tests per review: the nested and two-alias cases now
assert `meta.inner` and `other` still resolve to the anchor's original value,
not merely that visibility survived. The invariant is every key's resolved
value, so every alias should be pinned, not just the one that gates hiding.

Co-authored-by: multica-agent <github@multica.ai>

* fix(skills): detect quoted and valueless name keys in malformed frontmatter (MUL-5529)

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Steve Jobs (Multica Agent) <agent-steve-jobs@multica.ai>
2026-07-30 21:21:13 +08:00
Bohan Jiang
e45a8f6c12 fix(daemon): scan comment roots before bulk reads in agent catch-up (MUL-5372) (#6093)
* fix(daemon): scan comment roots before bulk reads in agent catch-up

The mandatory step-3 catch-up in the issue runtime brief asked for
`--recent 10`. `--recent N` caps THREADS, not comments: each returned
thread carries its root plus every descendant with no depth bound, so on
an issue with fewer than N root threads it returns the entire comment
history. Because the step is mandatory and fires on every run, every
reply turn re-read the whole issue -- and on comment-triggered turns it
duplicated the bounded thread read the per-turn message had already
pointed at, so the same bytes were fetched twice.

Lead the step with `--roots-only --summary` instead: every top-level
thread with reply_count and last_activity_at, contents clipped. That
keeps the property the step exists for -- the agent still sees every
thread that exists, so it cannot act on stale context -- and makes the
drill-down into `--thread <id> --tail 30` explicit. `--recent 10` stays
documented for when several complete threads really are needed, now with
its saturation semantics spelled out.

Measured on a live 2-thread issue: 21,249 -> 1,518 bytes for the
mandatory read (-93%), and the duplicate 11,082-byte thread read is gone.

The brief stays byte-identical across runs of a session (MUL-5377): the
new text interpolates only the issue id, no per-run state. The three
per-turn pointers that express the same rule move with it so the two
layers cannot drift.

MUL-5372

Co-authored-by: multica-agent <github@multica.ai>

* refactor(daemon): keep comment-read flag semantics in one place

The previous commit fixed the payload shape but restated the read surface
in four places: the workflow step, both per-turn prompt fallbacks, and the
cold-start hint each explained what `--recent 10` does. `## Available
Commands` is already the brief's single discovery point for these flags,
and `TestInjectRuntimeConfigStaticCatchUp` pins it as such -- so those
restatements were duplicated reference text, and the per-turn ones were
paid on every turn rather than once in the cached prefix.

Move the `--recent N` saturation warning into the `comment list` line in
Available Commands, next to the flags it qualifies, and add `--roots-only`
and `--summary` to that signature so the bounding options are discoverable
where an agent already looks. Workflow steps and per-turn hints now name
only the reads they actually want run.

Per-turn prompt sizes: assignment 1170 -> 749 bytes (-36%), cold-start
comment turn 1550 -> 1355 (-13%). Step 3 is 1065 bytes and no longer
carries a ready-to-paste bulk read.

MUL-5372

Co-authored-by: multica-agent <github@multica.ai>

* docs(daemon): address review nits on comment-catchup change

Three cosmetic follow-ups from review:

- `--recent N` saturation warning said it hands back "the entire
  history"; resolved threads are still folded by default on that read, so
  say so.
- Rename two tests whose names still advertised `--recent` after their
  assertions stopped mentioning it, plus the one added in this branch
  whose name referenced a bulk read the step no longer contains:
  MentionsRecent -> ScansRootsFirst, ScansRootsBeforeBulkRead ->
  ScansRootsFirst.
- Fix the stale doc comment that still described the mandatory read as
  bounded to "the recent active-thread window".

No behavior change.

MUL-5372

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 16:08:55 +08:00
Bohan Jiang
4eb80c8b77 fix(daemon): keep the runtime brief byte-stable across triggers (MUL-5377) (#6021)
* fix(daemon): keep the runtime brief byte-stable across triggers (MUL-5377)

Claude Code loads the runtime brief (CLAUDE.md / AGENTS.md) into messages[0],
ahead of the entire conversation. A cache breakpoint is all-or-nothing, so a
single differing byte there invalidates the prompt cache for the whole history
on every `--resume`. InjectRuntimeConfig rewrites that file on every run and
interpolated nine per-run values into it, so in practice the cache was thrown
away on the first comment that landed on any issue.

Measured on one issue over three runs: run 1 (cold) spent 89.9k cache-write
tokens building 105k of context; runs 2 and 3 each spent ~425k re-creating a
prefix they should have read. 842k of 946.5k cache-write tokens (89%) went into
re-creation, with only tools[]+system[] surviving each resume (a constant
18,085 tokens both times).

Fix: the brief now carries only what is stable for the lifetime of a resumed
session, and per-run state travels in the per-turn user message, which is
appended after the cached prefix.

- Merge kindCommentTriggered + kindAssignmentTriggered into kindIssue, and stop
  reading TriggerCommentID in classifyTask. The brief can no longer diverge by
  trigger type structurally, rather than by convention.
- Replace writeWorkflowComment/writeWorkflowAssignment with one writeWorkflowIssue
  that routes on the per-turn message. The mode-specific status rules live inside
  their own mode block, so "own the status arc" and "do not touch the status"
  can never be read as unarbitrated peers.
- Move Task Initiator, Session Continuity Notice and Connected Apps out of the
  brief into BuildPrompt via BuildTaskInitiatorBlock / SessionContinuityNotice /
  BuildConnectedAppsBlock.
- Drop TriggerCommentID, TriggerThreadID, NewCommentsSince, NewCommentCount,
  PriorSessionResumed and CommentReplyTargets from the brief; BuildPrompt already
  emitted all six from the same helpers, so this is de-duplication.
- Set PriorSessionResumeUnavailable on `task` as well as `taskCtx` in both local
  resume gates, or the notice would silently vanish on exactly the failure path
  it exists to disclose.

Tests: TestInjectRuntimeConfigByteIdenticalAcrossTriggers renders the brief
across nine per-run variants (trigger type, differing comment/thread ids, resume
delta, resume-unavailable, cross-thread fan-out, member/agent initiator,
connected apps) for two providers and requires bytes.Equal, with a non-vacuity
guard so it cannot pass on a function that ignores its input. Daemon-side tests
assert the moved sections still reach the agent through the per-turn prompt.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): route the issue workflow on an explicit turn-mode marker

Review follow-up on the mode router. The brief said Reply mode applies when
the per-turn message "opens with a [NEW COMMENT] block", but buildCommentPrompt
writes two paragraphs before that block and only emits it when
TriggerCommentContent is non-empty. Two ways to get it wrong:

- The message never literally opens with the block, so the router's own
  wording did not match the prompt it describes.
- A comment-triggered run with an empty comment body — or an older server that
  does not send one — emitted no block at all. An agent following the brief
  would fall through to Ownership mode and change the issue status on a turn
  whose rule is "do NOT change the issue status".

BuildPrompt now emits an unconditional `**Turn mode: Reply.**` /
`**Turn mode: Ownership.**` line from the same branches it uses to pick a code
path, and the brief routes on that marker. Brief and prompt can no longer
disagree about the mode, because the value that selects the path also states
it. The router also names a safe fallback (treat an unlabelled turn as Reply
mode and leave the status alone).

Tests: TestTurnModeMarkerAlwaysPresent covers comment-triggered with and
without comment content, plus both assignment shapes;
TestTurnModeMarkerAbsentOnIssuelessKinds keeps the marker off chat /
quick-create / autopilot; TestBriefModeRouterMatchesPromptMarkers fails if the
brief ever describes a marker the prompt does not emit.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-28 16:00:16 +08:00
Bohan Jiang
3c6bebaff0 MUL-5274: allow explicit persistent service handoffs (#5895)
* fix(runtime): allow explicit persistent service handoffs

Co-authored-by: multica-agent <github@multica.ai>

* fix(runtime): resolve review ambiguities in persistent-service handoff wording

Address MUL-5274 review findings on #5895:

- Drop the "The rules above apply only to work owned by the current run"
  scoping sentence: with the persistent-service exception inserted above
  it, it would have swept in work that is precisely no longer run-owned
  after handoff. The external-systems bullet carries the boundary on its
  own, and both pin tests now reject any "The rules above" reintroduction.
- Replace "detach it" (skill-level mechanism) with the lifecycle
  contract: hand off only once the service no longer depends on this run.
- End the negative-boundary bullet with "the CI-specific rules below
  still apply" instead of "must be collected before exit", which
  misread as license to start CI polling and collect it.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-24 18:21:28 +08:00
Jiayuan Zhang
a3fe6d91dd MUL-5150: add project context to Chat (#5765)
* feat(chat): add project context

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): resolve MUL-5150 review blockers

- Renumber project-context migrations to unique prefixes after current main:
  206_chat_session_project -> 212 (column), 207_chat_session_project_index ->
  213 (concurrent index). 206/207 collided with 206_agent_disabled_runtime_skills
  and main's 207-211 client_usage_daily set.
- Add the 4 missing chat input.project_context keys to ja/ko locales so the
  locale parity test passes (en/zh-Hans already had them).
- Lock the project-context control while a send is in flight (isSubmitting),
  not just while the agent is running. A brand-new chat creates its session
  lazily during send bound to the project at click time; switching project
  mid-send would create the session against the stale project and clear the
  editor as if the send landed on the new selection. Add a regression test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): complete project context handling

* fix(chat): pin fresh chat to open session's agent on project switch

Switching an existing session to a different project opens a fresh chat but
only cleared the active session, dropping selection back to the stored
`selectedAgentId`. When that preference was stale (open session belongs to
agent B while the persisted pick is still agent A), the lazily-created session
and its first send bound to the wrong agent (agent A).

Extract the project-switch decision into a shared `planProjectContextChange`
pure helper in use-chat-controller.ts and route both chat surfaces (the chat
tab controller and the floating ChatWindow) through it, so the fresh chat is
pinned to the open session's agent and the rule cannot drift between the two
copies. Add a dual-entry regression test (pure-fn guard + controller
integration) covering the stale selectedAgentId case.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* chore(ci): re-trigger required checks on latest head

The prior push updated the branch ref but GitHub did not emit a pull_request
synchronize for it (PR head-sync lag), so CI/Mobile Verify never ran on the
commit carrying the stale-agent project-switch fix. Empty commit to force a
fresh synchronize on a head that includes it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): renumber project migrations to 213/214 after main added 212

Current main added 212_agent_service_tier; the PR's 212/213 chat migrations
collided with it on the merge ref, failing TestMigrationNumericPrefixesStay
UniqueAfterLegacySet. Merge current main and move the chat column migration to
213 and the concurrent index migration to 214 (column before index preserved).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): lock ProjectPicker clear control during send (keyboard path)

The send-pending lock only put pointer-events-none on the wrapper, which
blocks the mouse but leaves ProjectPicker's inline clear button in the tab
order — a keyboard user could Tab to "Remove from project" and press Enter
mid-send, detaching the project after the lazily-created session already went
out with the old one (reopens the mid-send retarget path via keyboard).

Add an explicit `disabled` capability to the shared ProjectPicker that locks
the trigger, the menu (forced closed), and the inline clear button (disabled +
out of the tab order). Defaults to false, so issue/create/autopilot callers
keep their hover/keyboard clear. ChatInput passes disabled while the project
selection is locked.

Tests: real-ProjectPicker regression (keyboard activation of the clear control
is inert when disabled; still works when enabled) + ChatInput wiring assertion
that the picker is disabled mid-send.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Walt <walt@multica.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Naiyuan Qing <145280634+NevilleQingNY@users.noreply.github.com>
Co-authored-by: NevilleQingNY <nevilleqing@gmail.com>
2026-07-24 11:30:27 +08:00
Bohan Jiang
e2ce3a2da8 MUL-5223 fix(runtime): forbid blocking on external CI in the runtime brief (#5840)
* fix(runtime): forbid blocking on external CI in the brief (MUL-5223)

The external-work boundary added in #5803 did not stop agents from
waiting on GitHub Actions. Two holes: the section's only concrete
"how to wait" example was a blocking foreground call, which is exactly
the shape of `gh pr checks --watch`; and the "unless acceptance
criteria require it" escape was satisfied by the repo's own merge
requirement that CI be green.

Name the banned tool shapes, allow a single non-blocking status
snapshot, deny branch protection as an acceptance criterion, and give
the replacement hand-off phrasing (local test result + PR link).

Co-authored-by: multica-agent <github@multica.ai>

* fix(runtime): scope the CI-wait ban so the explicit exception stays executable (MUL-5223)

Review feedback on #5840:

- The ban read as absolute ("Blocking on external CI is never part of
  your deliverable") while the next bullet allowed waiting when the
  task explicitly asks for the CI result, leaving no way to satisfy
  both. The ban is now scoped to "unless the explicit exception below
  applies", and the exception names the one executable shape: a single
  foreground blocking watch inside the same turn.
- `gh pr merge --auto` enables auto-merge and returns; it is not a
  wait. Only waiting for it to land is banned.

Both hard-pin tests now also pin the exception so it cannot be dropped
or re-absolutised.

Co-authored-by: multica-agent <github@multica.ai>

* polish(runtime): group Background Task Safety into run-owned and external-CI clusters (MUL-5223)

Co-authored-by: multica-agent <github@multica.ai>

* polish(runtime): cut redundant phrasing from the external-CI cluster (MUL-5223)

The cluster said "report and finish" three different ways and carried
two rhetorical tails. Fold the delivery template into the post-push
playbook bullet, tighten the merge-gate denial, and drop filler.
5 bullets -> 4, -36 words, every behavioral fact and test pin intact.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-23 18:38:51 +08:00
Jiayuan Zhang
2c749ddf9f fix(runtime): clarify external background work (#5803) 2026-07-23 02:43:17 +08:00
Bohan Jiang
fbe00ca164 fix(daemon): run Codex unsandboxed on Windows to stop reject-by-policy (MUL-4957) (#5672)
* fix(daemon): run Codex unsandboxed on Windows to stop reject-by-policy (MUL-4957)

Windows has no Landlock/Seatbelt-equivalent filesystem sandbox that the
daemon configures, so the per-task `sandbox_mode = "workspace-write"` it
wrote was unenforceable. Worse than having no sandbox, it pushed Codex
into rejecting non-safe mutation commands "by policy": `multica issue
create` fails with "was rejected by policy" because Codex can neither
sandbox the command nor (under approval_policy = "never") escalate it to
the daemon's auto-approver, so the request never reaches the approver.

Mirror the existing macOS fallback and give Windows danger-full-access so
those commands run. Also generalize the danger-full-access warn log so it
no longer hardcodes "on macOS" and only surfaces the macOS-specific
upgrade hint on macOS (new codexSandboxPolicy.Hint field).

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): correct Windows sandbox rationale and respect user windows.sandbox (MUL-4957)

Addresses two review must-fixes on #5672:

1. Correct a false security fact. The comments and log Reason claimed
   Windows has no filesystem sandbox backend. Codex 0.144.5 does ship a
   native Windows sandbox (windows.sandbox = "unelevated"/"elevated"); it
   is experimental with open upstream reliability bugs, so the daemon
   defaults to danger-full-access as a deliberate compatibility choice.
   Enabling the native sandbox is tracked as separate follow-up work.

2. Stop silently downgrading users who opted into isolation. The fallback
   was unconditional. Add codexSandboxPolicyForConfig: on Windows an
   explicit windows.sandbox = unelevated|elevated keeps workspace-write so
   Codex enforces task isolation with the user's chosen backend;
   danger-full-access applies only when windows.sandbox is absent,
   disabled, or unparseable. This is also the branch point for a future
   native-sandbox rollout (flip the default; callers unchanged).

Adds fixture tests locking the priority (user opt-in kept vs. unconfigured
fallback) plus predicate coverage for codexSandboxPolicyForConfig.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): fail closed on undecidable Windows sandbox config, honor -c windows.sandbox (MUL-4957)

Second review round on #5672. Two must-fixes.

1. Undecidable config no longer fails open. The old bool detector collapsed
   "unparseable / invalid value / failed copy" into "unconfigured" and then
   loosened to danger-full-access. Replaced with a tri-state
   (absent/native/undecidable): only exact-lowercase unelevated|elevated (the
   sole values Codex accepts — verified: any other value makes Codex refuse to
   load the config) counts as native; any other present value, unparseable
   TOML, a read error, or a missing per-task config when a shared
   ~/.codex/config.toml exists (i.e. the copy failed) is undecidable and fails
   closed to workspace-write — it never loosens — logged at error level.

2. windows.sandbox set via `-c`/`--config` custom args is now honored. Such
   args never land in config.toml, so config-only detection silently
   downgraded those users' isolation. The effective Codex args (daemon
   defaults + profile-fixed + per-agent custom_args) are threaded through
   PrepareParams/ReuseParams/CodexHomeOptions into the sandbox decision and
   scanned for a windows.sandbox override (inline, two-token, quoted, spaced;
   last-wins).

Also drops issue-status-bound source comments (openai/codex#24098 has since
closed). Adds unit coverage for config/args classification, the fold
precedence (undecidable > native > absent), and the copy-failed fail-closed
path.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): fail closed on config-sync errors and honor shell-quoted -c windows.sandbox (MUL-4957)

Round-3 review must-fixes:

1. resolveWindowsSandboxState now takes the config.toml sync error and a
   tri-state shared-config presence instead of re-stat-ing inside. A failed
   sync (stale/absent per-task copy) or an un-stat-able shared source is
   undecidable and keeps workspace-write, closing the fail-open where a failed
   sync was read as "unconfigured". Splits IO from the decision so the paths
   are unit-testable without faulting the filesystem.

2. The Windows sandbox decision consumes agent.NormalizeCodexLaunchArgs (the
   shared helper buildCodexArgs now uses) so a shell-quoted -c windows.sandbox
   opt-in is normalized identically to launch, instead of being missed by a
   raw-token scan and silently downgraded.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): abort when the Codex sandbox block cannot be written (MUL-4957)

Round-4 review must-fix: ensureCodexSandboxConfig failures were warn-and-continue,
so a computed fail-closed workspace-write policy could stay only in memory while
config.toml kept a stale danger-full-access from a prior run — the decision
failed closed but the effective config failed open.

prepareCodexHomeWithOpts now returns the error, which blocks startup on both
paths: fresh Prepare fails the task, and Reuse leaves env.CodexHome unset, which
configureCodexTaskShellEnvironment already refuses to start.

Regression covers the full reuse scenario (stale danger-full-access + failed
config sync + failed managed-block write); it fails with "got nil" without the
fix.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: J <j@multica.ai>
2026-07-21 15:18:44 +08:00
Bowser
f8bf6cd8b9 feat(runtime): add Qwen Code runtime (MUL-5015)
Merge approved PR #5666.
2026-07-21 14:55:08 +08:00
Wangjue Yao
5664c0d755 MUL-4874: fix(execenv): seed Codex model cache in task homes (#5449)
* fix(execenv): reuse Codex startup caches

* fix(execenv): preserve task-local Codex model cache

* fix(codex): bind model cache to provider config

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-16 17:13:52 +08:00
YikaJ
54fd29ebdd MUL-4383 fix(daemon): stop routing CodeBuddy skills/memory through Claude's .claude paths (#5224)
* fix(daemon): stop routing CodeBuddy skills/memory through Claude's .claude paths

CodeBuddy Code is a Claude Code fork but ships its own native config
directory (~/.codebuddy, .codebuddy/) with its own memory filename
(CODEBUDDY.md). It only reads .claude/skills or CLAUDE.md if a user
manually symlinks/copies them during migration
(https://www.codebuddy.ai/docs/cli/troubleshooting#migrating-from-claude-code).

Multica's daemon/execenv code treated "codebuddy" as an alias for
"claude" in three places, so skills synced by Multica landed in
.claude/skills/ and CLAUDE.md — paths the default CodeBuddy install
never reads — instead of ~/.codebuddy/skills, .codebuddy/skills, and
CODEBUDDY.md as documented at
https://www.codebuddy.ai/docs/cli/codebuddy-dir and
https://www.codebuddy.ai/docs/cli/skills.

Split the "claude", "codebuddy" switch cases in:
- daemon/local_skills.go (user-level local skill discovery/import)
- daemon/execenv/context.go (per-task skill materialization)
- daemon/execenv/runtime_config.go (runtime brief target file)

Added regression tests locking in the new paths and updated the
install-agent-runtime / providers docs (all 4 locales) that had
documented the old .claude/skills behavior.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(daemon): cover CodeBuddy sidecar hygiene and lifecycle

Address PR #5224 review feedback: exclude CODEBUDDY.md/.codebuddy from
repo-cache worktrees, extend sidecar lifecycle matrices to codebuddy, and
make the local-skills CodeBuddy test exercise a true same-key collision.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Eve <eve@multica-ai.local>
2026-07-15 16:47:15 +08:00
Bohan Jiang
6cc553e5a3 fix(daemon): isolate Codex sessions per task to unblock initialize (MUL-4424) (#5360)
* fix(daemon): isolate Codex sessions per task to unblock initialize (MUL-4424)

Codex 0.143+ backfills a per-home session-state DB by enumerating every
rollout visible under sessions/ during `initialize`. The per-task
CODEX_HOME symlinked the shared ~/.codex/sessions in, so a machine with a
large accumulated history (one reporter: ~2000 rollouts / ~22 GiB) stalled
`initialize` for tens of seconds — the app-server started but the task
produced no output before it was cancelled (github #5273).

Give each task its own local sessions/ directory instead:

- Fresh task: create an empty local sessions/ so backfill is trivial.
- Reused task with a real sessions/ dir: it is authoritative — leave it.
- Reused task still holding a legacy symlink (older build): migrate in
  place. Replace the symlink with a real dir; when resuming, symlink only
  the single rollout being resumed (never copy — a rollout can be GiB and
  this is on initialize's critical path); and drop the stale, rebuildable
  session-state DB (state_*.sqlite*, session_index.jsonl) so Codex
  re-indexes the task-local sessions. Unrelated per-task DBs (goals_*,
  logs_*, memories_*) are left intact.

Also point the token-usage fallback scan at the backend's per-task
CODEX_HOME instead of the daemon-global home, so usage isn't lost now that
sessions are isolated there.

Complements the #5319 handshake watchdog (which turns a silent stall into a
loud, phased timeout); this removes the underlying cause.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): address Codex session isolation review (MUL-4424)

Resolves the three blockers from Elon's review of #5360:

1. local_directory context loss. local_directory tasks get a fresh
   codex-home per task ID (the daemon never reuses their workdir), so
   task-local isolation stranded every follow-up run with an empty
   sessions dir and silently restarted the conversation. Their only
   stable, GC-safe cross-task store is the user's own ~/.codex/sessions
   (a persistent store under WorkspacesRoot would be orphan-GC'd), so keep
   the shared-sessions symlink for them (IsLocalDirectory). Managed tasks
   stay isolated.

2. Migration resume robustness.
   - Rollout lookup now covers the flat layout and background-compressed
     .jsonl.zst rollouts, not just nested YYYY/MM/DD *.jsonl — both are
     legitimate Codex 0.144 history that were previously judged "not
     found", silently dropping resume.
   - Exposure hard-links first, then symlinks, never copies — hard links
     need no privilege and work on Windows within a volume, so the
     zero-copy path is exercised identically on CI.
   - The daemon now verifies the rollout is actually present in the task
     CODEX_HOME (execenv.CodexResumeRolloutPresent) before the brief is
     generated; if absent it clears the resume from both the backend and
     the brief instead of telling the agent it is continuing a lost thread.

3. session_index.jsonl is no longer deleted during migration — Codex uses
   it as the authoritative thread-id -> name store (not rebuildable from
   rollouts). Only the rebuildable state_*.sqlite* is reset.

Tests: 2-round local_directory resume across task IDs; compressed/flat
lookup; hard-link zero-copy (os.SameFile); session_index preserved;
CodexResumeRolloutPresent + the daemon gate helper (present keeps /
absent drops / non-codex + empty no-op).

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): scope Codex sessions to a per-issue store; disclose lost resumes (MUL-4424)

Addresses the three blockers from Elon's second review of #5360.

1. local_directory still enumerated the whole machine history. The prior
   fix re-linked the entire ~/.codex/sessions into every fresh local_directory
   codex-home, so Codex still backfilled from thousands of unrelated rollouts on
   `initialize` (measured ~8.3s with 3450 rollouts; the reporter's 22 GiB could
   exceed the 30s watchdog). Point sessions/ at a persistent, per-(agent, issue)
   store under the shared Codex home (multica-sessions/<agent>/<issue>) that holds
   only this issue's rollouts. It is keyed stably across task IDs and lives
   outside the task-scoped envRoot the GC reclaims, so follow-up runs resume it
   while `initialize` only ever sees this issue's history.

2. Windows cross-volume resume was lost. Exposing a single rollout by hard-link
   (same-volume only) then file symlink (needs Windows privilege) can't cross a
   volume boundary. The store now lives on the shared Codex volume, so the resume
   rollout is hard-linked there zero-copy, and sessions/ is exposed to the task
   home via a directory link — a symlink on Unix, a junction on Windows — which
   crosses volumes without privilege and never copies a (possibly GiB) rollout on
   initialize's critical path. There is no remaining per-file cross-volume link.

3. An unavailable resume was a silent downgrade. Both resume gates
   (gateResumeToReusedWorkdir, gateCodexResumeToRolloutPresence) now set
   PriorSessionResumeUnavailable, and the runtime brief renders a Session
   Continuity Notice telling the agent to disclose to the user, up front in its
   reply, that the previous conversation context could not be restored and this
   run starts fresh — turning a silent restart into a user-visible one. The task
   is not failed: it can still do useful work without the prior context.

Managed fresh / reused-real-dir tasks keep their task-local, GC-collected
sessions dir unchanged; only the legacy-symlink migration with a resume routes
through the store (cross-volume-safe), and a home already linked to the store is
treated as authoritative on reuse.

Tests: local_directory per-issue store (only this issue's history, no whole-
machine leak); no-key fallback to an empty dir; two-round resume across task IDs
through the store; legacy migration routed through the store with a zero-copy
hard link; reused store link stays authoritative; both gates set the
resume-unavailable flag; brief renders the continuity notice only when a resume
was lost. execenv + daemon + pkg/agent packages, go vet, and gofmt all pass.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): disclose live resume-RPC loss; bound Codex session store lifecycle (MUL-4424)

Addresses the two blockers from Elon's third review of #5360.

1. A real thread/resume failure was still a silent new session. The brief's
   Session Continuity Notice only covers losses the daemon detects before launch
   (workdir not reused, rollout absent). But when the rollout is present yet
   Codex rejects the live thread/resume (corrupt/incompatible rollout, server-side
   thread GC, schema drift), startOrResumeThread falls back to thread/start and
   the run succeeds on a fresh thread with no user-facing signal. Carry the
   original resume intent into the backend as ExecOptions.ResumeExpected (set from
   the post-gate PriorSessionID, so a pre-flight drop still routes through the
   brief and never double-notifies), and when a resume was expected but the
   backend landed on a fresh thread, prepend the same continuity notice to the
   first turn/start input. This also covers the daemon's transport-error
   fresh-session retry, which clears ResumeSessionID but not ResumeExpected.

2. The persistent per-issue store had no data lifecycle. multica-sessions stores
   live outside the task-scoped envRoot the GC reclaims (so resume survives across
   task IDs), which meant a done/abandoned issue's prompts and full rollouts (one
   reporter: a single 1.5 GiB rollout) accumulated forever and were never freed on
   issue/agent/workspace deletion. Add PruneCodexSessionStores: the daemon GC loop
   reclaims any store untouched for GCCodexSessionTTL (default 14 days, configurable
   via MULTICA_GC_CODEX_SESSION_TTL, 0 disables). A store's newest rollout mtime is
   its last activity, so an active or recently-resumed task keeps its store fresh
   and is never reclaimed, while a deleted issue's store ages out — an eventual
   reclamation guarantee without needing deletion events.

Tests: codexTurnInput discloses on resume fallback and stays silent on success /
fresh start (paired with the existing live-RPC fallback test); store pruning
reclaims aged stores, keeps recent ones, isolates issues, cleans empty agent
dirs, and is disable-able. execenv / daemon / pkg/agent, go vet, gofmt all pass.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): protect a reopened Codex session store from GC mid-mount (MUL-4424)

Addresses Elon's fourth-review blocker: reopening an issue idle past
GCCodexSessionTTL could lose context, because mounting its per-issue session
store (MkdirAll + rollout lookup + task-home link) never refreshed the store's
mtime, so a GC cycle firing before the resumed turn wrote its first rollout saw
a >TTL-old store and reclaimed it — a stat->remove race with no in-use guard.

Two complementary defenses:

- Activity refresh: linkCodexSessionsToStore now os.Chtimes the store to now
  after linking, so codexStoreStat (which reads the newest mtime as last
  activity) sees a just-used store. This fixes the sequential repro — a mount
  immediately followed by a prune keeps the store.

- In-process active-store guard: the daemon marks the per-issue store in-use
  (execenv.CodexSessionStorePath) from before Prepare/Reuse mounts it until the
  task ends, and PruneCodexSessionStores now takes an isActive predicate and
  skips any store a live task holds. Because prepare and prune run in the same
  process, this closes the remaining concurrent stat->remove window the mtime
  refresh alone cannot. Reference-counted, mirroring the env-root guard.

Tests: a reopened >TTL store survives a GC cycle after remount and stays
resumable; an idle-on-disk store marked active is skipped, then reclaimed once
inactive; the existing idle-reclaim / isolation / disable / empty-agent-dir
cases still pass. execenv + daemon, go vet, gofmt all pass.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): make Codex store delete atomic with mark-active (MUL-4424)

Addresses Elon's fifth-review blocker: the active-store guard's check and delete
were not one atomic step. PruneCodexSessionStores called isActive (which locked,
read, and unlocked) and only then RemoveAll'd, leaving a window where a task
could markActiveCodexStore between the check and the removal and still lose its
store — the exact mark-then-delete interleaving Elon reproduced.

Replace the point-in-time isActive predicate with a reserve-for-deletion
protocol that shares one lock with mark-active:

- reserveCodexStoreForDeletion(store) atomically refuses when a live task holds
  the store (or another delete already reserved it) and otherwise marks it
  reserved, all under one activeCodexStoresMu acquisition. PruneCodexSessionStores
  reserves before RemoveAll and commits after, so confirm-inactive and remove are
  effectively atomic against a concurrent mark.
- markActiveCodexStore now waits (on a sync.Cond) while a store is reserved, so a
  task never mounts a store mid-removal; committing the removal wakes it and the
  store is recreated fresh by Prepare (with the continuity notice).

So mark-before-reserve keeps the store (reserve refused); reserve-before-mark
removes it and blocks the late mark until the removal commits. The genuinely
idle case still reclaims.

Tests (daemon, run under -race): mark-then-reserve is refused; reserve blocks a
concurrent mark until commit then the store reads active; a second reserve is
refused mid-flight. The execenv prune tests move to the reserve seam; the
activity-refresh / reopen-then-prune / isolation / disable cases still pass.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): namespace Codex session stores per profile for cross-daemon safety (MUL-4424)

Addresses Elon's sixth-review blocker: the in-process reservation guard cannot
span processes, but Multica supports multiple profile daemons on one machine
(e.g. production + staging) that share the same ~/.codex. Each daemon's GC
scanned the whole multica-sessions root, so a staging daemon could reclaim a
store a production task was actively resuming — its reservation lived only in the
other process's memory.

Isolate by profile instead of trying to lock across processes:

- Store path is now <shared>/multica-sessions/<namespace>/<agent>/<issue>, where
  namespace is the daemon's profile (empty -> "default"). PrepareParams/ReuseParams
  carry Profile; codexSessionStoreKey and CodexSessionStorePath fold it in.
- PruneCodexSessionStores takes the profile and scans ONLY that namespace, so a
  daemon never even sees another profile's stores, let alone deletes them. The
  per-profile trees are disjoint, so the in-process guard is sufficient within a
  namespace (profiles get separate daemon state, so no two daemons share one).

Test: a "staging"-owned idle store is untouched by a default-profile prune and
reclaimed only by staging's own prune. Existing prune/guard/reopen tests move
under the namespace. execenv + daemon under -race, go vet, gofmt all pass.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): make the Codex store profile→namespace map injective (MUL-4424)

Addresses Elon's seventh-review blocker: the per-profile namespace was derived by
dropping unsafe characters, which is not injective. The CLI treats the empty
(default) profile and a profile literally named "default" as separate daemons,
yet both mapped to namespace "default"; likewise "staging.prod" and "stagingprod"
both mapped to "stagingprod". Two distinct daemons then shared one store tree, so
one could again reclaim the other's live session — the cross-process blocker
reopened for those profile names.

Make codexSessionStoreNamespace injective: the empty profile gets a reserved
bare literal "default", and every named profile is hex-encoded (bijective,
filesystem-safe) under a "p_" prefix a bare literal can never collide with. So
"" -> "default" while "default" -> "p_64656661756c74", and "staging.prod" /
"stagingprod" get distinct hex segments. sanitizeCodexPathSegment stays for the
UUID agent/issue segments (injective for real UUIDs); only the user-controlled
profile needed the encoding.

Tests: codexSessionStoreNamespace is distinct for "" vs "default", punctuation
variants, case variants, and an encoded-looking name; and end-to-end, pruning one
profile never reclaims the other's store for the "" vs "default" and
"staging.prod" vs "stagingprod" pairs. execenv + daemon under -race, go vet,
gofmt all pass.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): fixed-length Codex store namespace so long profiles fit (MUL-4424)

Addresses Elon's eighth-review blocker: hex-encoding the full profile doubled the
namespace segment length. A profile can be as long as a filesystem segment allows
(~255 bytes) and the CLI persists it as its own config dir, but the store
namespace "p_" + hex(profile) reached 2 + 127*2 = 256 bytes at 127 chars,
overflowing the 255-byte single-segment limit — the profile's own dir created
fine, then the session store failed with "file name too long".

Derive the namespace from a fixed-length hash instead: a named profile is now
"p_" + hex(sha256(profile)) — a constant 64 hex chars (66 with the prefix),
filesystem-safe and collision-resistant. The empty (default) profile keeps its
reserved bare literal "default", which the "p_"-prefixed 66-char segment can
never equal. Still injective across the CLI's distinct-daemon cases; just no
longer length-expanding.

Test: the namespace stays <=255 bytes and creatable for profiles up to the
255-byte segment limit (127- and 255-char cases that overflowed under hex); the
prior injectivity and cross-profile prune-isolation tests still hold. execenv +
daemon under -race, go vet, gofmt all pass.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-15 00:48:30 +08:00
Rusty Raven
ef9b334408 MUL-4398: fix Hermes bound-skill discovery with per-task overlay (#5308)
* fix(execenv): overlay per-task HERMES_HOME so Hermes discovers bound skills

Hermes has no workspace-relative skill discovery — it scans <HERMES_HOME>/skills
first, then skills.external_dirs from config.yaml (verified against the bundled
agent/skill_utils.py). The daemon wrote assigned skills to the generic
.agent_context/skills/ fallback, which Hermes never reads, so they silently never
took effect (#5242).

When (and only when) an agent has skills bound, redirect HERMES_HOME to a minimal
per-task compatibility overlay; a skill-less Hermes task keeps its real home and
original behavior:

- mirror every top-level entry of the shared home via symlink except the
  overlay-owned ones (denylist), reconciling entries deleted from the shared home;
- derive a task-local config.yaml whose skills.external_dirs references the shared
  skills dir plus the user's existing external_dirs, expanded against the
  sanitized effective child env (unknown vars preserved, blocklisted keys resolved
  to the process value) and normalized to absolute paths;
- write only the bound skills into the task-local skills/ dir (home skills scanned
  first, so they win); global skills are referenced, not copied;
- keep memories/ overlay-owned (fresh per-task dir) AND disable the external
  memory.provider, so neither on-disk memory nor a shared backend crosses tasks;
- keep active_profile/profiles out of the overlay so Hermes can't follow a sticky
  profile and redirect past it at startup.

Profile handling mirrors hermes_cli.profiles: the daemon reads -p/--profile with
agent.HermesProfileFromArgs and seeds the overlay from that profile's home via
ResolveHermesSourceHome (default/invalid -> base, valid name -> <base>/profiles/
<name>, validated; a missing named profile fails closed). The profile flags are
stripped from the acp argv ONLY when the overlay is active (hermesLaunchArgs), so
a skill-less task's profile passes through unchanged. HERMES_HOME is no longer
custom_env-blocklisted: no skills -> user value passes through; skills -> overlay
overrides after layering. Fail closed — Prepare errors, Reuse returns nil.
Task home 0700, derived config 0600 via atomic replace. Platform-native default
home (%LOCALAPPDATA%\hermes, incl. the LOCALAPPDATA-missing fallback, on Windows).

Tests span execenv/daemon/agent: no-skill no-op, child-env layering + env
sanitization, profile parse/unquote + conditional strip + final args/env per
scenario, custom/profile/default/invalid/missing/Windows source home, sticky-
profile not mirrored, memory dir isolation + external provider disable, mirror
reconciliation, external_dirs rebasing + sanitized/unknown-var expansion,
local-precedence slug, perms, fail-closed, resume teardown. Docs (en + ja/ko/zh).

Fixes #5242

* fix(hermes): make profile selection one resolver contract matching Hermes

Round 5 review: the profile chain approximated Hermes' semantics in three
separate places (argv parsing, source-home selection, arg filtering), so it
diverged from native Hermes in several merge-blocking cases. Collapse it into
one authoritative resolution:

- agent.ParseHermesProfileArgs replaces HermesProfileFromArgs/
  FilterHermesProfileArgs. It reproduces _apply_profile_override step 1/1b
  (first occurrence, value-flag skipping, `--` and `mcp add --args` boundaries,
  space-form profile-id guard) and returns the exact argv occurrence to consume;
  StripHermesProfileArgs removes only that occurrence.

- execenv.ResolveHermesProfile replaces ResolveHermesSourceHome. It derives the
  Hermes root exactly like get_default_hermes_root (an already-profile-scoped
  HERMES_HOME roots at its grandparent), selects an explicit profile first,
  otherwise trusts a profile-scoped home (step 1.5) and only then the sticky
  <root>/active_profile (step 2), and validates via normalize/validate_profile_name
  (reserved hermes/test/tmp/root/sudo and empty inline `--profile=` are hard
  errors). Profiles always resolve under the root, so `-p default` re-roots and
  `-p <sibling>` is a sibling, never nested.

- The daemon runs one parse + resolve, fails the task closed on a reserved/
  invalid selection (matching Hermes' sys.exit(1)), and exports the selected
  source home as the effective env's HERMES_HOME so ${HERMES_HOME} in a profile's
  skills.external_dirs expands against the selected profile home (as native
  Hermes does before loading config.yaml), not the root or the overlay.

Regressions added: root + sticky named profile selection; already-profile-scoped
home with no flag; that home with -p default and -p <sibling>; reserved and empty
inline profile values; and a selected profile whose external_dirs contains
${HERMES_HOME}.

* fix(hermes): overlay-owned derived .env + symlink-resolved root

Round 6 review, two remaining overlay-bypass paths:

1. A source `.env` could redirect HERMES_HOME after profile resolution. Hermes
   runs `_apply_profile_override()` then `load_hermes_dotenv()`, which loads
   `<HERMES_HOME>/.env` with override=True — so a mirrored source `.env` carrying
   an out-of-band `HERMES_HOME=` overwrote the overlay's home, repointing skill
   discovery and memory back at the source. `.env` is now overlay-owned and
   DERIVED (writeDerivedHermesEnv): it preserves the source's credentials/settings
   but strips any `HERMES_HOME` assignment and pins `HERMES_HOME` to the overlay
   last (single-quoted, literal), written 0600 via atomic replace. It is written
   even when the source has none, so Hermes' project-`.env` fallback (override=True
   only when no user `.env` loaded) can't relocate the home either.

2. Root derivation was lexical-only, diverging from `get_default_hermes_root`,
   which compares `env_path.resolve()` with `native_home.resolve()`. A HERMES_HOME
   symlinked into `<native>/profiles/<x>` was treated as its own root, so
   `-p default`/`-p <sibling>` resolved wrong. `hermesRootFromHomeFor` now resolves
   symlinks (Path.resolve(strict=False)-style best effort) for the containment
   decision while keeping the returned root unresolved, matching Hermes.

Regressions: source `.env` with HERMES_HOME replayed through the override=True
dotenv order (bound skill + task memory stay on the overlay; creds preserved);
minimal overlay `.env` created when the source has none; and a symlinked profile
home resolving `-p default`/`-p <sibling>` to the native root.
2026-07-14 15:05:18 +08:00
Multica Eve
0c2e48ded2 refactor: retire FF_RUNTIME_BRIEF_SLIM, make slim runtime brief the only path (MUL-4297)
The runtime_brief_slim feature flag has burned in; the slim runtime brief is now the sole path.

- execenv: buildMetaSkillContent / BuildCommentReplyInstructions delegate to the slim assembler unconditionally; delete the legacy verbose brief body and writeBackgroundTaskSafetyInstructions.
- Remove the runtime_brief_slim flag and the daemon-bound flag delivery subsystem built solely for it: execenv flag wiring (runtime_config_flag.go, server_snapshot_provider.go), the featureflagdispatch package, the DaemonFeatureFlagSnapshot heartbeat protocol field, and the server/daemon wiring in router.go, handler, daemon.go, main.go, cmd_daemon.go.
- Keep the generic server/pkg/featureflag engine (still used by composio_mcp_apps).
- Update tests to slim-only expectations and docs/feature-flags.md.

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-09 13:48:33 +08:00
Bohan Jiang
4c510dfef6 fix(daemon): harden background-task-safety brief against background-and-yield (MUL-4140) (#4998)
A Multica-managed run goes terminal the moment the top-level turn exits;
there is no "background work finishes later and wakes you up" step. When an
agent starts background work (a run_in_background shell, a Monitor, an async
subagent) and ends its turn to "wait for a completion notification", the work
is orphaned and the result comment it meant to post is never sent (MUL-4091 /
PR #4970).

The existing claude-only protocol guard forces run_in_background tool inputs
to foreground and fails loud on async_launched tool results, but it cannot
catch the actual MUL-4091 mechanism: a turn that ends cleanly with a
"Standing by, I'll report when CI finishes" message. That shape is only
addressable behaviorally, and it is harness-agnostic.

Harden the Background Task Safety brief (both the legacy/verbose production
path and the slim staging path) with explicit hard pins:
- never background-and-yield / expect a future wakeup that does not exist here;
- do every wait synchronously in a single foreground call (e.g. gh run watch);
- the standalone-harness "running in the background, keep working" hint does
  not apply in Multica-managed runs;
- never end a turn with a "standing by" / "I'll report back" sign-off.

Add verbose- and slim-path test coverage for the new pins so a future brief
trim cannot silently drop them.

Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-06 19:47:00 +08:00
ZeroIce
65269ef922 fix(daemon): copy Codex model catalog into task home
Fixes #4825

Co-authored-by: multica-agent <github@multica.ai>
2026-07-03 11:57:34 +08:00
beast
0c4c3ff038 fix(cli): prevent daemon-managed CLI from silently using user tokens (MUL-3922)
Treat MULTICA_DAEMON_PORT and a workdir daemon-task marker as daemon-managed signals so a task subprocess that loses MULTICA_TOKEN / MULTICA_AGENT_ID / MULTICA_TASK_ID fails closed instead of silently falling back to the user config-file PAT (which made agent writes land as the workspace owner). Adds an actionable error naming a leftover marker for local_directory recovery. Fixes #4204.
2026-07-02 15:08:06 +08:00
beast
20eecfb093 fix(projects): honor repo resource checkout refs (MUL-3593) (#4470) 2026-06-24 16:25:17 +08:00
Bohan Jiang
76c58a4ee8 MUL-3617: remove Gemini CLI runtime (#4503)
* fix: remove gemini cli runtime

Co-authored-by: multica-agent <github@multica.ai>

* fix: skip unsupported custom runtime profiles

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-06-24 15:15:42 +08:00
Bohan Jiang
da72e2fa22 feat(daemon): inject project description into the agent brief (MUL-3465) (#4395)
* feat(daemon): inject project description into the agent brief

Issues bound to a project only surfaced the project title in the runtime
brief; the project description (durable, project-wide context the owner
sets) was loaded but dropped. Carry it end-to-end:

- claim handler reads proj.Description onto the response (issue-bound and
  quick-create paths)
- new ProjectDescription field on AgentTaskResponse, daemon Task, and
  TaskContextForEnv
- rendered in the brief's `## Project Context` section and written to
  .multica/project/resources.json as project_description

Empty descriptions render nothing (no extra heading). Updated the
projects-and-resources built-in skill docs in the same change.

MUL-3465

Co-authored-by: multica-agent <github@multica.ai>

* feat(projects): clarify project description is injected as agent context

The project description is now durable context injected into every task's
brief, but the UI still presented it as a plain "Description" field, so
existing descriptions could silently become agent input. Add a hint under
the description editor on the project detail page and in the create-project
modal, in all four locales, stating it is shared with agents as context for
every task in the project. No data-semantics change.

Addresses review feedback on PR #4395. MUL-3465

Co-authored-by: multica-agent <github@multica.ai>

* test(handler): assert project description flows through task claim

The execenv tests cover brief rendering, but nothing pinned the claim
handler boundary where proj.Description is read onto the response. Add
two tests — issue-bound and quick-create paths — so a regression in that
assignment fails loudly instead of silently dropping the description.

Addresses review feedback on PR #4395. MUL-3465

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-06-22 23:39:27 +08:00
DylanLi
78342a39ce MUL-3305: feat(agent): add qoder CLI as a choice of agent provider. (#2461)
* feat(agent): Qoder ACP runtime, chat reconnect recovery, and task linkage

- Add Qoder CLI backend (ACP transport, model discovery, blocked-args policy)
- Wire daemon/runtime config, docs, and UI provider assets
- Retry terminal task reports; add backoff unit tests
- Chat: SQL attach user message to task; handler + optimistic cache reconcile
- Invalidate chat/task-messages caches on WS reconnect; extract helper + tests

Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: drop non-Qoder changes (chat reconnect, task link, terminal report retries)

Keep only Qoder runtime, docs, daemon config/execenv, and UI provider assets.

Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent): harden Qoder ACP drain and wire project skills path

- Stop streaming to msgCh after reader wait so grace timeout cannot race close
- Resolve injected skills to .qoder/skills per Qoder CLI discovery
- Update AGENTS.md skill copy and add execenv tests

Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(qoder): add provider logo and wire MCP config into ACP sessions

- Add inline SVG QoderLogo component to provider-logo.tsx, replacing
  the generic Monitor icon placeholder
- Add convertMcpConfigForACP helper to convert Claude-style MCP server
  config (object map) into ACP array format for session/new and
  session/resume
- Add unit tests for convertMcpConfigForACP covering stdio, SSE,
  empty/nil, and multi-server cases

Co-authored-by: Orca <help@stably.ai>

* fix(test): capture both return values from InjectRuntimeConfig in Qoder test

Co-authored-by: Orca <help@stably.ai>

* fix(qoder): preserve remote MCP headers and promote provider errors

Addresses review feedback on #2461 (Bohan-J): two runtime-correctness
issues in the Qoder ACP backend.

1. Remote MCP headers were dropped. The bespoke convertMcpConfigForACP
   only forwarded url/type, so an authenticated remote MCP server looked
   configured in Multica but failed inside the Qoder session. Replace it
   with the shared buildACPMcpServers helper (same path Hermes/Kimi/Kiro
   use), which preserves headers as [{name, value}], sorts for
   deterministic output, and handles remote transport aliases. Fail
   closed on malformed mcp_config instead of silently dropping servers.

2. Provider failures could report as completed tasks. stderr was wired
   via io.MultiWriter and the result was only promoted to failed when
   output was empty, so a terminal upstream error (HTTP 429 / expired
   token) racing a stopReason=end_turn with text still became
   "completed". Switch to StderrPipe + an explicit copier, drain it
   (bounded by the existing grace window, since qodercli can leave a
   child holding the inherited fds) before the decision, and run the
   shared promoteACPResultOnProviderError.

Tests: replace the convertMcpConfigForACP unit tests with two
end-to-end Qoder tests — one asserts the Authorization header reaches
the session/new payload as {name, value}, the other asserts a terminal
stderr error with non-empty output reports failed.

Co-authored-by: Orca <help@stably.ai>

* fix(qoder): align ACP session handling

Co-authored-by: Orca <help@stably.ai>

* fix(agent): guard qoder late output after drain

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-06-22 18:55:45 +08:00
BeliyDym
5fd3d01d13 MUL-3502: OST-1161: Bound assignment comment catch-up
Squashed PR #4392. Updates assignment/comment catch-up guidance to use recent 10 and aligns related examples.
2026-06-22 15:46:47 +08:00
Hzzzzzx
b4c9e4423c test: enable -race detector in Go test pipeline (WOR-61) (#4274)
* test: enable -race detector in Go test pipeline (WOR-61)

Add the -race flag to all three Go test invocation sites so the existing
concurrency regression harness (workdir_race_test.go for #3999,
runtime_gone_test.go, runtime_profile_drift_test.go) actually exercises
the race detector. The daemon package alone has 28+ goroutine launch
points with no automated race coverage before this change.

Sites updated:
- Makefile:299        (make test, local)
- .github/workflows/ci.yml:101      (CI backend job)
- .github/workflows/release.yml:55  (release verify job)

go test already runs a vet subset by default, so no separate -vet flag
is added. No production code touched.

Co-authored-by: multica-agent <github@multica.ai>

* test(execenv): serialize runtimeGOOS-mutating test (WOR-61)

TestInjectRuntimeConfigIssueMetadataCodexFormattingUnchanged called
t.Parallel() while mutating the package-level runtimeGOOS to drive the
windows/linux branches, racing with the other parallel tests that read
runtimeGOOS in buildMetaSkillContent. The -race flag enabled in the
prior commit surfaced it as 3 WARNING: DATA RACE reports and 11
"race detected" failures in CI (only the execenv package failed).

Drop t.Parallel() and add the "// Not parallel: mutates the package-level
runtimeGOOS." comment already used by the six sibling writer tests across
execenv_test.go and reply_instructions_test.go. This is test-isolation
only; no production code, no mutex/atomic, no signature change.

Verified locally:
  go test -race -count=1 ./internal/daemon/execenv/   -> ok 2.276s
  go test -race -count=1 ./internal/daemon/...        -> all 3 pkgs ok

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: hzz <331380069@qq.com>
Co-authored-by: multica-agent <github@multica.ai>
2026-06-18 15:50:24 +08:00
Bohan Jiang
1279f22d1c MUL-3325: add background task safety brief (#4257)
* fix(daemon): add background task safety brief

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): force Claude background tools foreground

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): narrow Claude async launch detection

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-06-18 10:45:51 +08:00
Multica Eve
18a58e80c0 MUL-3316: fix(execenv): switch agent prompt to --content-file to prevent heredoc flag swallowing (#4182) (#4191)
* fix(execenv): switch agent prompt to --content-file to prevent heredoc flag swallowing (#4182)

The Linux/macOS reply template recommended --content-stdin with a quoted
HEREDOC. That pattern is safe for the trivial single-flag comment-add case
that BuildCommentReplyInstructions emits, but as soon as a model wraps
extra flags around the heredoc on multica issue create / update — assignee,
project — the bash heredoc/flag boundary is fragile in two ways the model
cannot see:

  - A 'BODY \\' terminator with a trailing token is not recognised as the
    heredoc end, so flag lines after it are swallowed into the description
    (OXY-78: residual flag text leaked into the description, command exit 0).
  - A clean terminator turns the trailing '--assignee ...' line into a
    separate failing shell statement, while the create itself already exited
    0 with no assignee (OXY-76: assignee silently dropped, no residual text).

In both cases the CLI never receives the swallowed flags, the API request
omits the fields, and the daemon has no visibility. The created issue lands
with assignee_id: null / project_id: null.

This commit:

  * Switches the Linux/macOS branch of BuildCommentReplyInstructions to
    --content-file with a 3-step recipe (write file, post, rm) so the body
    never reaches the shell and all flags live on one shell-token line.
    There is no heredoc boundary for flags to leak across.
  * Adds a parallel cleanup step (Remove-Item) to the Windows branch so the
    cross-platform template is one shape.
  * Rewrites the runtime_config.go ## Comment Formatting non-Windows section
    to mandate --content-file and explicitly ban --content-stdin HEREDOC for
    agent-authored comments, citing #4182.
  * Reorders the Available Commands menu lines for issue create / update /
    comment add to put --content-file / --description-file ahead of the
    stdin variant and add a per-line note pointing at #4182.
  * Updates and renames the affected tests
    (TestBuildCommentReplyInstructionsCodexLinux,
    TestBuildCommentReplyInstructionsNonCodexLinux,
    TestInjectRuntimeConfigLinuxCommentFormattingEmphasizesFile,
    TestInjectRuntimeConfigIssueMetadataCodexFormattingUnchanged) so the
    new file-first contract is pinned and the old HEREDOC mandate is in the
    banned-strings lists.

This converges Linux/macOS with the long-standing Windows file-only path,
so the cross-platform guidance is now one shape. It also strictly improves
on the previous MUL-2904 guardrail by eliminating shell exposure of the
body entirely (no body ever reaches the shell, so backtick / $() / $VAR
substitution cannot corrupt it).

Closes GitHub multica-ai/multica#4182.

No CLI or backend changes — --content-file / --description-file already
exist.

Co-authored-by: multica-agent <github@multica.ai>

* docs(prompt): correct stale BuildPrompt comment to file-first (#4182)

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: CC-Girl <cc-girl@multica.ai>
2026-06-16 17:14:25 +08:00
Bohan Jiang
24b162cdbc feat(daemon): surface the real task initiator to the agent runtime (MUL-2645) (#3899)
* feat(daemon): surface the real task initiator to the agent runtime (MUL-2645)

In a multi-person workspace the agent runtime only ever saw the runtime
OWNER identity: the brief's `## Requesting User` is sourced from
runtime.OwnerID and the task-scoped token is owner-bound, so every
requester (whoever commented, @mentioned, or chatted) appeared to the
agent as the owner. Agents that route by initiator for permission,
privacy, or audit all misjudged.

Resolve the real task initiator at claim time and surface it distinctly
from the owner:
- comment / mention trigger -> triggering comment's author (member or agent)
- chat task -> chat session creator (sessions are creator-only)
- on-assign / autopilot / quick-create -> no attributable initiator (omitted)

Adds initiator_{type,id,name,email} to the claim response, the daemon
Task, and TaskContextForEnv, rendered into the brief as a new
`## Task Initiator` section. The section documents the privacy boundary:
the agent's credentials stay owner-scoped, so this is an attested
identity for the agent's own routing/privacy logic, not act-as. No DB
migration — both paths are derivable from existing rows.

Tests: brief rendering (member/agent/omit/sanitize) + email guard unit
tests, and claim-handler tests for the comment and chat paths.

Co-authored-by: multica-agent <github@multica.ai>

* fix(chat): store real sender as task initiator, not chat_session creator (MUL-2645)

Review fix (Niko, PR #3899). v1 resolved the chat task initiator from
chat_session.creator_id at claim time. That is correct for web chat and
Lark p2p (creator == sender), but WRONG for Lark group chats: the group
session creator is deliberately the installer (stable identity across
member churn), not the message sender. So in a Lark group, every member
who triggered the agent showed up in the brief as the installer/owner —
the exact bug this issue is about, still live at that entry point.

Capture the real sender at enqueue time instead of deriving it from the
session creator at claim time:

- migration 117: agent_task_queue.initiator_user_id (FK user, ON DELETE
  SET NULL); NULL for non-chat and pre-migration rows.
- EnqueueChatTask now takes an explicit initiatorUserID. Web chat passes
  the authenticated request user; the Lark dispatcher threads the inbound
  sender (binding.MulticaUserID) through scheduleRun -> flushChatRun. The
  debouncer keeps the latest scheduled flush per session, so in a multi-
  sender silence window the LATEST sender wins (documented + tested).
- claim handler resolves the initiator from task.initiator_user_id and
  drops the creator_id fallback entirely.

The Lark group session creator stays the installer (unchanged) — only the
task initiator is corrected, keeping the two concepts cleanly separate.

Tests: dispatcher group regression (initiator = sender, not installer),
latest-sender-wins, p2p initiator assertion; the chat claim handler test
now sets creator != initiator and asserts the stored sender wins.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-06-08 19:29:57 +08:00
Naiyuan Qing
b9334dd59f fix: anchor comment triggers to thread roots (#3746)
Co-authored-by: multica-agent <github@multica.ai>
2026-06-04 13:47:05 +08:00
Bohan Jiang
d6a556bdbf fix(execenv): refresh skills in place on reuse instead of accumulating duplicate dirs (#3716)
Re-dispatching the same agent on the same issue reuses the persistent workdir
via execenv.Reuse(), where the standard-provider skill refresh re-wrote skills
without clearing the prior dispatch's output, so allocateCollisionFreeSkillDir
dodged Multica's own directories into issue-review-multica-N.

On reuse, reclaim the platform-owned managed skill directories the prior
manifest recorded (removeReusedManagedSkillDirs) and roll back the remaining
sidecar files (CleanupSidecars) before refreshing, so each skill lands at its
canonical slug every dispatch. Mirrors the Codex hydrateCodexSkills wipe;
scoped to reuse, which never runs for local_directory tasks.

Fixes #3684 (MUL-2963).
2026-06-03 19:30:42 +08:00
Bohan Jiang
888186b183 fix(daemon): make comment-posting guardrail provider-agnostic (MUL-2904) (#3654)
* fix(daemon): make comment-posting guardrail provider-agnostic (MUL-2904)

Agents inlining a backtick-wrapped token into `multica issue comment add
--content "..."` had the shell run it as a command substitution, silently
deleting the token; the stored comment never matched the model's intent, so
it retried forever — spamming OKK-497 with duplicate comments.

The corruption is shell-driven, not provider-driven, so extend the
"never inline --content; use --content-file / quoted-HEREDOC --content-stdin"
rule from Codex-only to ALL providers:

- BuildCommentReplyInstructions: collapse the Linux/macOS non-Codex inline
  branch into the unified quoted-HEREDOC stdin template.
- buildMetaSkillContent: rename "Codex-Specific Comment Formatting" ->
  "Comment Formatting" and emit it for every provider; strengthen the
  Available Commands entry and the assignment step-6 examples to steer away
  from inline --content.
- Windows behavior unchanged (file-only; avoids PowerShell ASCII drop).

Tests: flip the non-Codex Linux reply test into a MUL-2904 regression,
broaden the stdin-emphasis test across providers, and pin the
provider-agnostic guardrail.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): keep Windows assignment brief file-only (address review)

Review catch on #3654: the previous commit added platform-agnostic prose
recommending "--content-file or --content-stdin" in the Available Commands
entry and the assignment-triggered step-6 example. The assignment path has
no BuildCommentReplyInstructions OS override, so on Windows an agent following
step 6 literally would pipe its final comment through PowerShell and drop
non-ASCII bytes (#2198 / #2236 / #2376) — contradicting this PR's own
Windows file-only rule in the ## Comment Formatting section.

Make the platform-agnostic surfaces defer to the OS-aware ## Comment
Formatting section (the single source of truth) instead of naming stdin.
The flag synopsis still lists all three modes.

Add TestInjectRuntimeConfigWindowsAssignmentBriefStaysFileOnly: a Windows
assignment-triggered brief must not contain any prescriptive "... or
--content-stdin" recommendation.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-06-02 17:47:09 +08:00
Naiyuan Qing
3c8645e546 feat(cli): add squad member set-role (#3583)
Co-authored-by: multica-agent <github@multica.ai>
2026-06-01 12:51:15 +08:00
Naiyuan Qing
4ae4722ef0 fix(comments): preserve direct parent on replies (#3579)
Co-authored-by: multica-agent <github@multica.ai>
2026-06-01 08:28:15 +08:00
Naiyuan Qing
973a43923f fix(comments): revert since-delta to issue-wide, steer to parent thread first (#3535)
#3509/#3523 scoped the comment-trigger since-delta count to the triggering
thread, so an agent resuming a busy issue only saw "+N in this thread" and
lost visibility of new comments in other threads. Revert the count to
issue-wide (every thread), keeping the trigger-comment + agent-own
exclusions, and reshape the warm-path hint to:

  - report the issue-wide new-comment volume,
  - steer the agent to read the triggering (parent) thread FIRST
    (`--thread <trigger> --since`, or `--tail 30` for full context),
  - demote the issue-wide `--since` catch-up to an only-if-needed fallback
    ("don't read them all blindly").

Also fixes the now-stale "scoped to the triggering thread" wording in the
resumed-session no-delta hint (it's issue-wide zero now).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>
2026-05-29 20:13:23 +08:00
Multica Eve
d1c7d478e1 MUL-2785: clarify thread-scoped comment delta (#3523)
Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-05-29 17:16:58 +08:00
Multica Eve
9616d78e47 MUL-2785: optimize resumed comment reads (#3509)
* feat(comments): skip default thread read on resumed comment sessions

Co-authored-by: multica-agent <github@multica.ai>

* fix(comments): scope since delta to trigger thread

Co-authored-by: multica-agent <github@multica.ai>

* chore(comments): address thread delta review nits

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-05-29 14:57:14 +08:00
Naiyuan Qing
3187bbf90c feat(comments): re-add since-delta + cold-start thread read + parent-root write normalization (#3494)
* feat(comments): since-delta new-comment hint + default-on comment session resume (#3432)

* feat(db): add unresolved comment count + list filter queries

Add CountUnresolvedComments (excludes the agent's own comments) and
ListUnresolvedCommentsForIssue. Both are additive — existing callers stay
on the unfiltered queries — so old clients are unaffected.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(handler): support unresolved-only comment listing

Wire an additive `unresolved` query param into ListComments. Defaults off
so an old CLI that never sends it gets unchanged behavior; only true/1
enable it. Rejects combining unresolved with thread/recent (whole-issue
filter vs navigation models). Includes filter + count query tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(handler): plumb unresolved count + thread root into claim, gate comment resume

Populate trigger_parent_id (thread root of the trigger comment) and
unresolved_count (excludes the agent's own comments) on comment-triggered
claim responses. Both fields are omitempty so old daemons ignore them.

Gate comment-triggered session resume behind MULTICA_RESUME_COMMENT_SESSION
(default off): resumed comment turns can inherit the prior turn's "Done."
final message, so this stays an explicit rollout switch. The runtime-match
and poisoned-session guards still apply regardless of the flag.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(daemon): inject unresolved-comments hint + resolve step into agent brief

Add a shared BuildUnresolvedCommentsHint helper rendered on both the
per-turn prompt and the CLAUDE.md workflow (kept in sync per PR #2816). It
ships only the count and the relevant CLI call — never comment bodies — so
the server stays cheap. Thread case points at --thread <root>; issue case
points at --unresolved. Suppressed when the count is 0.

Also add a workflow step telling the agent to `multica comment resolve
<thread-root>` once a thread is fully handled, so the unresolved set
converges.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(cli): add comment list --unresolved and comment resolve command

Add an --unresolved filter to `issue comment list` (wired to the server's
unresolved param, rejected when combined with --thread/--recent) and a
top-level `comment resolve <id>` command that POSTs to the existing
/api/comments/{id}/resolve endpoint, letting an agent close threads it has
fully handled.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(comments): since-delta new-comment hint + default-on comment resume

Simplifies the comment-triggered agent flow down to what's actually needed:

- New-comment awareness is now a pure time delta: the claim response carries
  new_comment_count + new_comments_since (anchored on the prior run's
  started_at, never completed_at so a long run can't miss comments). The
  per-turn prompt and CLAUDE.md workflow render one line — "N new comment(s)
  since your last run, --since <ts>" — via a shared BuildNewCommentsHint so the
  two surfaces can't drift. Cold start (no prior run) falls back to a plain read.
- Comment-triggered tasks resume the prior session by default (same runtime),
  dropping the MULTICA_RESUME_COMMENT_SESSION rollout gate. The "Focus on THIS
  comment" prompt guard defends against inheriting the prior turn's "Done."
  marker; GetLastTaskSession still excludes poisoned sessions.
- Drops the resolved-based machinery from the first draft: CountUnresolvedComments
  / ListUnresolvedCommentsForIssue queries, the `comment list --unresolved`
  flag, the `multica comment resolve` command, and the resolve workflow step.
- Removes the verbose cursor-pagination paragraph from the comment prompt; the
  --thread/--recent/--since flags stay in the CLI/API, just no longer explained
  inline every turn.

Compatibility: new claim fields are omitempty (old daemons ignore them).
Comment resume is default-on and affects even old daemons, which already
consume prior_session_id.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(comments): collapse reply parent_id to thread root on write

Comment threads are a 2-level model (root + flat replies, like Linear/Slack),
enforced today only by the UI and the agent path — the CreateComment handler
stored whatever parent_id it was handed, and the agent-side flatten walked just
one level, so a reply-to-a-reply could land at depth 3+. Add GetThreadRoot (a
recursive walk to the parent_id=NULL root) and run both write paths
(handler.CreateComment, service.createAgentComment) through it, so every stored
reply's parent_id IS its thread root. Readers can now treat parent_id as the
thread root without re-walking. The agent-drift guard still compares the raw
parent_id to the trigger comment before normalization.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(comments): cold-start reads triggering thread, warm keeps --thread pointer

The since-delta rework dropped the thread-first read on the COLD path: a
first-time agent fell back to the flat `comment list` dump (oldest-first, cap
2000), burying the trigger's context in ancient chatter. Point cold start at the
triggering conversation instead via a shared BuildColdCommentsHint
(`--thread <trigger> --tail 30` + a --recent pointer for cross-thread
background). On the WARM path, --since is a pure time delta and can miss the
triggering thread's pre-anchor history, so BuildNewCommentsHint now also emits a
--thread pointer. Both surfaces (per-turn prompt + CLAUDE.md workflow) render via
the shared helpers so they cannot drift (PR #2816 rule).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-29 10:38:37 +08:00