mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-09 06:21:51 +02:00
codex/mika-issue-first-onboarding
116 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
62a2d70afd |
refactor(daemon): stage-2 — Mentions and Comment Formatting judgment rewrite (MUL-5442) (#6453)
* refactor(daemon): rewrite Mentions and Comment Formatting to their judgment form (MUL-5442) Stage 2, final two sections, bundled per the small-sections rule. Mentions: the four-link side-effect table stays verbatim (platform facts); the two H3 subsections merge into one policy paragraph keeping every anti-loop anchor — the no-mention default with its cost mechanism, the no-sign-off-mention ban, the end-with-no-mention rule, the three mention-warranted cases, and the silence closer. The retired headings' pins re-anchor to the policy phrases. Comment Formatting: both variants keep the full operational contract (file-first sentence verbatim, both bans with incident ids, workdir scope, --parent continuity, cleanup, newline rule); what goes is the mechanism narration (what the shell rewrites, how flags get swallowed, PowerShell version/encoding detail — the consequence stays). Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): correct the PowerShell version claim, pin the mention scope qualifiers, section-scope the formatting assertions (MUL-5442) Review catches by Elon on #6453, all three accepted: 1. The compressed Windows rationale over-generalized a version-specific fact: '$OutputEncoding drops non-ASCII' is true of Windows PowerShell 5.1 (ASCII default), false of PowerShell 6+ (utf8NoBOM). Now reads 'Windows PowerShell 5.1 ... may replace non-ASCII characters with ?'; the Go comment documents the version split and why file-first stays version-agnostic (agents cannot rely on which shell services the pipe). 2. The merged Mentions paragraph was pinned only at the list head — the scope qualifiers ARE the anti-repeat-notify boundary: 'not yet involved', 'for the first time', 'explicitly asks to loop someone in', and the loop-cost mechanism are each pinned individually now. 3. The Comment Formatting assertions ran against the whole file, where '#4182' also appears in Available Commands — the HEREDOC ban could vanish with green tests. The assertions now slice the section (matched at the line-start heading, since Available Commands references the heading inline) and cover all seven contract elements within it. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
f44df04f96 |
refactor(daemon): rewrite the Output section to its judgment form (MUL-5442) (#6425)
Stage 2 section 2. The delivery contract keeps every platform fact and boundary: comment-add as the only delivery channel, the invisible terminal, exactly one comment per run before turn exit, --attachment as the only file path, the runtime-local-path ban with the code-location form and the say-so-in-words fallback. What goes is derivation: the invisible-task consequence restatement, the plans-in-your-reasoning elaboration, the good/bad style examples, the exists-right-now clause, and two quick-create explanation tails. One pin re-anchor: 'Do not assume any workspace issue prefix' follows the rewrite to 'never assume a workspace issue prefix'. Every other Output and delivery-invariant pin passes unchanged, including the MUL-4899 trio and the anti-dangling pointer target the Attachments section names. Issue-kind Output: 1,493 -> 978. Brief (real-UUID fixture): 13,382 -> 12,907 (-475); quick-create variant sheds ~160 more on its own kind. Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
2bb13667c6 |
refactor(daemon): cross-channel command dedup — brief states the loop, per-turn carries the commands (MUL-5442) (#6421)
* refactor(daemon): move the workflow's command templates to the per-turn channel (MUL-5442) Cross-channel dedup, brief side. MUL-5721 measured that every issue-turn variant of the per-turn message already carries the ready-to-run commands with real ids (issue-get line, reading hints, reply cookbook), fresher than the brief's static copies. The six workflow steps now state the loop shape — command names and flag mnemonics only; full templates and every embedded issue UUID leave the brief. The Ownership status commands and the squad-activity call switch to <issue-id> placeholder form. Doctrine pins stay on the brief side (mandatory catch-up, scan-first order, both motivation anchors); command-template pins re-anchor to names and mnemonics. Two design-consequence test updates: the static-catch-up assertion repoints at the doctrine (the full command now lives in every per-turn variant), and the byte-identity non-vacuity guard varies agent identity instead of issue id — because the brief is now deliberately issue-id-independent, which opens the cross-issue prefix-cache door. Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): point the reply cookbook at Comment Formatting instead of restating it (MUL-5442) Cross-channel dedup, per-turn side. The cookbook's prose restated the brief's Comment Formatting rules (file-first rationale, inline-content and HEREDOC hazards, the full Windows $OutputEncoding mechanics) around the command that already demonstrates the shape. It now keeps the file-first order, the ready-to-run command, the literal-newline rule, and the pointer; the hazard mechanics live once, in the brief section the pointer names. Default cookbook: 639 -> 494 bytes (-145 per reply turn, uncacheable channel); the Windows variant sheds ~200 more. Pins re-anchor from the deleted prose to the surviving anchors (file-first order, the pointer, the command form); the MUL-2904/#4182 banned-shape negative guards are untouched. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): narrow the per-turn coverage claim, test the cross-issue invariant, strip the command shape from the guest-visible prohibition (MUL-5442) Review catches by Elon on #6421 plus the CI failure, all three related: 1. The steps header overstated what the per-turn message carries — it ships the issue id and ready-to-run context-read commands; other calls are assembled from Available Commands. Reworded to say exactly that (a factual claim in a prompt cannot rely on the reader discovering it is wrong). 2. The cross-issue byte-equality this PR claims as a design benefit is now asserted directly per provider (same stable inputs, different 36-char issue UUIDs, byte-identical brief) — a Contains-based negative cannot catch truncated/transformed/conditional id use. 3. CI: the squad guest-leader contract test bans any runnable 'issue status <issue-id> in_review' shape in guest-visible text; the placeholder rewrite made the leader dispatch rule's NEGATIVE sentence match that ban. The prohibition now states itself without a command form ('do NOT move it to in_review or done on this turn') — negative sentences should not carry copy-pasteable command shapes at all. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
b98f760f1f |
MUL-5604: 支持 Reasonix runtime (#6370)
* feat(agent): add Reasonix runtime Co-authored-by: multica-agent <github@multica.ai> * fix(agent): make Reasonix sandbox host-adaptive Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
c3aab05a38 |
refactor(daemon): rewrite Background Task Safety to its judgment form (MUL-5442) (#6381)
* refactor(daemon): rewrite Background Task Safety to its judgment form (MUL-5442) Owner-authorized judgment rewrite per the Claude 5 context-engineering guidance: state the platform fact and the boundaries an agent cannot infer, drop the enforcement details a frontier model derives. The section is now three paragraphs: (1) the fact everything derives from — turn exit is task-terminal, no wakeup exists, never background-and-yield — with the foreground-collect, synchronous-fallback and no-standing-by consequences stated once; (2) the external-systems/CI boundary: not run-owned, the named --watch ban (kept because MUL-5223 proved the principle alone did not stop CI-watching), merge-gate is not acceptance criteria, the hand-off phrasing, and the single explicit-ask exception; (3) the persistent-service handoff contract, review-locked verbatim. Deliberately dropped as derivable: the run-owned work enumeration, the tool-promise enumeration, the wait/collect split rule, the persistent-service scope bullet, auto-merge and snapshot elaborations. Both pin lists rewritten to the surviving semantic anchors (21 each); retired pins are listed in the test comments with the rationale. 2,980 -> 1,931 (-1,049). Incident history stays in the Go comment where it costs no context. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): delete the duplicated CI exception, pin the full ban clause, sync the comment ledger (MUL-5442) Review catches by Elon on the judgment rewrite, all three accepted: 1. The old exception bullet survived below the handoff paragraph — the CI exception appeared twice, the second copy scope-widened by its position, and every substring pin stayed green. Deleted; both BTS tests now count-guard exactly one 'The one exception' occurrence. 2. The compound watch/poll ban was pinned only by its first command; the full clause is the MUL-5223 boundary, so the pin now covers all three members. 3. The Go comment ledger still described the pre-rewrite bullet model (retired pins as kept, a forward reference that no longer exists). Rewritten to the three-paragraph model so the safety-decision history reads true. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
18f8f78af5 |
refactor(daemon): per-turn message slimming — OPT-1 UUID dedup + OPT-2 initiator prose (MUL-5721) (#6382)
* refactor(daemon): drop duplicate UUID commands from per-turn comment hints (MUL-5721) OPT-1 from the MUL-5721 measurement: within one hint, a second full `multica issue comment list <uuid> ...` command restated the issue UUID (and, in the resumed hint, both anchors) for no routing value. The warm hint's issue-wide catch-up and the cold hint's roots scan become flag swaps on the thread command they follow; the resumed hint loses its anchor-restating sentence (the read command carries the thread anchor, the reply cookbook carries the trigger id as --parent). Fixture-measured: warm hint 555->481, resumed 551->417, cold 429->396; reply-turn totals -74/-134/-33 B. Ownership prompt and the reply cookbook (with its concrete --parent command) are untouched. Every dropped duplicate gets a deny assertion so it cannot silently return; the demoted fallbacks get semantic anchors in its place. Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): compress Task Initiator attribution prose (MUL-5721) OPT-2: the block's 400 B attribution paragraph carried two derivable restatements ("in a workspace many people can reach", "so this attribution does not widen what you can read or write"). The compressed paragraph keeps every rule: per-person privacy/access rules apply, the initiator (not the runtime owner) is who you are answering, credentials stay scoped to the runtime owner, and the initiator may not see everything you can. Both MUL-2645 test-pinned phrases survive verbatim; the fact line and heading are untouched. Fixture-measured: block 513->370 B (member+email form). The two previously unpinned surviving rules gain anchors ("is who you are answering", "do not assume the initiator can see everything you can"). Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): restore initiator read/write authorization boundary (MUL-5721) Review must-fix (#6382): the compression treated "attribution does not widen what you can read or write" as derivable from credential scoping and dropped it. It is an independent boundary: credential scoping names whose credentials run, the visibility sentence constrains disclosure — neither stops the model from treating an initiator's request as extra authorization. Restored as a short explicit negation with a semantic anchor. Initiator block 370 -> 440 B (still -73 vs the original 513); typical reply turn 2,191 B (net -147 for the PR). Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
1e661bb04b |
refactor(daemon): stage-1 judgment rewrite — mode router and small guardrail sections (MUL-5442) (#6365)
* refactor(daemon): compress the mode router and Reply-block preamble to their facts (MUL-5442) Stage 1 batch A of the judgment rewrite. The mode-router paragraph spent ~800 bytes explaining WHY the file cannot know the turn mode; the fact an agent needs is only WHERE the mode is named and that exactly one mode block applies. Compressed to a single Turn-mode lead carrying every routing fact: both mode lines, the shared-steps note, the one-block rule, and the no-mode-line default. The Reply block's first bullet drops its derivations (session history, confusion warnings) and keeps the id-source rule; the reply-warranted bullet keeps both pinned anchors and every behavioural outcome, dropping only the worked examples. Re-pins: the 'Mode router' heading anchor becomes '**Turn mode.**'. Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): trim the mechanism prose from the small guardrail sections (MUL-5442) Stage 1 batch B. Comment Formatting (both variants), Issue Body Formatting, and Always Use the CLI keep every ban, every incident id, and every pinned phrase verbatim; what goes is the mechanism narration (what the shell rewrites, why the CLI rejects outside paths, the resource-type enumeration). Zero test changes — every pin survives as-is. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): restore two platform boundaries and pin the routing rules (MUL-5442) Stage-1 review catches by Elon on #6365: 1. Two behavioural boundaries had been deleted as explanation. The issue-body H1 rule narrowed 'body or description' to 'body' — but 'description' is the CLI/API field name, so the alias is a cross-surface mapping. And step 2 lost 'CLI failures are normal' — platform failure semantics (a failed metadata read must not block the task), not tool mechanics. The second deletion also went undisclosed in the PR description. Both restored; both now pinned. 2. The compressed Turn-mode lead was guarded only by its markers: every routing rule (shared steps, apply-exactly-one, status difference, no-mode-line fallback, no status change) could have been dropped with green tests. All five now pinned as semantic anchors, and both Reply outcomes (work-produced arm, silent-exit arm) pinned individually alongside the existing heading anchors. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): pin the full no-mode-line mapping and gofmt the touched tests (MUL-5442) Final-review catches on #6365: runtime_config_kind_test.go:99 was not gofmt-clean (a scripted insertion), and the no-mode-line fallback was pinned as two separable halves — 'No mode line' plus a 'Reply mode' that other file content satisfies — so the text could mutate to route the fallback at Ownership mode with green tests. The anchor is now the exact mapping clause 'No mode line → Reply mode'; the status anchor stays. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
dc60429366 |
refactor(daemon): slim Issue Metadata to its judgment core, defer discipline to the skill (MUL-5442) (#6351)
* refactor(daemon): fold the metadata ban list into the write rule (MUL-5442) The Issue Metadata section spent 1,024 bytes teaching a KV bag. The What-NOT-to-pin bullet was a separate heading restating the write rule's negative space; it folds into Write on exit as one sentence with every ban kept explicitly (secrets/tokens/API keys, logs or comment summaries, runtime bookkeeping, single-run details). The parenthetical examples and connective prose go; every rule stays. Pins updated in kind: the merged ban list is pinned as one sentence plus the result-comment redirect, so the next compression pass cannot drop a ban without failing CI. -202 bytes (16,099 -> 15,897 on the standard fixture). Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): restore the runtime-bookkeeping examples that define the ban (MUL-5442) Review catch by Elon on #6351: 'runtime bookkeeping' has no other definition anywhere in the brief or skills, so its examples (attempts, run timestamps, agent IDs) are the category's boundary, not connective prose — and all three are high-probability miswrites that can look like re-readable diagnostics or durable facts. Restored inside the merged bullet. Also adopts the review's pin advice, which is our own anchor-pin doctrine applied properly: the whole-sentence pin from the previous commit is replaced with separate semantic anchors (each ban plus the bookkeeping examples plus the result-comment redirect), so the sentence can be reworded later without CI churn while dropping any single ban still fails. +52 bytes; the section lands at 874 (from 1,024), the PR nets -150. Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): defer the metadata write discipline to the working-on-issues skill (MUL-5442) Owner decision on MUL-5442: metadata is deliberately free-form custom key-value state — the recommended-keys block never matched the feature's intent and is removed outright, not relocated. The full ban list defers to the multica-working-on-issues skill, which already carried a near-complete copy; the skill gains the two bans it lacked (secrets/tokens/API keys, agent ids) so nothing loses its home. The brief keeps only what the interface cannot express: the read-as-hints stance, the will-a-future-run-re-read-it bar, and the two write-time boundaries (never secrets, never long content). Section: 874 -> 505 bytes (1,024 at the round's start). Pins move with the content: the brief side slims to the surviving semantics plus the skill pointer; the skill-side contract test gains the relocated ban anchors so the pointer cannot dangle. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): remove the curated key list from the skill and restore the ban categories (MUL-5442) Round-3 review catches by Elon on #6351, both accepted: 1. The recommended-keys concept survived in the skill — which loads exactly when an agent is about to write metadata, so it still steered free-form KV into a platform vocabulary. The owner's ruling was that the concept should not exist, not that it should move. The key list, the 'high-signal keys only' heading, and the pr_url-specific example are gone; the section now describes free-form durable custom state and a generic set example. A mustNotContain guard keeps the curation from creeping back. 2. The relocated ban list had dropped its two defining categories (runtime bookkeeping, other single-run details), leaving unlisted values looking writable. Restored with the reviewer's structure: each category names its examples, and the test pins every category AND every example as separate line-safe anchors — no item can be silently dropped again. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): remove the last key-vocabulary residue from the property-vs-metadata note (MUL-5442) Round-4 review catch by Elon on #6351: the property-vs-metadata bullet still read 'free-form scratchpad for run state (pr_url, waiting_on, ...)' — recommending the ruled-out fields AND calling metadata a home for run state, in direct conflict with the runtime-bookkeeping ban restored two sections above. Now reads 'free-form bag for durable custom issue state', consistent with both the owner ruling and the ban list. Test side per the review: the broad pr_url anchor narrows to the full stale-warning phrase (the one sanctioned pr_url reference — a negative compatibility note, not a write recommendation), and the curation guard gains the removed residue phrases so neither the vocabulary nor the run-state framing can creep back. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
766c7a7fa2 |
refactor(daemon): retreat deep read semantics to --help, trim BTS prose (MUL-5442) (#6347)
* feat(cli): carry the comment-read contract in comment list --help (MUL-5442) The --recent flag help gains the MUL-5372 saturation semantics (N caps threads, not comments; every thread returns uncapped; small issues return the whole history) and a pointer at the bounded alternative. The --before help names the stderr cursor labels alongside the response header. This is the relocation target for the brief's deep read semantics: the contract follows the flag, and TestIssueCommentListHelpCarriesReadContract pins it here so the brief-side pointer cannot dangle. Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): retreat deep read semantics to --help and trim BTS prose (MUL-5442) Unblocked by #6309: the platform no longer recommends --recent anywhere, so the flag's deep semantics no longer need to ride in every brief. - Available Commands comment-list entry: keep the signature, the two bounded read shapes, and the one-clause saturation warning; per-thread cap detail, folding rules, and cursor labels retreat to --help (pinned there by TestIssueCommentListHelpCarriesReadContract). - Workflow step 3: drop the worked examples and the redundant trailing pointer; the mandatory two-step read and both motivation pins stay. - Background Task Safety: tighten connective prose only — every pinned phrase and every behavioural constraint is untouched. - multica-squads SKILL.md quick-start: replace the last remaining --recent recommendation on the platform with the bounded scan — same class of fix as #6309, found while relocating the warning. -676 bytes on the standard fixture (16,766 -> 16,090). Co-authored-by: multica-agent <github@multica.ai> * fix(cli,daemon): repair the help rendering, restore the BTS orphan scope, complete the squads read (MUL-5442) Three review catches by Elon on #6347: 1. The --before usage string wrapped the cursor labels in backticks, and pflag's UnquoteUsage hijacked the first pair as the flag's value placeholder — rendered help showed '--before Next thread cursor' instead of '--before string' (same regression class as TestLoginTokenHelpOutputRendersCleanly). Labels now use double quotes, and the --recent help no longer suggests composing mutually exclusive flags ('--roots-only + --thread --tail' -> an explicit two-step read). TestIssueCommentListHelpCarriesReadContract now asserts on the RENDERED FlagUsages output, pinning '--before string' and the two-step order. 2. The compressed BTS opening said 'anything still running is orphaned', which swept externally-owned work (GitHub Actions) into the orphan rule the same section later scopes out. Restored: 'any run-owned work still active is orphaned', pinned. 3. The squads quick-start's roots-only scan never returns reply bodies, where mention triggers and failure reasons usually live. Added the bounded drill-down step and the sequence rationale; both reads pinned in TestSquadsSkillCoversLeaderRoutingContract, which also guards against a regression back to --recent. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
e92d828f0c |
refactor(daemon): compress brief prose and demote the sub-issue playbook (MUL-5442) (#6310)
* refactor(daemon): compress brief prose and demote the sub-issue playbook (MUL-5442) Six static edits to the issue brief, -1,249 bytes on the standard fixture (17,977 -> 16,728). All edits are identical for every issue run; prompt-cache byte stability (#6008) is untouched. - BTS persistent-service bullet: keep the handoff contract (deliverable-only, lifecycle detached, readiness verified, best-effort survival), drop the operational walkthrough. Pins re-anchored to concepts (durable logs, recorded PID, verify readiness). - BTS CI-ban bullet: keep the command blacklist and the complete-hand-off rule, drop the auto-merge/snapshot elaboration. - Ownership mode: state the identity-forbids clause once on the header instead of once per status bullet. - Delivery invariant: merge the three sub-bullets into the lead paragraph; also fixes the stale '(below)' — the per-surface delivery line renders above the invariant, not below it. - Sub-issue Creation: demote the todo/backlog/stage playbook to the multica-working-on-issues skill; the brief keeps a one-line flag map plus the skill pointer. Skill-side anchors added to TestWorkingOnIssuesSkillCoversIssueLoopContracts so the pointer cannot dangle. - Attachments: collapse the two-sentence CLI-fetch intro into one line. Every pinned behavioral phrase is either carried verbatim or re-pinned to an equivalent semantic anchor in the same assertion; no assertion is deleted without a replacement. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): restore the full persistent-service handoff contract (MUL-5442) Review catch by Elon on #6310: the compressed bullet weakened the contract two ways — the reply requirement dropped 'logs' from the URL/logs/stop triple (durable logs alone are unobservable if the user is never told where they are), and the general ownership/cleanup handle narrowed to a bare PID (a supervisor/profile-managed service has no single stable PID). Both test pin sets had been updated to the weakened phrases, which would have made the regression look legitimate. Keeps the compressed sentence shape; restores both halves of the contract and pins them ("cleanup handle such as PID/profile", "URL, logs, and stop instructions") so they cannot be compressed away again. +38 bytes; the PR still nets -1,211 against main. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
3c6176cbcd |
refactor(daemon): de-duplicate cross-section rules in the runtime brief (MUL-5442) (#6302)
* refactor(daemon): fold four restatements of the end-of-turn rule into one (MUL-5442) Background Task Safety opened with four bullets that were four views of the same rule -- do not end the turn with run-owned work outstanding: the general ban, the "tool says it will notify you" case, the unobservable-result case, and the "standing by" sign-off. Separating them cost bytes without adding a distinct behaviour. Fold them into the leading bullet. Every phrase the behaviour tests pin is carried over verbatim, including "Do NOT end your turn while background tasks", "Never background-and-yield", "wait for a future notification/reminder", "running in the background so you can keep working", "run the work synchronously instead" and "standing by". MUL-5442 Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): demote workflow-step restatements of delivery, mention and metadata policy to pointers (MUL-5442) The six workflow steps and the Reply mode block restated policy that already has a dedicated section. Each restatement is a second place to edit when the policy changes, which is how they drifted apart in the first place. Give each rule one canonical home and leave a pointer at the call site: - delivery ("only a comment reaches the user") -> ## Output; steps 5 and the Reply block point at it. - mention discipline -> ## Mentions. The reply-time phrasing the loop-hardening test pins moves into that section rather than being duplicated in the Reply block, so the anti-loop signal is unchanged. - metadata read/write bar -> ## Issue Metadata; steps 2 and 6 point at it instead of paraphrasing the bar. Tests: the metadata scope test pinned the old pointer wording, and the mention test's comment claimed the sign-off rule lived in the workflow steps. Both are updated to the new placement; every behavioural phrase they guard is still asserted, file-wide. MUL-5442 Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): give the comment-read surface and the file-safety rules one home each (MUL-5442) Three cross-section duplications, each resolved toward the section that owns the rule: - comment reads: the workflow step and the Available Commands entry both explained what --roots-only / --summary / --thread --tail do. The step keeps the two reads it mandates and its anti-stale motive; flag semantics stay in Available Commands, which is the single discovery point. The saturation trap ("caps THREADS, not comments") and the pagination cursor labels stay put -- they are load-bearing after MUL-5372 and remain asserted. - --content-file: the comment add entry restated the guardrail Comment Formatting owns; it now names the rule and points there for the rationale. - workdir path rule (MUL-4252): issue create carried its own copy of the stale-file rationale; it keeps the rule and defers the why. - inbound attachments: trimmed to the pinned rule plus a pointer. MUL-5442 Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): merge the Agent Identity action list, gate squad maintenance, trim the attachment restatement (MUL-5442) Three follow-ups from review, each re-examined against what the text guards today rather than against the fact that a test pinned it. 1. Agent Identity enumeration. Instruction Precedence and workflow step 4 were added in the same commit (#3802) and each carried its own list of actions Agent Identity can forbid -- and the lists disagreed: one named status changes, the other named issue create/update and delegation, neither contained the other. Merge them into Instruction Precedence, which owns the rule. Step 4 keeps only what that section cannot express: a delegation-only role stops once its delegation is delivered. 2. Squad maintenance. `multica squad member set-role` shipped to every run, including every agent that leads no squad and therefore has no squad whose roles it could change. Gate it on IsSquadLeader -- agent configuration, not per-run state, so the brief stays byte-stable across runs of one session (MUL-5377), the same predicate the workflow already branches on. 3. Inbound attachments. The section restated Output's no-clickable-local-path rule verbatim. Keep the framing Output cannot express -- a downloaded attachment feels shared but landed in a private workdir -- and point at Output for the rule. The delivery test now also asserts the pointed-at rule is present, so the pointer cannot dangle. Brief size, plain issue task: 19,231 -> 18,009 bytes for an ordinary agent (-6.4%), 20,718 -> 19,759 for a squad leader (-4.6%). MUL-5442 Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
a54662f30e |
fix(daemon): provision config-referenced Codex instruction files into task home (#6291)
* fix(daemon): provision config-referenced Codex instruction files into task home A per-task CODEX_HOME copies the user's config.toml verbatim, so a model_instructions_file reference survived the move while the file it names did not. Codex resolves the relative value against CODEX_HOME — now the task home — and failed loading its configuration before the task prompt was ever delivered (#6271). Generalize the existing model_catalog_json provisioning into a keyed table of path-valued config keys and add model_instructions_file plus its deprecated experimental_instructions_file alias. Semantics are unchanged per key: absolute/~ values are left for Codex to read directly, relative values must stay inside the task home, the copy refreshes on reuse, and a missing source fails during environment preparation with a diagnostic naming the key instead of an opaque os error 2 at thread/start. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): root-scope task-home writes for config-referenced Codex files Review must-fix: filepath.IsLocal only proves the config string has no lexical '..'. A task home is reused, and the task that ran in it can replace an intermediate directory of the copy destination with a symlink, which made the daemon's next MkdirAll/Remove/copy follow that link and delete or overwrite a file outside the task home. Do the mkdir, stale-copy removal, and write through os.OpenRoot(codexHome) so links leaving the task home are rejected (links staying inside it are harmless), and refuse outright when codexHome itself is a symlink, since OpenRoot would resolve it before confining anything below. The pre-existing model_catalog_json path is covered by the same helper. Also stop reporting every source stat failure as a missing file — a permission or IO error now says so — and document the source-side symlink policy plus the stale-copy contract. Co-authored-by: multica-agent <github@multica.ai> * test(daemon): pin stale-copy contract when a referenced Codex file is repointed Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): bind codex-home root check to the opened handle Review must-fix: checking the path with os.Lstat and then opening it with os.OpenRoot leaves a window where the directory is swapped for a link to somewhere else. OpenRoot resolves the path it is given, so the resulting root is confined to the wrong tree and every root-scoped write lands outside the task home without ever escaping its root. Task homes are reused and Windows cannot confirm descendant cleanup, so a leftover process that knows its old CODEX_HOME can create that window. Open first, then prove identity against the handle: compare root.Stat(".") with a no-follow os.Lstat of codexHome via os.SameFile, and reject a symlink outright. A swap before the open now fails the check; a swap after it cannot matter because all writes go through the verified handle. verifyCodexHomeRoot is split out so the swap is tested deterministically instead of raced, plus an end-to-end test for a task home that is already a link to an outside directory. Co-authored-by: multica-agent <github@multica.ai> * docs(daemon): scope the codex-home symlink test name and root-handle claim Review nit: the end-to-end test asserted only that the referenced-file copy refuses a symlinked task home, but its name read as though the whole prepare were safe. Rename it accordingly and state in both the test and openVerifiedCodexHomeRoot that the earlier path-addressed steps of prepareCodexHomeWithOpts are out of scope, tracked in MUL-5647. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
e4b6f7a31b |
MUL-5581: add Qoder CN CLI runtime (#6232)
* feat(agent): add Qoder CN CLI runtime Co-authored-by: multica-agent <github@multica.ai> * fix(agent): address Qoder CN review nits Co-authored-by: multica-agent <github@multica.ai> * fix(agent): defer Qoder CN version gate Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
98766cc06a |
feat(daemon): remove the Linux Codex per-task HOME; default Linux to danger-full-access (MUL-5578) (#6233)
* feat(daemon): default Linux Codex to danger-full-access on the real HOME (MUL-5578) Linux Codex tasks ran under the `workspace-write` Landlock sandbox with a generated per-task HOME: the daemon rewrote HOME/XDG_*/npm_config_cache into `<envRoot>/home`, symlinked a hand-maintained allowlist of seven credential paths back into it, and granted that directory as a `writable_roots` entry. That mechanism could not converge. Any host CLI outside the allowlist (aws, kubectl, gcloud, glab, rclone, cargo, …) started every task as if unconfigured even though it works in the daemon user's shell, and three open issues asked to extend it in three incompatible directions (#5636, #5573, #3867). The containment it bought was also narrower than it looked: workspace-write restricts writes only — reads and network were already unrestricted, and `.ssh` was seeded into the task home — so credentials under the real HOME were readable and exfiltratable regardless. Linux now runs `danger-full-access` on the daemon user's real HOME and inherited XDG environment, matching macOS and Windows. The task filesystem boundary is the boundary the daemon itself runs inside (VM, container, or dedicated Unix user), which is what the other providers already assumed — Claude Code runs with `--permission-mode bypassPermissions` today. This also removes the split-brain the mechanism could enter: the task-HOME decision read the platform default only, so a `-c sandbox_mode=danger-full-access` override (which passes arg filtering and wins over config.toml) left a task running unsandboxed while the daemon still redirected HOME and emitted writable_roots for a sandbox that was not in effect. With no HOME rewrite there is only one environment contract left to disagree about. Removed: prepareTaskHome, prepareCodexSandboxHome, TaskHomeEnv, Env.TaskHome, both seed allowlists, and the now-unfed WritableRoots plumbing through codexSandboxPolicy / CodexHomeOptions. Env roots created by older daemons keep working; their leftover `home/` directory is simply ignored and reclaimed with the env root. Task-scoped CODEX_HOME is untouched — it is managed Codex state, not the Unix HOME. macOS, Windows, and non-Codex providers are unchanged. Co-authored-by: multica-agent <github@multica.ai> * docs: document the agent execution security model (MUL-5578) The docs had no page describing what a task can reach on the machine that runs it — the sandbox posture was only discoverable from source. Adds one, stating plainly that tasks run with the full permissions of the daemon user and that isolation must come from a dedicated Unix user, container, or VM. Also separates what Multica genuinely isolates (per-task workdir, task-scoped CODEX_HOME, agent+task-bound API tokens) from what is not a boundary (the coding tool's own sandbox and approval settings), and links it from step 5 of the self-host quickstart, where the daemon is first installed. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): address review on the Linux full-access default (MUL-5578) Docs, both must-fix: - The Chinese page rendered the product concept as `Agent`; the repo glossary in developers/conventions.zh.mdx mandates 智能体. Fixed all nine occurrences. ja/ko already used エージェント / 에이전트 and needed no change. - The security page asserted without qualification that every task runs with the daemon user's full permissions and that the Codex filesystem sandbox is always off. codexSandboxPolicyForWindows still keeps workspace-write when a user explicitly opts into a native windows.sandbox, so an authoritative security page contradicted the code. It now states that Multica makes no filesystem-sandbox guarantee, names the Windows opt-in as the one current exception, and says which combinations sandbox anything is a compatibility detail that moves with tool versions. All four locales updated, plus the historical callout which now says Linux matches the macOS/Windows *default*. Both stale comments from the non-blocking notes: - prepareCodexHome no longer claims to assume workspace-write + network_access; it pins GOOS=linux, which now resolves to danger-full-access. - ensureCodexSandboxConfig's warn-level logging is no longer described as a macOS-only fallback; it fires for every danger-full-access resolution. Adds TestCodexTaskShellEnvInheritsRealHome at the daemon env-assembly layer: HOME and the XDG base dirs must reach a Codex task's shell tools from the inherited daemon environment. Verified it fails when that pass-through breaks. It guards the pass-through, not runTask's decision not to inject a HOME of its own — that decision is inline in runTask and has no unit seam. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
abdfd3e28c |
refactor(skills): make the brief's skill list a names-only index (MUL-5529) (#6207)
* refactor(skills): make the brief's skill list a names-only index (MUL-5529)
Step 3 of MUL-5529. Every runtime CLI discovers the SKILL.md files the daemon
writes and builds its own listing from their frontmatter — verified against 11
locally installed CLIs plus official docs for 5 more. The brief's copy of those
descriptions was therefore the same routing signal paid for twice: measured on
a real task, `## Skills` was 13,295 chars, 40% of the entire brief, against a
16,304-char CLI listing of the same 28 skills.
Now 850 chars for that same set — roughly 3,100 tokens back per brief.
The index itself stays. It is the one skill listing Multica controls; each
CLI's own listing is theirs, and its format — or its existence — can change
with any release.
Three changes:
- Descriptions dropped from the `## Skills` entries.
- The per-provider branch is gone. Its fallback told providers outside a
hardcoded list to read `.agent_context/skills/`, but the only providers
that ever reached it were grok and traecli, whose files are written to
`.grok/skills` and `.traecli/skills` and which discover natively. The
pointer was wrong for everyone it addressed, so removing the branch
deletes the bug rather than relocating it. This closes MUL-5537.
- issue_context.md and its quick-create / autopilot variants no longer render
`## Agent Skills`. That copy duplicated the brief once both were
names-only, and nothing ever read it: no prompt references the path, and
grepping the server finds only the writer. `.agent_context/skills/` had the
same fate for hermes (issue #5242). Quick-create, previously skipped in the
brief and served only by that unread copy, now gets the brief section like
every other kind — one index, one place.
Not included: skills carrying `disable-model-invocation` are still written to
disk for every provider. The plan assumed that key needed provider-specific
handling for everything except claude; probing the installed CLIs shows 9 of 11
honor it, and only opencode and hermes do not. The remaining question is
narrow and a genuine product tradeoff — withholding the file honors the
author's intent but also removes explicit invocation — so it is left to a
separate decision rather than folded in here.
Co-authored-by: multica-agent <github@multica.ai>
* docs(skills): align stale comments with the names-only brief contract (MUL-5529)
Co-authored-by: multica-agent <github@multica.ai>
---------
Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Steve Jobs (Multica Agent) <agent-steve-jobs@multica.ai>
|
||
|
|
5199278780 |
fix(skills): give every skill one name across brief, directory, and frontmatter (MUL-5529) (#6189)
* fix(skills): give every skill one name across brief, directory, and frontmatter (MUL-5529)
A skill could answer to three different names at once. The runtime brief listed
`AgentSkillData.Name` verbatim — a workspace skill's human display name ("PR
review") — while the invocable identity on disk is the sanitized slug
(`pr-review`). Separately, ensureSkillFrontmatter returned valid upstream
frontmatter untouched, so a SKILL.md could declare `name: multica-dev-workflow`
inside a directory called `multica-git-workflow`.
That last divergence is the sharp one: runtimes disagree on which field
identifies a skill. Claude routes on the directory name, OpenCode on the
frontmatter `name`. So the same skill is invocable under different names
depending on where it runs, and the brief's instruction to use "only names from
the listing" pointed at names that resolve nowhere.
The slug is authoritative: it is what lands on disk, it derives from the name
users see in the product, and it is the only value with a uniqueness guarantee
(allocateCollisionFreeSkillDir). A frontmatter `name` is author-supplied and two
imported skills may both claim the same one.
- modelVisibleSkills normalizes Name to the slug. All four model-visible
listings (runtime brief + the three issue_context renderers) already route
through it, so they cannot drift apart.
- ensureSkillFrontmatter rewrites `name` to the allocated slug and keeps every
other key byte-identical, so deliberately shaped upstream frontmatter still
survives. The rewrite follows the collision fallback slug too.
- Name matching is now top-level only. An indented `name:` belongs to a nested
mapping; treating it as the skill's identity both missed that the block had no
top-level name and would have spliced a top-level key into the nested one.
Known gap, tracked separately: the listings render the natural slug, so a
collision fallback to `<slug>-multica` still leaves the brief naming the user's
skill. Closing it needs the allocated slug threaded back from Prepare, which the
renderers cannot reach without giving up the byte-identical-brief guarantee.
Co-authored-by: multica-agent <github@multica.ai>
* fix(skills): cover multi-line name values and in-batch slug collisions (MUL-5529)
Two holes in the name-unification change, both found in review, both
reproduced before fixing.
1. setFrontmatterName replaced only the `name:` line, not the rest of the YAML
value. A value may continue onto indented lines — `name: >-\n upstream`,
multi-line plain scalars, wrapped quoted scalars — so the continuation
survived and YAML folded it into the new value: the block parsed as
"my-slug upstream-name", not "my-slug". The directory == frontmatter-name
invariant this change set exists to establish was still broken, just less
visibly. frontmatterNameSpan now covers the whole value.
As a side effect this also fixes `name:` with the value entirely on the
following line, which previously read as "no name" and got a second
top-level `name` injected above it — two `name` keys, which strict loaders
reject outright.
2. The listings derived slugs with sanitizeSkillName alone, which is not
injective: "A B" and "A-B" both reduce to "a-b". writeSkillFiles resolved
that at write time, so the second skill landed in `a-b-multica` while both
were listed as `a-b` — the second skill had no invocable name and the model
was pointed at the first. This needed no user-installed skill and no
local_directory; two such skills bound to one agent reproduce it in a clean
workdir. resolveSkillSlugs now deduplicates the batch up front and both the
listings and the writer derive from it, with skillSlugCandidate shared so
the in-memory and filesystem allocators cannot disagree on the suffix
sequence.
Slugs are allocated over the unfiltered batch: hidden
(disable-model-invocation) skills are still written to disk and still
consume a slug, so filtering first would shift every later suffix.
Filesystem-dependent collisions against user-installed directories remain out
of scope and are still tracked in MUL-5550.
Regression tests assert the parsed YAML value rather than the output text —
the first line looked correct in every one of these cases — and set-equality
between listed names and the directories actually written. Both were confirmed
to fail against the previous implementation.
Co-authored-by: multica-agent <github@multica.ai>
* fix(skills): bound the frontmatter name value with the YAML parser (MUL-5529)
Review round two found a valid multi-line name the indentation rule still
mangled. A quoted scalar may wrap onto a line at the *same* indentation as its
key:
name: "upstream
continued"
The rule stopped at the key's line, stranding `continued"` and producing
invalid YAML. Probing the parser showed the same holds for flow collections
(`name: [a,\nb]`, `name: {a: 1,\nb: 2}`), so this was not a quoting special
case: indentation simply does not bound a YAML value, and no amount of
patching the heuristic would have made it one.
Value extent now comes from yaml.v3's own line numbers. frontmatterNameValueSpan
locates the `name` key node and ends the span where the next top-level key
begins, stepping back over blank lines and unindented comments so they survive
the rewrite. Unindented is the operative word: block scalar content is always
indented, so an unindented `#` can only be a comment, while ` # text` inside a
block scalar is value and stays in the span.
Detection stays lexical, in lexicalFrontmatterNameSpan. It runs on malformed
blocks too, where there is no parse to consult, and only decides which branch
to take.
setFrontmatterName now re-parses its own output and returns verified=false
unless `name` really is the slug; ensureSkillFrontmatter then routes to the
existing re-synthesis path rather than emitting a block it cannot vouch for.
This adds no new fallback — it feeds an already-present one. The invariant is
the whole point of the change, so an unprovable rewrite is worth less than a
reformatted block: with the span deliberately broken, output stays valid YAML
carrying the right name and loses only upstream formatting.
Tests extend the table to same-indent single/double quoted scalars, flow
sequences and mappings, and name-as-last-key, and now assert the *input* parses
so a case cannot pass by silently taking the re-synthesis path. Two more cover
comment survival and a `#` line inside a block scalar name. All four
same-indent cases fail against the previous implementation.
Co-authored-by: multica-agent <github@multica.ai>
* fix(skills): preserve policy keys when the surgical name rewrite fails (MUL-5529)
Review round three: the post-condition check added last round was routing valid
YAML into bare re-synthesis, and re-synthesis emits only name and description.
An anchor on the name value is one way to get there. Replacing the line drops
`&skill_name`, so `description: *skill_name` no longer resolves, the check
correctly rejects the rewrite — and the block was then rebuilt as just:
name: my-slug
`disable-model-invocation: true` went with it. That key is the author's
instruction that a runtime must not surface the skill on its own, and the
SKILL.md we write is what native discovery reads, so losing it advertises a
skill that was deliberately hidden. Confirmed on the written output:
skillDisablesModelInvocation went from true to false. That is a semantic
regression, not the formatting loss the fallback was justified by.
The two failure modes are now separate:
- invalid YAML → re-synthesize, unchanged; there is nothing to preserve.
- valid YAML, rewrite unprovable → renameFrontmatterNameViaNode rebuilds the
block from the parsed node with `name` set to the slug, keeping every other
key. Formatting normalizes; semantics survive.
The anchor is deliberately kept on the name node. Dropping it is what
invalidates a document that aliases it; keeping it means the alias resolves to
the slug — that value changes, but the document still loads and every policy
key is intact. The rebuilt block is re-parsed and checked like the surgical
path, so an unprovable result still falls through to re-synthesis.
Regression test asserts name, disable-model-invocation, and a custom key all
survive, and that the written file still reads as hidden to
skillDisablesModelInvocation. It fails without the new path.
Co-authored-by: multica-agent <github@multica.ai>
* fix(skills): materialize aliases before renaming an anchored name (MUL-5529)
Round three kept the anchor on the name node so aliases would not dangle. That
avoided an invalid document but created a worse one: every alias resolves
through the anchor, so renaming the anchored value silently rewrote whatever
those fields meant.
name: &shared "true"
disable-model-invocation: *shared
skillDisablesModelInvocation went true -> false across the rewrite. Same
skill-exposing regression as round three, reached by keeping the key instead of
dropping it — which is the lesson: preserving a key is not the invariant,
preserving each key's resolved value is.
Aliases pointing at the name node are now materialized to the value they
resolved to *before* the rename, and the anchor is dropped afterwards as
unreferenced. Ordering matters: the clones are taken first, so they capture the
original value rather than the slug. Copies are per-alias, since sharing one
node would make the encoder re-emit an anchor/alias pair.
Nested aliases are covered by walking the whole document, not just the top
mapping.
Tests: three shapes (alias carrying the policy value, alias nested in another
mapping, two aliases of one anchor) each assert the fixture starts hidden and
stays hidden, that no anchor or alias survives, and that name is the slug. The
round-three test now also pins its aliased `description` to the pre-rename
value instead of merely tolerating the slug. All four fail without the change.
Co-authored-by: multica-agent <github@multica.ai>
* fix(skills): let the parser decide whether a name key exists (MUL-5529)
The branch that chooses between "rewrite the name" and "inject a name" was
still gated on a lexical scan, which only recognizes a bare `name:` carrying a
value on the same line. Two valid spellings therefore read as nameless:
"name": upstream # quoting is syntax, not identity
name: # a key with no value is still the key
Both got a second `name` injected above the existing one, and a duplicate
mapping key is rejected outright — `mapping key "name" already defined`. The
skill does not end up misnamed, it fails to load. Single-quoted keys have the
same problem.
For valid YAML the parsed top-level mapping now answers the question, so any
spelling of the key routes to the rewrite. The lexical scan is confined to the
invalid-YAML branch, where there is no parse to consult and it is only choosing
between re-synthesis and injection. Nesting still reads as absent: a `name`
under another mapping is not the skill's name, so one is added.
Once past the gate the existing paths handle both shapes unchanged, since
frontmatterNameValueSpan already bounds the entry by node line numbers rather
than by how the key is written.
Also tightens the alias tests per review: the nested and two-alias cases now
assert `meta.inner` and `other` still resolve to the anchor's original value,
not merely that visibility survived. The invariant is every key's resolved
value, so every alias should be pinned, not just the one that gates hiding.
Co-authored-by: multica-agent <github@multica.ai>
* fix(skills): detect quoted and valueless name keys in malformed frontmatter (MUL-5529)
Co-authored-by: multica-agent <github@multica.ai>
---------
Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Steve Jobs (Multica Agent) <agent-steve-jobs@multica.ai>
|
||
|
|
e45a8f6c12 |
fix(daemon): scan comment roots before bulk reads in agent catch-up (MUL-5372) (#6093)
* fix(daemon): scan comment roots before bulk reads in agent catch-up The mandatory step-3 catch-up in the issue runtime brief asked for `--recent 10`. `--recent N` caps THREADS, not comments: each returned thread carries its root plus every descendant with no depth bound, so on an issue with fewer than N root threads it returns the entire comment history. Because the step is mandatory and fires on every run, every reply turn re-read the whole issue -- and on comment-triggered turns it duplicated the bounded thread read the per-turn message had already pointed at, so the same bytes were fetched twice. Lead the step with `--roots-only --summary` instead: every top-level thread with reply_count and last_activity_at, contents clipped. That keeps the property the step exists for -- the agent still sees every thread that exists, so it cannot act on stale context -- and makes the drill-down into `--thread <id> --tail 30` explicit. `--recent 10` stays documented for when several complete threads really are needed, now with its saturation semantics spelled out. Measured on a live 2-thread issue: 21,249 -> 1,518 bytes for the mandatory read (-93%), and the duplicate 11,082-byte thread read is gone. The brief stays byte-identical across runs of a session (MUL-5377): the new text interpolates only the issue id, no per-run state. The three per-turn pointers that express the same rule move with it so the two layers cannot drift. MUL-5372 Co-authored-by: multica-agent <github@multica.ai> * refactor(daemon): keep comment-read flag semantics in one place The previous commit fixed the payload shape but restated the read surface in four places: the workflow step, both per-turn prompt fallbacks, and the cold-start hint each explained what `--recent 10` does. `## Available Commands` is already the brief's single discovery point for these flags, and `TestInjectRuntimeConfigStaticCatchUp` pins it as such -- so those restatements were duplicated reference text, and the per-turn ones were paid on every turn rather than once in the cached prefix. Move the `--recent N` saturation warning into the `comment list` line in Available Commands, next to the flags it qualifies, and add `--roots-only` and `--summary` to that signature so the bounding options are discoverable where an agent already looks. Workflow steps and per-turn hints now name only the reads they actually want run. Per-turn prompt sizes: assignment 1170 -> 749 bytes (-36%), cold-start comment turn 1550 -> 1355 (-13%). Step 3 is 1065 bytes and no longer carries a ready-to-paste bulk read. MUL-5372 Co-authored-by: multica-agent <github@multica.ai> * docs(daemon): address review nits on comment-catchup change Three cosmetic follow-ups from review: - `--recent N` saturation warning said it hands back "the entire history"; resolved threads are still folded by default on that read, so say so. - Rename two tests whose names still advertised `--recent` after their assertions stopped mentioning it, plus the one added in this branch whose name referenced a bulk read the step no longer contains: MentionsRecent -> ScansRootsFirst, ScansRootsBeforeBulkRead -> ScansRootsFirst. - Fix the stale doc comment that still described the mandatory read as bounded to "the recent active-thread window". No behavior change. MUL-5372 Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
4eb80c8b77 |
fix(daemon): keep the runtime brief byte-stable across triggers (MUL-5377) (#6021)
* fix(daemon): keep the runtime brief byte-stable across triggers (MUL-5377) Claude Code loads the runtime brief (CLAUDE.md / AGENTS.md) into messages[0], ahead of the entire conversation. A cache breakpoint is all-or-nothing, so a single differing byte there invalidates the prompt cache for the whole history on every `--resume`. InjectRuntimeConfig rewrites that file on every run and interpolated nine per-run values into it, so in practice the cache was thrown away on the first comment that landed on any issue. Measured on one issue over three runs: run 1 (cold) spent 89.9k cache-write tokens building 105k of context; runs 2 and 3 each spent ~425k re-creating a prefix they should have read. 842k of 946.5k cache-write tokens (89%) went into re-creation, with only tools[]+system[] surviving each resume (a constant 18,085 tokens both times). Fix: the brief now carries only what is stable for the lifetime of a resumed session, and per-run state travels in the per-turn user message, which is appended after the cached prefix. - Merge kindCommentTriggered + kindAssignmentTriggered into kindIssue, and stop reading TriggerCommentID in classifyTask. The brief can no longer diverge by trigger type structurally, rather than by convention. - Replace writeWorkflowComment/writeWorkflowAssignment with one writeWorkflowIssue that routes on the per-turn message. The mode-specific status rules live inside their own mode block, so "own the status arc" and "do not touch the status" can never be read as unarbitrated peers. - Move Task Initiator, Session Continuity Notice and Connected Apps out of the brief into BuildPrompt via BuildTaskInitiatorBlock / SessionContinuityNotice / BuildConnectedAppsBlock. - Drop TriggerCommentID, TriggerThreadID, NewCommentsSince, NewCommentCount, PriorSessionResumed and CommentReplyTargets from the brief; BuildPrompt already emitted all six from the same helpers, so this is de-duplication. - Set PriorSessionResumeUnavailable on `task` as well as `taskCtx` in both local resume gates, or the notice would silently vanish on exactly the failure path it exists to disclose. Tests: TestInjectRuntimeConfigByteIdenticalAcrossTriggers renders the brief across nine per-run variants (trigger type, differing comment/thread ids, resume delta, resume-unavailable, cross-thread fan-out, member/agent initiator, connected apps) for two providers and requires bytes.Equal, with a non-vacuity guard so it cannot pass on a function that ignores its input. Daemon-side tests assert the moved sections still reach the agent through the per-turn prompt. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): route the issue workflow on an explicit turn-mode marker Review follow-up on the mode router. The brief said Reply mode applies when the per-turn message "opens with a [NEW COMMENT] block", but buildCommentPrompt writes two paragraphs before that block and only emits it when TriggerCommentContent is non-empty. Two ways to get it wrong: - The message never literally opens with the block, so the router's own wording did not match the prompt it describes. - A comment-triggered run with an empty comment body — or an older server that does not send one — emitted no block at all. An agent following the brief would fall through to Ownership mode and change the issue status on a turn whose rule is "do NOT change the issue status". BuildPrompt now emits an unconditional `**Turn mode: Reply.**` / `**Turn mode: Ownership.**` line from the same branches it uses to pick a code path, and the brief routes on that marker. Brief and prompt can no longer disagree about the mode, because the value that selects the path also states it. The router also names a safe fallback (treat an unlabelled turn as Reply mode and leave the status alone). Tests: TestTurnModeMarkerAlwaysPresent covers comment-triggered with and without comment content, plus both assignment shapes; TestTurnModeMarkerAbsentOnIssuelessKinds keeps the marker off chat / quick-create / autopilot; TestBriefModeRouterMatchesPromptMarkers fails if the brief ever describes a marker the prompt does not emit. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
3c6bebaff0 |
MUL-5274: allow explicit persistent service handoffs (#5895)
* fix(runtime): allow explicit persistent service handoffs Co-authored-by: multica-agent <github@multica.ai> * fix(runtime): resolve review ambiguities in persistent-service handoff wording Address MUL-5274 review findings on #5895: - Drop the "The rules above apply only to work owned by the current run" scoping sentence: with the persistent-service exception inserted above it, it would have swept in work that is precisely no longer run-owned after handoff. The external-systems bullet carries the boundary on its own, and both pin tests now reject any "The rules above" reintroduction. - Replace "detach it" (skill-level mechanism) with the lifecycle contract: hand off only once the service no longer depends on this run. - End the negative-boundary bullet with "the CI-specific rules below still apply" instead of "must be collected before exit", which misread as license to start CI polling and collect it. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
a3fe6d91dd |
MUL-5150: add project context to Chat (#5765)
* feat(chat): add project context Co-authored-by: multica-agent <github@multica.ai> * fix(chat): resolve MUL-5150 review blockers - Renumber project-context migrations to unique prefixes after current main: 206_chat_session_project -> 212 (column), 207_chat_session_project_index -> 213 (concurrent index). 206/207 collided with 206_agent_disabled_runtime_skills and main's 207-211 client_usage_daily set. - Add the 4 missing chat input.project_context keys to ja/ko locales so the locale parity test passes (en/zh-Hans already had them). - Lock the project-context control while a send is in flight (isSubmitting), not just while the agent is running. A brand-new chat creates its session lazily during send bound to the project at click time; switching project mid-send would create the session against the stale project and clear the editor as if the send landed on the new selection. Add a regression test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: multica-agent <github@multica.ai> * fix(chat): complete project context handling * fix(chat): pin fresh chat to open session's agent on project switch Switching an existing session to a different project opens a fresh chat but only cleared the active session, dropping selection back to the stored `selectedAgentId`. When that preference was stale (open session belongs to agent B while the persisted pick is still agent A), the lazily-created session and its first send bound to the wrong agent (agent A). Extract the project-switch decision into a shared `planProjectContextChange` pure helper in use-chat-controller.ts and route both chat surfaces (the chat tab controller and the floating ChatWindow) through it, so the fresh chat is pinned to the open session's agent and the rule cannot drift between the two copies. Add a dual-entry regression test (pure-fn guard + controller integration) covering the stale selectedAgentId case. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: multica-agent <github@multica.ai> * chore(ci): re-trigger required checks on latest head The prior push updated the branch ref but GitHub did not emit a pull_request synchronize for it (PR head-sync lag), so CI/Mobile Verify never ran on the commit carrying the stale-agent project-switch fix. Empty commit to force a fresh synchronize on a head that includes it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: multica-agent <github@multica.ai> * fix(chat): renumber project migrations to 213/214 after main added 212 Current main added 212_agent_service_tier; the PR's 212/213 chat migrations collided with it on the merge ref, failing TestMigrationNumericPrefixesStay UniqueAfterLegacySet. Merge current main and move the chat column migration to 213 and the concurrent index migration to 214 (column before index preserved). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: multica-agent <github@multica.ai> * fix(chat): lock ProjectPicker clear control during send (keyboard path) The send-pending lock only put pointer-events-none on the wrapper, which blocks the mouse but leaves ProjectPicker's inline clear button in the tab order — a keyboard user could Tab to "Remove from project" and press Enter mid-send, detaching the project after the lazily-created session already went out with the old one (reopens the mid-send retarget path via keyboard). Add an explicit `disabled` capability to the shared ProjectPicker that locks the trigger, the menu (forced closed), and the inline clear button (disabled + out of the tab order). Defaults to false, so issue/create/autopilot callers keep their hover/keyboard clear. ChatInput passes disabled while the project selection is locked. Tests: real-ProjectPicker regression (keyboard activation of the clear control is inert when disabled; still works when enabled) + ChatInput wiring assertion that the picker is disabled mid-send. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Lambda <lambda@multica.ai> Co-authored-by: multica-agent <github@multica.ai> Co-authored-by: Walt <walt@multica.ai> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Naiyuan Qing <145280634+NevilleQingNY@users.noreply.github.com> Co-authored-by: NevilleQingNY <nevilleqing@gmail.com> |
||
|
|
e2ce3a2da8 |
MUL-5223 fix(runtime): forbid blocking on external CI in the runtime brief (#5840)
* fix(runtime): forbid blocking on external CI in the brief (MUL-5223) The external-work boundary added in #5803 did not stop agents from waiting on GitHub Actions. Two holes: the section's only concrete "how to wait" example was a blocking foreground call, which is exactly the shape of `gh pr checks --watch`; and the "unless acceptance criteria require it" escape was satisfied by the repo's own merge requirement that CI be green. Name the banned tool shapes, allow a single non-blocking status snapshot, deny branch protection as an acceptance criterion, and give the replacement hand-off phrasing (local test result + PR link). Co-authored-by: multica-agent <github@multica.ai> * fix(runtime): scope the CI-wait ban so the explicit exception stays executable (MUL-5223) Review feedback on #5840: - The ban read as absolute ("Blocking on external CI is never part of your deliverable") while the next bullet allowed waiting when the task explicitly asks for the CI result, leaving no way to satisfy both. The ban is now scoped to "unless the explicit exception below applies", and the exception names the one executable shape: a single foreground blocking watch inside the same turn. - `gh pr merge --auto` enables auto-merge and returns; it is not a wait. Only waiting for it to land is banned. Both hard-pin tests now also pin the exception so it cannot be dropped or re-absolutised. Co-authored-by: multica-agent <github@multica.ai> * polish(runtime): group Background Task Safety into run-owned and external-CI clusters (MUL-5223) Co-authored-by: multica-agent <github@multica.ai> * polish(runtime): cut redundant phrasing from the external-CI cluster (MUL-5223) The cluster said "report and finish" three different ways and carried two rhetorical tails. Fold the delivery template into the post-push playbook bullet, tighten the merge-gate denial, and drop filler. 5 bullets -> 4, -36 words, every behavioral fact and test pin intact. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
2c749ddf9f | fix(runtime): clarify external background work (#5803) | ||
|
|
fbe00ca164 |
fix(daemon): run Codex unsandboxed on Windows to stop reject-by-policy (MUL-4957) (#5672)
* fix(daemon): run Codex unsandboxed on Windows to stop reject-by-policy (MUL-4957) Windows has no Landlock/Seatbelt-equivalent filesystem sandbox that the daemon configures, so the per-task `sandbox_mode = "workspace-write"` it wrote was unenforceable. Worse than having no sandbox, it pushed Codex into rejecting non-safe mutation commands "by policy": `multica issue create` fails with "was rejected by policy" because Codex can neither sandbox the command nor (under approval_policy = "never") escalate it to the daemon's auto-approver, so the request never reaches the approver. Mirror the existing macOS fallback and give Windows danger-full-access so those commands run. Also generalize the danger-full-access warn log so it no longer hardcodes "on macOS" and only surfaces the macOS-specific upgrade hint on macOS (new codexSandboxPolicy.Hint field). Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): correct Windows sandbox rationale and respect user windows.sandbox (MUL-4957) Addresses two review must-fixes on #5672: 1. Correct a false security fact. The comments and log Reason claimed Windows has no filesystem sandbox backend. Codex 0.144.5 does ship a native Windows sandbox (windows.sandbox = "unelevated"/"elevated"); it is experimental with open upstream reliability bugs, so the daemon defaults to danger-full-access as a deliberate compatibility choice. Enabling the native sandbox is tracked as separate follow-up work. 2. Stop silently downgrading users who opted into isolation. The fallback was unconditional. Add codexSandboxPolicyForConfig: on Windows an explicit windows.sandbox = unelevated|elevated keeps workspace-write so Codex enforces task isolation with the user's chosen backend; danger-full-access applies only when windows.sandbox is absent, disabled, or unparseable. This is also the branch point for a future native-sandbox rollout (flip the default; callers unchanged). Adds fixture tests locking the priority (user opt-in kept vs. unconfigured fallback) plus predicate coverage for codexSandboxPolicyForConfig. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): fail closed on undecidable Windows sandbox config, honor -c windows.sandbox (MUL-4957) Second review round on #5672. Two must-fixes. 1. Undecidable config no longer fails open. The old bool detector collapsed "unparseable / invalid value / failed copy" into "unconfigured" and then loosened to danger-full-access. Replaced with a tri-state (absent/native/undecidable): only exact-lowercase unelevated|elevated (the sole values Codex accepts — verified: any other value makes Codex refuse to load the config) counts as native; any other present value, unparseable TOML, a read error, or a missing per-task config when a shared ~/.codex/config.toml exists (i.e. the copy failed) is undecidable and fails closed to workspace-write — it never loosens — logged at error level. 2. windows.sandbox set via `-c`/`--config` custom args is now honored. Such args never land in config.toml, so config-only detection silently downgraded those users' isolation. The effective Codex args (daemon defaults + profile-fixed + per-agent custom_args) are threaded through PrepareParams/ReuseParams/CodexHomeOptions into the sandbox decision and scanned for a windows.sandbox override (inline, two-token, quoted, spaced; last-wins). Also drops issue-status-bound source comments (openai/codex#24098 has since closed). Adds unit coverage for config/args classification, the fold precedence (undecidable > native > absent), and the copy-failed fail-closed path. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): fail closed on config-sync errors and honor shell-quoted -c windows.sandbox (MUL-4957) Round-3 review must-fixes: 1. resolveWindowsSandboxState now takes the config.toml sync error and a tri-state shared-config presence instead of re-stat-ing inside. A failed sync (stale/absent per-task copy) or an un-stat-able shared source is undecidable and keeps workspace-write, closing the fail-open where a failed sync was read as "unconfigured". Splits IO from the decision so the paths are unit-testable without faulting the filesystem. 2. The Windows sandbox decision consumes agent.NormalizeCodexLaunchArgs (the shared helper buildCodexArgs now uses) so a shell-quoted -c windows.sandbox opt-in is normalized identically to launch, instead of being missed by a raw-token scan and silently downgraded. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): abort when the Codex sandbox block cannot be written (MUL-4957) Round-4 review must-fix: ensureCodexSandboxConfig failures were warn-and-continue, so a computed fail-closed workspace-write policy could stay only in memory while config.toml kept a stale danger-full-access from a prior run — the decision failed closed but the effective config failed open. prepareCodexHomeWithOpts now returns the error, which blocks startup on both paths: fresh Prepare fails the task, and Reuse leaves env.CodexHome unset, which configureCodexTaskShellEnvironment already refuses to start. Regression covers the full reuse scenario (stale danger-full-access + failed config sync + failed managed-block write); it fails with "got nil" without the fix. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> Co-authored-by: J <j@multica.ai> |
||
|
|
f8bf6cd8b9 |
feat(runtime): add Qwen Code runtime (MUL-5015)
Merge approved PR #5666. |
||
|
|
5664c0d755 |
MUL-4874: fix(execenv): seed Codex model cache in task homes (#5449)
* fix(execenv): reuse Codex startup caches * fix(execenv): preserve task-local Codex model cache * fix(codex): bind model cache to provider config Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
54fd29ebdd |
MUL-4383 fix(daemon): stop routing CodeBuddy skills/memory through Claude's .claude paths (#5224)
* fix(daemon): stop routing CodeBuddy skills/memory through Claude's .claude paths CodeBuddy Code is a Claude Code fork but ships its own native config directory (~/.codebuddy, .codebuddy/) with its own memory filename (CODEBUDDY.md). It only reads .claude/skills or CLAUDE.md if a user manually symlinks/copies them during migration (https://www.codebuddy.ai/docs/cli/troubleshooting#migrating-from-claude-code). Multica's daemon/execenv code treated "codebuddy" as an alias for "claude" in three places, so skills synced by Multica landed in .claude/skills/ and CLAUDE.md — paths the default CodeBuddy install never reads — instead of ~/.codebuddy/skills, .codebuddy/skills, and CODEBUDDY.md as documented at https://www.codebuddy.ai/docs/cli/codebuddy-dir and https://www.codebuddy.ai/docs/cli/skills. Split the "claude", "codebuddy" switch cases in: - daemon/local_skills.go (user-level local skill discovery/import) - daemon/execenv/context.go (per-task skill materialization) - daemon/execenv/runtime_config.go (runtime brief target file) Added regression tests locking in the new paths and updated the install-agent-runtime / providers docs (all 4 locales) that had documented the old .claude/skills behavior. Co-authored-by: Cursor <cursoragent@cursor.com> * test(daemon): cover CodeBuddy sidecar hygiene and lifecycle Address PR #5224 review feedback: exclude CODEBUDDY.md/.codebuddy from repo-cache worktrees, extend sidecar lifecycle matrices to codebuddy, and make the local-skills CodeBuddy test exercise a true same-key collision. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Eve <eve@multica-ai.local> |
||
|
|
6cc553e5a3 |
fix(daemon): isolate Codex sessions per task to unblock initialize (MUL-4424) (#5360)
* fix(daemon): isolate Codex sessions per task to unblock initialize (MUL-4424) Codex 0.143+ backfills a per-home session-state DB by enumerating every rollout visible under sessions/ during `initialize`. The per-task CODEX_HOME symlinked the shared ~/.codex/sessions in, so a machine with a large accumulated history (one reporter: ~2000 rollouts / ~22 GiB) stalled `initialize` for tens of seconds — the app-server started but the task produced no output before it was cancelled (github #5273). Give each task its own local sessions/ directory instead: - Fresh task: create an empty local sessions/ so backfill is trivial. - Reused task with a real sessions/ dir: it is authoritative — leave it. - Reused task still holding a legacy symlink (older build): migrate in place. Replace the symlink with a real dir; when resuming, symlink only the single rollout being resumed (never copy — a rollout can be GiB and this is on initialize's critical path); and drop the stale, rebuildable session-state DB (state_*.sqlite*, session_index.jsonl) so Codex re-indexes the task-local sessions. Unrelated per-task DBs (goals_*, logs_*, memories_*) are left intact. Also point the token-usage fallback scan at the backend's per-task CODEX_HOME instead of the daemon-global home, so usage isn't lost now that sessions are isolated there. Complements the #5319 handshake watchdog (which turns a silent stall into a loud, phased timeout); this removes the underlying cause. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): address Codex session isolation review (MUL-4424) Resolves the three blockers from Elon's review of #5360: 1. local_directory context loss. local_directory tasks get a fresh codex-home per task ID (the daemon never reuses their workdir), so task-local isolation stranded every follow-up run with an empty sessions dir and silently restarted the conversation. Their only stable, GC-safe cross-task store is the user's own ~/.codex/sessions (a persistent store under WorkspacesRoot would be orphan-GC'd), so keep the shared-sessions symlink for them (IsLocalDirectory). Managed tasks stay isolated. 2. Migration resume robustness. - Rollout lookup now covers the flat layout and background-compressed .jsonl.zst rollouts, not just nested YYYY/MM/DD *.jsonl — both are legitimate Codex 0.144 history that were previously judged "not found", silently dropping resume. - Exposure hard-links first, then symlinks, never copies — hard links need no privilege and work on Windows within a volume, so the zero-copy path is exercised identically on CI. - The daemon now verifies the rollout is actually present in the task CODEX_HOME (execenv.CodexResumeRolloutPresent) before the brief is generated; if absent it clears the resume from both the backend and the brief instead of telling the agent it is continuing a lost thread. 3. session_index.jsonl is no longer deleted during migration — Codex uses it as the authoritative thread-id -> name store (not rebuildable from rollouts). Only the rebuildable state_*.sqlite* is reset. Tests: 2-round local_directory resume across task IDs; compressed/flat lookup; hard-link zero-copy (os.SameFile); session_index preserved; CodexResumeRolloutPresent + the daemon gate helper (present keeps / absent drops / non-codex + empty no-op). Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): scope Codex sessions to a per-issue store; disclose lost resumes (MUL-4424) Addresses the three blockers from Elon's second review of #5360. 1. local_directory still enumerated the whole machine history. The prior fix re-linked the entire ~/.codex/sessions into every fresh local_directory codex-home, so Codex still backfilled from thousands of unrelated rollouts on `initialize` (measured ~8.3s with 3450 rollouts; the reporter's 22 GiB could exceed the 30s watchdog). Point sessions/ at a persistent, per-(agent, issue) store under the shared Codex home (multica-sessions/<agent>/<issue>) that holds only this issue's rollouts. It is keyed stably across task IDs and lives outside the task-scoped envRoot the GC reclaims, so follow-up runs resume it while `initialize` only ever sees this issue's history. 2. Windows cross-volume resume was lost. Exposing a single rollout by hard-link (same-volume only) then file symlink (needs Windows privilege) can't cross a volume boundary. The store now lives on the shared Codex volume, so the resume rollout is hard-linked there zero-copy, and sessions/ is exposed to the task home via a directory link — a symlink on Unix, a junction on Windows — which crosses volumes without privilege and never copies a (possibly GiB) rollout on initialize's critical path. There is no remaining per-file cross-volume link. 3. An unavailable resume was a silent downgrade. Both resume gates (gateResumeToReusedWorkdir, gateCodexResumeToRolloutPresence) now set PriorSessionResumeUnavailable, and the runtime brief renders a Session Continuity Notice telling the agent to disclose to the user, up front in its reply, that the previous conversation context could not be restored and this run starts fresh — turning a silent restart into a user-visible one. The task is not failed: it can still do useful work without the prior context. Managed fresh / reused-real-dir tasks keep their task-local, GC-collected sessions dir unchanged; only the legacy-symlink migration with a resume routes through the store (cross-volume-safe), and a home already linked to the store is treated as authoritative on reuse. Tests: local_directory per-issue store (only this issue's history, no whole- machine leak); no-key fallback to an empty dir; two-round resume across task IDs through the store; legacy migration routed through the store with a zero-copy hard link; reused store link stays authoritative; both gates set the resume-unavailable flag; brief renders the continuity notice only when a resume was lost. execenv + daemon + pkg/agent packages, go vet, and gofmt all pass. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): disclose live resume-RPC loss; bound Codex session store lifecycle (MUL-4424) Addresses the two blockers from Elon's third review of #5360. 1. A real thread/resume failure was still a silent new session. The brief's Session Continuity Notice only covers losses the daemon detects before launch (workdir not reused, rollout absent). But when the rollout is present yet Codex rejects the live thread/resume (corrupt/incompatible rollout, server-side thread GC, schema drift), startOrResumeThread falls back to thread/start and the run succeeds on a fresh thread with no user-facing signal. Carry the original resume intent into the backend as ExecOptions.ResumeExpected (set from the post-gate PriorSessionID, so a pre-flight drop still routes through the brief and never double-notifies), and when a resume was expected but the backend landed on a fresh thread, prepend the same continuity notice to the first turn/start input. This also covers the daemon's transport-error fresh-session retry, which clears ResumeSessionID but not ResumeExpected. 2. The persistent per-issue store had no data lifecycle. multica-sessions stores live outside the task-scoped envRoot the GC reclaims (so resume survives across task IDs), which meant a done/abandoned issue's prompts and full rollouts (one reporter: a single 1.5 GiB rollout) accumulated forever and were never freed on issue/agent/workspace deletion. Add PruneCodexSessionStores: the daemon GC loop reclaims any store untouched for GCCodexSessionTTL (default 14 days, configurable via MULTICA_GC_CODEX_SESSION_TTL, 0 disables). A store's newest rollout mtime is its last activity, so an active or recently-resumed task keeps its store fresh and is never reclaimed, while a deleted issue's store ages out — an eventual reclamation guarantee without needing deletion events. Tests: codexTurnInput discloses on resume fallback and stays silent on success / fresh start (paired with the existing live-RPC fallback test); store pruning reclaims aged stores, keeps recent ones, isolates issues, cleans empty agent dirs, and is disable-able. execenv / daemon / pkg/agent, go vet, gofmt all pass. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): protect a reopened Codex session store from GC mid-mount (MUL-4424) Addresses Elon's fourth-review blocker: reopening an issue idle past GCCodexSessionTTL could lose context, because mounting its per-issue session store (MkdirAll + rollout lookup + task-home link) never refreshed the store's mtime, so a GC cycle firing before the resumed turn wrote its first rollout saw a >TTL-old store and reclaimed it — a stat->remove race with no in-use guard. Two complementary defenses: - Activity refresh: linkCodexSessionsToStore now os.Chtimes the store to now after linking, so codexStoreStat (which reads the newest mtime as last activity) sees a just-used store. This fixes the sequential repro — a mount immediately followed by a prune keeps the store. - In-process active-store guard: the daemon marks the per-issue store in-use (execenv.CodexSessionStorePath) from before Prepare/Reuse mounts it until the task ends, and PruneCodexSessionStores now takes an isActive predicate and skips any store a live task holds. Because prepare and prune run in the same process, this closes the remaining concurrent stat->remove window the mtime refresh alone cannot. Reference-counted, mirroring the env-root guard. Tests: a reopened >TTL store survives a GC cycle after remount and stays resumable; an idle-on-disk store marked active is skipped, then reclaimed once inactive; the existing idle-reclaim / isolation / disable / empty-agent-dir cases still pass. execenv + daemon, go vet, gofmt all pass. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): make Codex store delete atomic with mark-active (MUL-4424) Addresses Elon's fifth-review blocker: the active-store guard's check and delete were not one atomic step. PruneCodexSessionStores called isActive (which locked, read, and unlocked) and only then RemoveAll'd, leaving a window where a task could markActiveCodexStore between the check and the removal and still lose its store — the exact mark-then-delete interleaving Elon reproduced. Replace the point-in-time isActive predicate with a reserve-for-deletion protocol that shares one lock with mark-active: - reserveCodexStoreForDeletion(store) atomically refuses when a live task holds the store (or another delete already reserved it) and otherwise marks it reserved, all under one activeCodexStoresMu acquisition. PruneCodexSessionStores reserves before RemoveAll and commits after, so confirm-inactive and remove are effectively atomic against a concurrent mark. - markActiveCodexStore now waits (on a sync.Cond) while a store is reserved, so a task never mounts a store mid-removal; committing the removal wakes it and the store is recreated fresh by Prepare (with the continuity notice). So mark-before-reserve keeps the store (reserve refused); reserve-before-mark removes it and blocks the late mark until the removal commits. The genuinely idle case still reclaims. Tests (daemon, run under -race): mark-then-reserve is refused; reserve blocks a concurrent mark until commit then the store reads active; a second reserve is refused mid-flight. The execenv prune tests move to the reserve seam; the activity-refresh / reopen-then-prune / isolation / disable cases still pass. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): namespace Codex session stores per profile for cross-daemon safety (MUL-4424) Addresses Elon's sixth-review blocker: the in-process reservation guard cannot span processes, but Multica supports multiple profile daemons on one machine (e.g. production + staging) that share the same ~/.codex. Each daemon's GC scanned the whole multica-sessions root, so a staging daemon could reclaim a store a production task was actively resuming — its reservation lived only in the other process's memory. Isolate by profile instead of trying to lock across processes: - Store path is now <shared>/multica-sessions/<namespace>/<agent>/<issue>, where namespace is the daemon's profile (empty -> "default"). PrepareParams/ReuseParams carry Profile; codexSessionStoreKey and CodexSessionStorePath fold it in. - PruneCodexSessionStores takes the profile and scans ONLY that namespace, so a daemon never even sees another profile's stores, let alone deletes them. The per-profile trees are disjoint, so the in-process guard is sufficient within a namespace (profiles get separate daemon state, so no two daemons share one). Test: a "staging"-owned idle store is untouched by a default-profile prune and reclaimed only by staging's own prune. Existing prune/guard/reopen tests move under the namespace. execenv + daemon under -race, go vet, gofmt all pass. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): make the Codex store profile→namespace map injective (MUL-4424) Addresses Elon's seventh-review blocker: the per-profile namespace was derived by dropping unsafe characters, which is not injective. The CLI treats the empty (default) profile and a profile literally named "default" as separate daemons, yet both mapped to namespace "default"; likewise "staging.prod" and "stagingprod" both mapped to "stagingprod". Two distinct daemons then shared one store tree, so one could again reclaim the other's live session — the cross-process blocker reopened for those profile names. Make codexSessionStoreNamespace injective: the empty profile gets a reserved bare literal "default", and every named profile is hex-encoded (bijective, filesystem-safe) under a "p_" prefix a bare literal can never collide with. So "" -> "default" while "default" -> "p_64656661756c74", and "staging.prod" / "stagingprod" get distinct hex segments. sanitizeCodexPathSegment stays for the UUID agent/issue segments (injective for real UUIDs); only the user-controlled profile needed the encoding. Tests: codexSessionStoreNamespace is distinct for "" vs "default", punctuation variants, case variants, and an encoded-looking name; and end-to-end, pruning one profile never reclaims the other's store for the "" vs "default" and "staging.prod" vs "stagingprod" pairs. execenv + daemon under -race, go vet, gofmt all pass. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): fixed-length Codex store namespace so long profiles fit (MUL-4424) Addresses Elon's eighth-review blocker: hex-encoding the full profile doubled the namespace segment length. A profile can be as long as a filesystem segment allows (~255 bytes) and the CLI persists it as its own config dir, but the store namespace "p_" + hex(profile) reached 2 + 127*2 = 256 bytes at 127 chars, overflowing the 255-byte single-segment limit — the profile's own dir created fine, then the session store failed with "file name too long". Derive the namespace from a fixed-length hash instead: a named profile is now "p_" + hex(sha256(profile)) — a constant 64 hex chars (66 with the prefix), filesystem-safe and collision-resistant. The empty (default) profile keeps its reserved bare literal "default", which the "p_"-prefixed 66-char segment can never equal. Still injective across the CLI's distinct-daemon cases; just no longer length-expanding. Test: the namespace stays <=255 bytes and creatable for profiles up to the 255-byte segment limit (127- and 255-char cases that overflowed under hex); the prior injectivity and cross-profile prune-isolation tests still hold. execenv + daemon under -race, go vet, gofmt all pass. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: J <j@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
ef9b334408 |
MUL-4398: fix Hermes bound-skill discovery with per-task overlay (#5308)
* fix(execenv): overlay per-task HERMES_HOME so Hermes discovers bound skills Hermes has no workspace-relative skill discovery — it scans <HERMES_HOME>/skills first, then skills.external_dirs from config.yaml (verified against the bundled agent/skill_utils.py). The daemon wrote assigned skills to the generic .agent_context/skills/ fallback, which Hermes never reads, so they silently never took effect (#5242). When (and only when) an agent has skills bound, redirect HERMES_HOME to a minimal per-task compatibility overlay; a skill-less Hermes task keeps its real home and original behavior: - mirror every top-level entry of the shared home via symlink except the overlay-owned ones (denylist), reconciling entries deleted from the shared home; - derive a task-local config.yaml whose skills.external_dirs references the shared skills dir plus the user's existing external_dirs, expanded against the sanitized effective child env (unknown vars preserved, blocklisted keys resolved to the process value) and normalized to absolute paths; - write only the bound skills into the task-local skills/ dir (home skills scanned first, so they win); global skills are referenced, not copied; - keep memories/ overlay-owned (fresh per-task dir) AND disable the external memory.provider, so neither on-disk memory nor a shared backend crosses tasks; - keep active_profile/profiles out of the overlay so Hermes can't follow a sticky profile and redirect past it at startup. Profile handling mirrors hermes_cli.profiles: the daemon reads -p/--profile with agent.HermesProfileFromArgs and seeds the overlay from that profile's home via ResolveHermesSourceHome (default/invalid -> base, valid name -> <base>/profiles/ <name>, validated; a missing named profile fails closed). The profile flags are stripped from the acp argv ONLY when the overlay is active (hermesLaunchArgs), so a skill-less task's profile passes through unchanged. HERMES_HOME is no longer custom_env-blocklisted: no skills -> user value passes through; skills -> overlay overrides after layering. Fail closed — Prepare errors, Reuse returns nil. Task home 0700, derived config 0600 via atomic replace. Platform-native default home (%LOCALAPPDATA%\hermes, incl. the LOCALAPPDATA-missing fallback, on Windows). Tests span execenv/daemon/agent: no-skill no-op, child-env layering + env sanitization, profile parse/unquote + conditional strip + final args/env per scenario, custom/profile/default/invalid/missing/Windows source home, sticky- profile not mirrored, memory dir isolation + external provider disable, mirror reconciliation, external_dirs rebasing + sanitized/unknown-var expansion, local-precedence slug, perms, fail-closed, resume teardown. Docs (en + ja/ko/zh). Fixes #5242 * fix(hermes): make profile selection one resolver contract matching Hermes Round 5 review: the profile chain approximated Hermes' semantics in three separate places (argv parsing, source-home selection, arg filtering), so it diverged from native Hermes in several merge-blocking cases. Collapse it into one authoritative resolution: - agent.ParseHermesProfileArgs replaces HermesProfileFromArgs/ FilterHermesProfileArgs. It reproduces _apply_profile_override step 1/1b (first occurrence, value-flag skipping, `--` and `mcp add --args` boundaries, space-form profile-id guard) and returns the exact argv occurrence to consume; StripHermesProfileArgs removes only that occurrence. - execenv.ResolveHermesProfile replaces ResolveHermesSourceHome. It derives the Hermes root exactly like get_default_hermes_root (an already-profile-scoped HERMES_HOME roots at its grandparent), selects an explicit profile first, otherwise trusts a profile-scoped home (step 1.5) and only then the sticky <root>/active_profile (step 2), and validates via normalize/validate_profile_name (reserved hermes/test/tmp/root/sudo and empty inline `--profile=` are hard errors). Profiles always resolve under the root, so `-p default` re-roots and `-p <sibling>` is a sibling, never nested. - The daemon runs one parse + resolve, fails the task closed on a reserved/ invalid selection (matching Hermes' sys.exit(1)), and exports the selected source home as the effective env's HERMES_HOME so ${HERMES_HOME} in a profile's skills.external_dirs expands against the selected profile home (as native Hermes does before loading config.yaml), not the root or the overlay. Regressions added: root + sticky named profile selection; already-profile-scoped home with no flag; that home with -p default and -p <sibling>; reserved and empty inline profile values; and a selected profile whose external_dirs contains ${HERMES_HOME}. * fix(hermes): overlay-owned derived .env + symlink-resolved root Round 6 review, two remaining overlay-bypass paths: 1. A source `.env` could redirect HERMES_HOME after profile resolution. Hermes runs `_apply_profile_override()` then `load_hermes_dotenv()`, which loads `<HERMES_HOME>/.env` with override=True — so a mirrored source `.env` carrying an out-of-band `HERMES_HOME=` overwrote the overlay's home, repointing skill discovery and memory back at the source. `.env` is now overlay-owned and DERIVED (writeDerivedHermesEnv): it preserves the source's credentials/settings but strips any `HERMES_HOME` assignment and pins `HERMES_HOME` to the overlay last (single-quoted, literal), written 0600 via atomic replace. It is written even when the source has none, so Hermes' project-`.env` fallback (override=True only when no user `.env` loaded) can't relocate the home either. 2. Root derivation was lexical-only, diverging from `get_default_hermes_root`, which compares `env_path.resolve()` with `native_home.resolve()`. A HERMES_HOME symlinked into `<native>/profiles/<x>` was treated as its own root, so `-p default`/`-p <sibling>` resolved wrong. `hermesRootFromHomeFor` now resolves symlinks (Path.resolve(strict=False)-style best effort) for the containment decision while keeping the returned root unresolved, matching Hermes. Regressions: source `.env` with HERMES_HOME replayed through the override=True dotenv order (bound skill + task memory stay on the overlay; creds preserved); minimal overlay `.env` created when the source has none; and a symlinked profile home resolving `-p default`/`-p <sibling>` to the native root. |
||
|
|
0c2e48ded2 |
refactor: retire FF_RUNTIME_BRIEF_SLIM, make slim runtime brief the only path (MUL-4297)
The runtime_brief_slim feature flag has burned in; the slim runtime brief is now the sole path. - execenv: buildMetaSkillContent / BuildCommentReplyInstructions delegate to the slim assembler unconditionally; delete the legacy verbose brief body and writeBackgroundTaskSafetyInstructions. - Remove the runtime_brief_slim flag and the daemon-bound flag delivery subsystem built solely for it: execenv flag wiring (runtime_config_flag.go, server_snapshot_provider.go), the featureflagdispatch package, the DaemonFeatureFlagSnapshot heartbeat protocol field, and the server/daemon wiring in router.go, handler, daemon.go, main.go, cmd_daemon.go. - Keep the generic server/pkg/featureflag engine (still used by composio_mcp_apps). - Update tests to slim-only expectations and docs/feature-flags.md. Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
4c510dfef6 |
fix(daemon): harden background-task-safety brief against background-and-yield (MUL-4140) (#4998)
A Multica-managed run goes terminal the moment the top-level turn exits; there is no "background work finishes later and wakes you up" step. When an agent starts background work (a run_in_background shell, a Monitor, an async subagent) and ends its turn to "wait for a completion notification", the work is orphaned and the result comment it meant to post is never sent (MUL-4091 / PR #4970). The existing claude-only protocol guard forces run_in_background tool inputs to foreground and fails loud on async_launched tool results, but it cannot catch the actual MUL-4091 mechanism: a turn that ends cleanly with a "Standing by, I'll report when CI finishes" message. That shape is only addressable behaviorally, and it is harness-agnostic. Harden the Background Task Safety brief (both the legacy/verbose production path and the slim staging path) with explicit hard pins: - never background-and-yield / expect a future wakeup that does not exist here; - do every wait synchronously in a single foreground call (e.g. gh run watch); - the standalone-harness "running in the background, keep working" hint does not apply in Multica-managed runs; - never end a turn with a "standing by" / "I'll report back" sign-off. Add verbose- and slim-path test coverage for the new pins so a future brief trim cannot silently drop them. Co-authored-by: J <j@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
65269ef922 |
fix(daemon): copy Codex model catalog into task home
Fixes #4825 Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
0c4c3ff038 |
fix(cli): prevent daemon-managed CLI from silently using user tokens (MUL-3922)
Treat MULTICA_DAEMON_PORT and a workdir daemon-task marker as daemon-managed signals so a task subprocess that loses MULTICA_TOKEN / MULTICA_AGENT_ID / MULTICA_TASK_ID fails closed instead of silently falling back to the user config-file PAT (which made agent writes land as the workspace owner). Adds an actionable error naming a leftover marker for local_directory recovery. Fixes #4204. |
||
|
|
20eecfb093 | fix(projects): honor repo resource checkout refs (MUL-3593) (#4470) | ||
|
|
76c58a4ee8 |
MUL-3617: remove Gemini CLI runtime (#4503)
* fix: remove gemini cli runtime Co-authored-by: multica-agent <github@multica.ai> * fix: skip unsupported custom runtime profiles Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: J <j@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
da72e2fa22 |
feat(daemon): inject project description into the agent brief (MUL-3465) (#4395)
* feat(daemon): inject project description into the agent brief Issues bound to a project only surfaced the project title in the runtime brief; the project description (durable, project-wide context the owner sets) was loaded but dropped. Carry it end-to-end: - claim handler reads proj.Description onto the response (issue-bound and quick-create paths) - new ProjectDescription field on AgentTaskResponse, daemon Task, and TaskContextForEnv - rendered in the brief's `## Project Context` section and written to .multica/project/resources.json as project_description Empty descriptions render nothing (no extra heading). Updated the projects-and-resources built-in skill docs in the same change. MUL-3465 Co-authored-by: multica-agent <github@multica.ai> * feat(projects): clarify project description is injected as agent context The project description is now durable context injected into every task's brief, but the UI still presented it as a plain "Description" field, so existing descriptions could silently become agent input. Add a hint under the description editor on the project detail page and in the create-project modal, in all four locales, stating it is shared with agents as context for every task in the project. No data-semantics change. Addresses review feedback on PR #4395. MUL-3465 Co-authored-by: multica-agent <github@multica.ai> * test(handler): assert project description flows through task claim The execenv tests cover brief rendering, but nothing pinned the claim handler boundary where proj.Description is read onto the response. Add two tests — issue-bound and quick-create paths — so a regression in that assignment fails loudly instead of silently dropping the description. Addresses review feedback on PR #4395. MUL-3465 Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: J <j@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
78342a39ce |
MUL-3305: feat(agent): add qoder CLI as a choice of agent provider. (#2461)
* feat(agent): Qoder ACP runtime, chat reconnect recovery, and task linkage - Add Qoder CLI backend (ACP transport, model discovery, blocked-args policy) - Wire daemon/runtime config, docs, and UI provider assets - Retry terminal task reports; add backoff unit tests - Chat: SQL attach user message to task; handler + optimistic cache reconcile - Invalidate chat/task-messages caches on WS reconnect; extract helper + tests Co-authored-by: Orca <help@stably.ai> Co-authored-by: Cursor <cursoragent@cursor.com> * chore: drop non-Qoder changes (chat reconnect, task link, terminal report retries) Keep only Qoder runtime, docs, daemon config/execenv, and UI provider assets. Co-authored-by: Orca <help@stably.ai> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(agent): harden Qoder ACP drain and wire project skills path - Stop streaming to msgCh after reader wait so grace timeout cannot race close - Resolve injected skills to .qoder/skills per Qoder CLI discovery - Update AGENTS.md skill copy and add execenv tests Co-authored-by: Orca <help@stably.ai> Co-authored-by: Cursor <cursoragent@cursor.com> * feat(qoder): add provider logo and wire MCP config into ACP sessions - Add inline SVG QoderLogo component to provider-logo.tsx, replacing the generic Monitor icon placeholder - Add convertMcpConfigForACP helper to convert Claude-style MCP server config (object map) into ACP array format for session/new and session/resume - Add unit tests for convertMcpConfigForACP covering stdio, SSE, empty/nil, and multi-server cases Co-authored-by: Orca <help@stably.ai> * fix(test): capture both return values from InjectRuntimeConfig in Qoder test Co-authored-by: Orca <help@stably.ai> * fix(qoder): preserve remote MCP headers and promote provider errors Addresses review feedback on #2461 (Bohan-J): two runtime-correctness issues in the Qoder ACP backend. 1. Remote MCP headers were dropped. The bespoke convertMcpConfigForACP only forwarded url/type, so an authenticated remote MCP server looked configured in Multica but failed inside the Qoder session. Replace it with the shared buildACPMcpServers helper (same path Hermes/Kimi/Kiro use), which preserves headers as [{name, value}], sorts for deterministic output, and handles remote transport aliases. Fail closed on malformed mcp_config instead of silently dropping servers. 2. Provider failures could report as completed tasks. stderr was wired via io.MultiWriter and the result was only promoted to failed when output was empty, so a terminal upstream error (HTTP 429 / expired token) racing a stopReason=end_turn with text still became "completed". Switch to StderrPipe + an explicit copier, drain it (bounded by the existing grace window, since qodercli can leave a child holding the inherited fds) before the decision, and run the shared promoteACPResultOnProviderError. Tests: replace the convertMcpConfigForACP unit tests with two end-to-end Qoder tests — one asserts the Authorization header reaches the session/new payload as {name, value}, the other asserts a terminal stderr error with non-empty output reports failed. Co-authored-by: Orca <help@stably.ai> * fix(qoder): align ACP session handling Co-authored-by: Orca <help@stably.ai> * fix(agent): guard qoder late output after drain Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Orca <help@stably.ai> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: J <j@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
5fd3d01d13 |
MUL-3502: OST-1161: Bound assignment comment catch-up
Squashed PR #4392. Updates assignment/comment catch-up guidance to use recent 10 and aligns related examples. |
||
|
|
b4c9e4423c |
test: enable -race detector in Go test pipeline (WOR-61) (#4274)
* test: enable -race detector in Go test pipeline (WOR-61) Add the -race flag to all three Go test invocation sites so the existing concurrency regression harness (workdir_race_test.go for #3999, runtime_gone_test.go, runtime_profile_drift_test.go) actually exercises the race detector. The daemon package alone has 28+ goroutine launch points with no automated race coverage before this change. Sites updated: - Makefile:299 (make test, local) - .github/workflows/ci.yml:101 (CI backend job) - .github/workflows/release.yml:55 (release verify job) go test already runs a vet subset by default, so no separate -vet flag is added. No production code touched. Co-authored-by: multica-agent <github@multica.ai> * test(execenv): serialize runtimeGOOS-mutating test (WOR-61) TestInjectRuntimeConfigIssueMetadataCodexFormattingUnchanged called t.Parallel() while mutating the package-level runtimeGOOS to drive the windows/linux branches, racing with the other parallel tests that read runtimeGOOS in buildMetaSkillContent. The -race flag enabled in the prior commit surfaced it as 3 WARNING: DATA RACE reports and 11 "race detected" failures in CI (only the execenv package failed). Drop t.Parallel() and add the "// Not parallel: mutates the package-level runtimeGOOS." comment already used by the six sibling writer tests across execenv_test.go and reply_instructions_test.go. This is test-isolation only; no production code, no mutex/atomic, no signature change. Verified locally: go test -race -count=1 ./internal/daemon/execenv/ -> ok 2.276s go test -race -count=1 ./internal/daemon/... -> all 3 pkgs ok Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: hzz <331380069@qq.com> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
1279f22d1c |
MUL-3325: add background task safety brief (#4257)
* fix(daemon): add background task safety brief Co-authored-by: multica-agent <github@multica.ai> * fix(agent): force Claude background tools foreground Co-authored-by: multica-agent <github@multica.ai> * fix(agent): narrow Claude async launch detection Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: J <j@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
18a58e80c0 |
MUL-3316: fix(execenv): switch agent prompt to --content-file to prevent heredoc flag swallowing (#4182) (#4191)
* fix(execenv): switch agent prompt to --content-file to prevent heredoc flag swallowing (#4182) The Linux/macOS reply template recommended --content-stdin with a quoted HEREDOC. That pattern is safe for the trivial single-flag comment-add case that BuildCommentReplyInstructions emits, but as soon as a model wraps extra flags around the heredoc on multica issue create / update — assignee, project — the bash heredoc/flag boundary is fragile in two ways the model cannot see: - A 'BODY \\' terminator with a trailing token is not recognised as the heredoc end, so flag lines after it are swallowed into the description (OXY-78: residual flag text leaked into the description, command exit 0). - A clean terminator turns the trailing '--assignee ...' line into a separate failing shell statement, while the create itself already exited 0 with no assignee (OXY-76: assignee silently dropped, no residual text). In both cases the CLI never receives the swallowed flags, the API request omits the fields, and the daemon has no visibility. The created issue lands with assignee_id: null / project_id: null. This commit: * Switches the Linux/macOS branch of BuildCommentReplyInstructions to --content-file with a 3-step recipe (write file, post, rm) so the body never reaches the shell and all flags live on one shell-token line. There is no heredoc boundary for flags to leak across. * Adds a parallel cleanup step (Remove-Item) to the Windows branch so the cross-platform template is one shape. * Rewrites the runtime_config.go ## Comment Formatting non-Windows section to mandate --content-file and explicitly ban --content-stdin HEREDOC for agent-authored comments, citing #4182. * Reorders the Available Commands menu lines for issue create / update / comment add to put --content-file / --description-file ahead of the stdin variant and add a per-line note pointing at #4182. * Updates and renames the affected tests (TestBuildCommentReplyInstructionsCodexLinux, TestBuildCommentReplyInstructionsNonCodexLinux, TestInjectRuntimeConfigLinuxCommentFormattingEmphasizesFile, TestInjectRuntimeConfigIssueMetadataCodexFormattingUnchanged) so the new file-first contract is pinned and the old HEREDOC mandate is in the banned-strings lists. This converges Linux/macOS with the long-standing Windows file-only path, so the cross-platform guidance is now one shape. It also strictly improves on the previous MUL-2904 guardrail by eliminating shell exposure of the body entirely (no body ever reaches the shell, so backtick / $() / $VAR substitution cannot corrupt it). Closes GitHub multica-ai/multica#4182. No CLI or backend changes — --content-file / --description-file already exist. Co-authored-by: multica-agent <github@multica.ai> * docs(prompt): correct stale BuildPrompt comment to file-first (#4182) --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> Co-authored-by: CC-Girl <cc-girl@multica.ai> |
||
|
|
24b162cdbc |
feat(daemon): surface the real task initiator to the agent runtime (MUL-2645) (#3899)
* feat(daemon): surface the real task initiator to the agent runtime (MUL-2645)
In a multi-person workspace the agent runtime only ever saw the runtime
OWNER identity: the brief's `## Requesting User` is sourced from
runtime.OwnerID and the task-scoped token is owner-bound, so every
requester (whoever commented, @mentioned, or chatted) appeared to the
agent as the owner. Agents that route by initiator for permission,
privacy, or audit all misjudged.
Resolve the real task initiator at claim time and surface it distinctly
from the owner:
- comment / mention trigger -> triggering comment's author (member or agent)
- chat task -> chat session creator (sessions are creator-only)
- on-assign / autopilot / quick-create -> no attributable initiator (omitted)
Adds initiator_{type,id,name,email} to the claim response, the daemon
Task, and TaskContextForEnv, rendered into the brief as a new
`## Task Initiator` section. The section documents the privacy boundary:
the agent's credentials stay owner-scoped, so this is an attested
identity for the agent's own routing/privacy logic, not act-as. No DB
migration — both paths are derivable from existing rows.
Tests: brief rendering (member/agent/omit/sanitize) + email guard unit
tests, and claim-handler tests for the comment and chat paths.
Co-authored-by: multica-agent <github@multica.ai>
* fix(chat): store real sender as task initiator, not chat_session creator (MUL-2645)
Review fix (Niko, PR #3899). v1 resolved the chat task initiator from
chat_session.creator_id at claim time. That is correct for web chat and
Lark p2p (creator == sender), but WRONG for Lark group chats: the group
session creator is deliberately the installer (stable identity across
member churn), not the message sender. So in a Lark group, every member
who triggered the agent showed up in the brief as the installer/owner —
the exact bug this issue is about, still live at that entry point.
Capture the real sender at enqueue time instead of deriving it from the
session creator at claim time:
- migration 117: agent_task_queue.initiator_user_id (FK user, ON DELETE
SET NULL); NULL for non-chat and pre-migration rows.
- EnqueueChatTask now takes an explicit initiatorUserID. Web chat passes
the authenticated request user; the Lark dispatcher threads the inbound
sender (binding.MulticaUserID) through scheduleRun -> flushChatRun. The
debouncer keeps the latest scheduled flush per session, so in a multi-
sender silence window the LATEST sender wins (documented + tested).
- claim handler resolves the initiator from task.initiator_user_id and
drops the creator_id fallback entirely.
The Lark group session creator stays the installer (unchanged) — only the
task initiator is corrected, keeping the two concepts cleanly separate.
Tests: dispatcher group regression (initiator = sender, not installer),
latest-sender-wins, p2p initiator assertion; the chat claim handler test
now sets creator != initiator and asserts the stored sender wins.
Co-authored-by: multica-agent <github@multica.ai>
---------
Co-authored-by: J <j@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
|
||
|
|
b9334dd59f |
fix: anchor comment triggers to thread roots (#3746)
Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
d6a556bdbf |
fix(execenv): refresh skills in place on reuse instead of accumulating duplicate dirs (#3716)
Re-dispatching the same agent on the same issue reuses the persistent workdir via execenv.Reuse(), where the standard-provider skill refresh re-wrote skills without clearing the prior dispatch's output, so allocateCollisionFreeSkillDir dodged Multica's own directories into issue-review-multica-N. On reuse, reclaim the platform-owned managed skill directories the prior manifest recorded (removeReusedManagedSkillDirs) and roll back the remaining sidecar files (CleanupSidecars) before refreshing, so each skill lands at its canonical slug every dispatch. Mirrors the Codex hydrateCodexSkills wipe; scoped to reuse, which never runs for local_directory tasks. Fixes #3684 (MUL-2963). |
||
|
|
888186b183 |
fix(daemon): make comment-posting guardrail provider-agnostic (MUL-2904) (#3654)
* fix(daemon): make comment-posting guardrail provider-agnostic (MUL-2904) Agents inlining a backtick-wrapped token into `multica issue comment add --content "..."` had the shell run it as a command substitution, silently deleting the token; the stored comment never matched the model's intent, so it retried forever — spamming OKK-497 with duplicate comments. The corruption is shell-driven, not provider-driven, so extend the "never inline --content; use --content-file / quoted-HEREDOC --content-stdin" rule from Codex-only to ALL providers: - BuildCommentReplyInstructions: collapse the Linux/macOS non-Codex inline branch into the unified quoted-HEREDOC stdin template. - buildMetaSkillContent: rename "Codex-Specific Comment Formatting" -> "Comment Formatting" and emit it for every provider; strengthen the Available Commands entry and the assignment step-6 examples to steer away from inline --content. - Windows behavior unchanged (file-only; avoids PowerShell ASCII drop). Tests: flip the non-Codex Linux reply test into a MUL-2904 regression, broaden the stdin-emphasis test across providers, and pin the provider-agnostic guardrail. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): keep Windows assignment brief file-only (address review) Review catch on #3654: the previous commit added platform-agnostic prose recommending "--content-file or --content-stdin" in the Available Commands entry and the assignment-triggered step-6 example. The assignment path has no BuildCommentReplyInstructions OS override, so on Windows an agent following step 6 literally would pipe its final comment through PowerShell and drop non-ASCII bytes (#2198 / #2236 / #2376) — contradicting this PR's own Windows file-only rule in the ## Comment Formatting section. Make the platform-agnostic surfaces defer to the OS-aware ## Comment Formatting section (the single source of truth) instead of naming stdin. The flag synopsis still lists all three modes. Add TestInjectRuntimeConfigWindowsAssignmentBriefStaysFileOnly: a Windows assignment-triggered brief must not contain any prescriptive "... or --content-stdin" recommendation. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: J <j@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
3c8645e546 |
feat(cli): add squad member set-role (#3583)
Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
4ae4722ef0 |
fix(comments): preserve direct parent on replies (#3579)
Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
973a43923f |
fix(comments): revert since-delta to issue-wide, steer to parent thread first (#3535)
#3509/#3523 scoped the comment-trigger since-delta count to the triggering
thread, so an agent resuming a busy issue only saw "+N in this thread" and
lost visibility of new comments in other threads. Revert the count to
issue-wide (every thread), keeping the trigger-comment + agent-own
exclusions, and reshape the warm-path hint to:
- report the issue-wide new-comment volume,
- steer the agent to read the triggering (parent) thread FIRST
(`--thread <trigger> --since`, or `--tail 30` for full context),
- demote the issue-wide `--since` catch-up to an only-if-needed fallback
("don't read them all blindly").
Also fixes the now-stale "scoped to the triggering thread" wording in the
resumed-session no-delta hint (it's issue-wide zero now).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>
|
||
|
|
d1c7d478e1 |
MUL-2785: clarify thread-scoped comment delta (#3523)
Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
9616d78e47 |
MUL-2785: optimize resumed comment reads (#3509)
* feat(comments): skip default thread read on resumed comment sessions Co-authored-by: multica-agent <github@multica.ai> * fix(comments): scope since delta to trigger thread Co-authored-by: multica-agent <github@multica.ai> * chore(comments): address thread delta review nits Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
3187bbf90c |
feat(comments): re-add since-delta + cold-start thread read + parent-root write normalization (#3494)
* feat(comments): since-delta new-comment hint + default-on comment session resume (#3432) * feat(db): add unresolved comment count + list filter queries Add CountUnresolvedComments (excludes the agent's own comments) and ListUnresolvedCommentsForIssue. Both are additive — existing callers stay on the unfiltered queries — so old clients are unaffected. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(handler): support unresolved-only comment listing Wire an additive `unresolved` query param into ListComments. Defaults off so an old CLI that never sends it gets unchanged behavior; only true/1 enable it. Rejects combining unresolved with thread/recent (whole-issue filter vs navigation models). Includes filter + count query tests. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(handler): plumb unresolved count + thread root into claim, gate comment resume Populate trigger_parent_id (thread root of the trigger comment) and unresolved_count (excludes the agent's own comments) on comment-triggered claim responses. Both fields are omitempty so old daemons ignore them. Gate comment-triggered session resume behind MULTICA_RESUME_COMMENT_SESSION (default off): resumed comment turns can inherit the prior turn's "Done." final message, so this stays an explicit rollout switch. The runtime-match and poisoned-session guards still apply regardless of the flag. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(daemon): inject unresolved-comments hint + resolve step into agent brief Add a shared BuildUnresolvedCommentsHint helper rendered on both the per-turn prompt and the CLAUDE.md workflow (kept in sync per PR #2816). It ships only the count and the relevant CLI call — never comment bodies — so the server stays cheap. Thread case points at --thread <root>; issue case points at --unresolved. Suppressed when the count is 0. Also add a workflow step telling the agent to `multica comment resolve <thread-root>` once a thread is fully handled, so the unresolved set converges. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(cli): add comment list --unresolved and comment resolve command Add an --unresolved filter to `issue comment list` (wired to the server's unresolved param, rejected when combined with --thread/--recent) and a top-level `comment resolve <id>` command that POSTs to the existing /api/comments/{id}/resolve endpoint, letting an agent close threads it has fully handled. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(comments): since-delta new-comment hint + default-on comment resume Simplifies the comment-triggered agent flow down to what's actually needed: - New-comment awareness is now a pure time delta: the claim response carries new_comment_count + new_comments_since (anchored on the prior run's started_at, never completed_at so a long run can't miss comments). The per-turn prompt and CLAUDE.md workflow render one line — "N new comment(s) since your last run, --since <ts>" — via a shared BuildNewCommentsHint so the two surfaces can't drift. Cold start (no prior run) falls back to a plain read. - Comment-triggered tasks resume the prior session by default (same runtime), dropping the MULTICA_RESUME_COMMENT_SESSION rollout gate. The "Focus on THIS comment" prompt guard defends against inheriting the prior turn's "Done." marker; GetLastTaskSession still excludes poisoned sessions. - Drops the resolved-based machinery from the first draft: CountUnresolvedComments / ListUnresolvedCommentsForIssue queries, the `comment list --unresolved` flag, the `multica comment resolve` command, and the resolve workflow step. - Removes the verbose cursor-pagination paragraph from the comment prompt; the --thread/--recent/--since flags stay in the CLI/API, just no longer explained inline every turn. Compatibility: new claim fields are omitempty (old daemons ignore them). Comment resume is default-on and affects even old daemons, which already consume prior_session_id. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(comments): collapse reply parent_id to thread root on write Comment threads are a 2-level model (root + flat replies, like Linear/Slack), enforced today only by the UI and the agent path — the CreateComment handler stored whatever parent_id it was handed, and the agent-side flatten walked just one level, so a reply-to-a-reply could land at depth 3+. Add GetThreadRoot (a recursive walk to the parent_id=NULL root) and run both write paths (handler.CreateComment, service.createAgentComment) through it, so every stored reply's parent_id IS its thread root. Readers can now treat parent_id as the thread root without re-walking. The agent-drift guard still compares the raw parent_id to the trigger comment before normalization. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(comments): cold-start reads triggering thread, warm keeps --thread pointer The since-delta rework dropped the thread-first read on the COLD path: a first-time agent fell back to the flat `comment list` dump (oldest-first, cap 2000), burying the trigger's context in ancient chatter. Point cold start at the triggering conversation instead via a shared BuildColdCommentsHint (`--thread <trigger> --tail 30` + a --recent pointer for cross-thread background). On the WARM path, --since is a pure time delta and can miss the triggering thread's pre-anchor history, so BuildNewCommentsHint now also emits a --thread pointer. Both surfaces (per-turn prompt + CLAUDE.md workflow) render via the shared helpers so they cannot drift (PR #2816 rule). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |