mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-05 17:40:11 +02:00
v0.4.16
234 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cd9b956269 |
fix(agent): spawn Copilot's native binary on Windows so the prompt survives (MUL-5586) (#6236)
On Windows the daemon passes the full multi-line prompt as `-p <prompt>` but spawns npm's `copilot.cmd`, which we already rewrite to `powershell -File copilot.ps1`. Neither launcher can carry that argument: - `copilot.cmd` forwards with `%*`, which cmd.exe expands by re-tokenising the raw command line. - `copilot.ps1` ends in `& node.exe npm-loader.js $args`, and PowerShell re-serialises `$args` onto node's command line. Under Windows PowerShell 5.1 (and pwsh <= 7.2, which default to Legacy native argument passing) embedded double quotes are not re-escaped, so the prompt is re-tokenised. Copilot then sees several argv tokens where one was intended and refuses the run with "It looks like your prompt was not quoted, so the extra words were treated as separate arguments" — the same defect class already fixed for cursor-agent in #5649, except Copilot has no stdin prompt channel to escape through, so the prompt must stay on the command line and the launchers have to go. Copilot CLI ships a native per-platform binary and `npm-loader.js` does nothing but `spawnSync` it with argv untouched, so resolve `copilot-win32-{x64,arm64}\copilot.exe` out of the npm layout and spawn it directly. That leaves exactly one hop, Go -> native binary, and Go's syscall.EscapeArg is the exact inverse of the CRT parsing that binary uses. This mirrors resolveOpenCodeNativeFromShim / resolveDevecoNativeFromShim. Both the nested (current npm) and hoisted (older npm) platform-package locations are probed; when neither resolves, we keep falling back to the PowerShell launcher, which is still better than cmd.exe. Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
e4b6f7a31b |
MUL-5581: add Qoder CN CLI runtime (#6232)
* feat(agent): add Qoder CN CLI runtime Co-authored-by: multica-agent <github@multica.ai> * fix(agent): address Qoder CN review nits Co-authored-by: multica-agent <github@multica.ai> * fix(agent): defer Qoder CN version gate Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
cee985592e | fix(agent): drain trailing ACP notifications in kimi, kiro, qoder, and traecli (#5951) | ||
|
|
75c11db048 |
MUL-5549: feat(agent): discover codebuddy models over ACP instead of scraping --help (#6203)
* feat(agent): discover codebuddy models over ACP instead of scraping --help (MUL-5549) CodeBuddy speaks ACP, and `session/new` answers with a structured catalog under models.availableModels plus a currentModelId — exactly the shape the shared parseACPSessionNewModels already reads for Copilot / Kimi / Kiro / Qoder / Grok / TRAE. Scraping the `--model` line out of `codebuddy --help` was never necessary. The help text carried IDs and nothing else, which cost us three things: - Labels were guessed from the ID and were simply wrong. `kimi-k3-1` rendered as "Kimi K3 1" where the CLI says Kimi-K3; `deepseek-v3-2-volc` as "Deepseek V3 2 Volc" where the CLI says DeepSeek-V3.2. - The default model was a "first entry wins" guess rather than the advertised currentModelId. - The effort catalog needed a second regex over the same output. All three come from the handshake now. The effort catalog rides along in the same session/new response as the `thought_level` config option, so it costs no extra process — which also retires the "at most one --help per request" constraint added in #6196, because --help is no longer run at all. One trap worth naming: thought_level advertises `enabled` ("On (default)") alongside the six real levels, but `--effort enabled` is not a valid command line — the daemon passes the selected level straight to the flag. Advertised levels are filtered against the flag's accepted set, and a currentValue outside that set (the default `enabled`) becomes an empty DefaultLevel, which the UI renders as a generic "Default" instead of a value we cannot pass through. Two adjacent inaccuracies surfaced while confirming the real level set against CodeBuddy 2.130.0, both fixed here: the static effort fallback omitted `minimal` and `max`, and so did the server-side IsKnownThinkingValue gate — so the server rejected two levels the CLI genuinely accepts. Discovery keeps its fallback, still marked Fallback so it can never be cached as authoritative (#6196). That covers the not-logged-in case, which is deliberately NOT special-cased with an auth step: the catalog came back without calling authenticate on a logged-in CLI, and inventing an auth branch we cannot exercise would be speculation. Removes codebuddyModelRe, parseCodebuddyModels, codebuddyModelLabel, codebuddyModelProvider, codebuddyEffortRe, parseCodebuddyEffortHelp, codebuddyEffortSuperset, codebuddyHelpOutput and its 60s help cache. Co-authored-by: multica-agent <github@multica.ai> * fix(agent): keep codebuddy's vendor grouping after the ACP migration (MUL-5549) Review nit, and a real regression in the previous commit. Dropping codebuddyModelProvider looked like removing dead code, but it was the only thing populating Model.Provider for CodeBuddy — and the picker groups on that field. acpModelEntry can only recover a vendor from a `vendor:model` id. CodeBuddy's are bare (`glm-5.2`, `kimi-k3-1`), so every model came back with an empty Provider, and model-dropdown renders the empty group with no header at all: all 16 models would have collapsed into one unlabelled list where main shows Zhipu / Kimi / MiniMax / DeepSeek / Hunyuan sections. Restores the prefix inference as a post-pass over the ACP catalog, exactly the shape discoverCopilotModels already uses for the same reason. Verified against the real CLI: all 16 models land in five vendor groups with none ungrouped. Tests assert the vendor for every id CodeBuddy 2.130.0 advertises plus the static fallback ids, and that the fallback entries' hardcoded providers agree with the inference. Removing the post-pass fails them. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
44ce16d9b8 |
MUL-5549: fix(agent): stop reporting a failed model discovery as a real catalog (#6196)
* fix(agent): stop reporting a failed model discovery as a real catalog (MUL-5549) Selecting the CodeBuddy runtime showed a model list that shares no IDs with what the CLI actually supports, so every pick was an ID codebuddy rejects (GH #6180). The list in the report is codebuddyStaticModels() verbatim: the daemon had fallen back, but nothing downstream could tell. discoverCodebuddyModels returned (staticModels, nil) on all three failure paths, and copilot/cursor/grok do the same. A failed discovery therefore arrived as a successful one, which defeated every guard built to catch it: the daemon reported status "completed", the picker's discovery_failed hint only renders on isError, and cacheableModelCatalog — whose own comment says an empty list means transient failure — waves through a non-empty stand-in and stores it as last-known-good for the full 24h serve window. One blip got pinned as the answer for a day. Discovery now returns a Catalog carrying a Fallback marker, which the daemon forwards as an additive `fallback` field (older servers ignore it; an older daemon omitting it keeps the previous behaviour). A fallback catalog is still rendered — the picker stays populated and manual entry still works — but it is kept out of both the daemon's 60s discovery cache and the server's catalog cache. On the server it maps to Keep rather than Drop: a stand-in is no grounds to evict a real catalog, matching how a `failed` report is treated. Also stop codebuddyHelpOutput swallowing the exec error. CombinedOutput folds in stderr, so a codebuddy whose `#!/usr/bin/env node` interpreter is missing from a GUI-launched daemon's PATH had `env: node: No such file or directory` parsed as help text — and cached as such for 60s. Verified against CodeBuddy CLI v2.130.0: the parser itself is fine (16 models from real --help), so this fixes the reporting of the failure, not the parse. Co-authored-by: multica-agent <github@multica.ai> * fix(agent): run codebuddy --help at most once per model-list request (MUL-5549) Review catch on the previous commit. Model discovery and effort discovery both read `codebuddy --help`, and the effort pass called it independently. That was free while a failed --help was (wrongly) memoised, but once failures correctly stopped being cached, the failure path ran the 35s command twice in a single request — past the server's 60s running timeout, so the request timed out and the late report was then discarded as stale. The user got nothing, not even the fallback list the previous commit exists to preserve. discoverCodebuddyModels now owns the thinking annotation, so the one help capture feeds both catalogs, and the failure path uses codebuddyFallbackCatalog to apply the static effort levels without exec'ing at all: whatever broke --help for the model catalog breaks it for the effort catalog too. Also strengthen the handler tests. They decoded into a struct declared in the test rather than calling ReportModelListResult, so a wrong JSON tag or a mis-wired cache branch would have passed. They now drive the real endpoint with daemon auth and chi params, covering: a fallback report leaving a previously discovered catalog intact, an older daemon omitting the field still warming the cache, and an authoritative empty catalog still dropping the snapshot. Both fixes are mutation-tested — reverting either makes the new tests fail. Co-authored-by: multica-agent <github@multica.ai> * docs(agent): correct codebuddy --help comments after the single-capture refactor (MUL-5549) Review nit. The comments still described the pre-refactor call graph, where both discoverCodebuddyModels and codebuddyEffortSuperset called codebuddyHelpOutput and the cache was what stopped the duplicate run. The effort parser now takes an already-captured string, and the single-invocation guarantee is structural rather than cache-dependent — which matters, because a failed --help is deliberately not cached, so a second caller would re-run the full 35s timeout. Also note on codebuddyHelpOutput that it has exactly one caller and why a new one would reintroduce the bug. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
99d8f29bde |
fix(codex): raise the first-turn no-progress ceiling to 60s (MUL-5542) (#6192)
The first-turn watchdog killed healthy turns. Two independent field reports measured the first progress event landing just past 30s on gpt-5.5, and ~39s for a WSL app-server, all inside the window the watchdog treats as "stuck". Raise the ceiling to 60s. The window's only job is to fail fast instead of waiting out the 10 minute semantic inactivity backstop, so 60s keeps that value while clearing the observed evidence with margin. Add a regression test for the clamp, which had none: the configured semantic inactivity timeout can only shrink this window, never raise it. Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
1d2a3499c9 |
test(agent): stop the cursor fixtures racing their own prompt write (MUL-5536) (#6174)
TestCursorExecuteFailsOnCleanEOFWithoutResult failed on main with "cursor-agent prompt write failed: write |1: broken pipe" where it expects "stream ended without terminal result". The fake cursor-agent exits without reading stdin, so the prompt write races the child's exit: win and the pipe buffer swallows it, lose and the read end is gone and the write returns EPIPE. writeErr outranks both exitErr and the generic no-terminal-result error in cursor.go, so a lost race replaces the asserted failure with the EPIPE one. A real cursor-agent reads stdin to EOF, so the fixtures now do too. That removes the race rather than reordering production error precedence, which is deliberate. Only the two fixtures whose expected error ranks below writeErr need the drain. Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
999e9f93c7 |
fix(codex): record file edit payload for patch_apply events (#6158)
* fix(redact): scrub secrets nested inside tool input maps and slices
InputMap only passed top-level string values through Text and documented
non-string values as "preserved as-is". Any secret one level down reached
the database and the WebSocket broadcast untouched:
flat -> [REDACTED ...] (scrubbed)
nested -> [map[diff:token=ghp_... path:a.go]] (leaked verbatim)
This is a prerequisite for recording structured file-edit payloads. Codex
reports an edit as changes[]{path, diff, content}, and the legacy protocol
reports a deletion as the whole outgoing file — so without this, deleting a
.env would persist its full contents in cleartext.
redactValue now walks the composite shapes json.Unmarshal produces, plus
[]string and map[string]string for argv-style inputs. Composites are copied
rather than scrubbed in place, because the caller keeps using the map it
passed in.
Nesting depth comes from daemon-supplied JSON, so the walk is bounded at 32
levels; a pathologically nested payload would otherwise recurse until the
stack blows. Hitting the bound yields a placeholder rather than the raw
value, keeping the fail-safe direction.
Verified: the five new tests each fail against the previous top-level-only
implementation and pass now; full ./pkg/redact suite green.
Co-authored-by: multica-agent <github@multica.ai>
* fix(codex): record file edit payload for patch_apply events
Both Codex protocol paths recorded a file edit as a bare call ID and set no
payload, so a run that edited six files left six blank, unexpandable rows in
the transcript. The same task on Claude or Grok showed a readable diff, and
the branch Codex pushed was the only surviving record of what it changed
(GH #6157).
The omission was specific to this one tool, not to the adapter: the
exec_command handlers directly above already captured command and output.
Both paths are fixed, since the protocol is sniffed at runtime. Their wire
shapes differ more than they appear, and the normalizer reconciles that:
- legacy patch_apply_begin/end carry map[path]FileChange, internally tagged
on `type`, where add/delete hold whole-file `content` and only update
holds a `unified_diff` plus `move_path`. There is no diff for every case,
so the normalized form keeps diff and content as alternatives.
- v2 fileChange items carry an ordered array of {path, kind, diff} where
`kind` is an object, not a string — reading it as a string silently
yields "" and loses the add/delete/update distinction.
- status spellings differ too: legacy is snake_case, v2 is camelCase and
adds inProgress. Both normalize onto one vocabulary, and a legacy event
predating `status` falls back to its `success` bool.
Legacy map iteration is sorted by path so a replayed event does not reshuffle
the file list.
Completion events now also produce a non-empty output (status, file count,
and any apply_patch stdout/stderr), because an empty output renders as an
unexpandable blank row just like a missing input.
Anything unrecognised — absent, wrongly typed, or malformed changes — returns
no payload, preserving exactly the previous degradation rather than risking
the transcript.
Total diff/content bytes are bounded at 64 KiB with UTF-8-safe truncation,
recording `truncated` and `original_bytes`; paths and kinds always survive,
since they are what a reviewer needs when the body is gone. The bound is
deliberately scoped to this new payload: other providers stream tool inputs
through unbounded, and clamping them here would silently truncate
transcripts that render correctly today. Unifying the limit at the
persistence boundary is left as a follow-up.
Verified: the new tests reproduce the reported symptom (Input:map[],
Output:"") against the previous call sites and pass now; ./pkg/agent and
./pkg/redact green, go vet and gofmt clean.
Co-authored-by: multica-agent <github@multica.ai>
* feat(transcript): render Codex multi-file patch payloads as diffs
The presenter identified an edit by input shape — a top-level file_path plus
old_string/new_string or content — which is Claude's and Grok's shape. Codex
records one patch_apply covering several files as changes[], so even with the
payload now populated it fell through to pretty JSON instead of a diff.
A new `patch` detail kind carries one entry per file, since collapsing them
into a single body would lose which change belongs where. Each file reuses the
existing single-file surfaces, so all bodies behave alike inside the
virtualized list.
Codex hands over a ready-made unified diff, so parseUnifiedDiff maps it onto
diff rows rather than recomputing one — there is no before/after pair to
compare, and reconstructing both sides from the diff just to diff them again
would be circular. Hunk headers become `gap` rows, which is what they denote:
skipped unchanged content.
A deletion renders as all-removals rather than as a whole-file write, because
the legacy protocol reports it as the outgoing file's content and a green
"+N" gutter would state the opposite of what happened.
The collapsed row needed its own fix: with no single path field, the summary
fell through the preference chain and came back empty. It now reads as the
first path plus "+N more".
Anything that is not this shape still falls back to pretty JSON, so a payload
this presenter does not understand stays readable.
Verified: 17 new tests (43 in the presenter suite) pass; repo typecheck and
lint clean. The one failing views test, layout/sidebar-resize, fails
identically on an untouched checkout.
Co-authored-by: multica-agent <github@multica.ai>
* fix(codex): route v2 add/delete payloads as content, not diff
Addresses review on #6158.
Upstream's format_file_change_diff only produces a unified diff for `update`.
For `add` and `delete` it returns the whole file's contents under the same
`diff` field, and for a moved `update` it appends a trailing
"\n\nMoved to: <path>" line:
FileChange::Add { content } => content.clone(),
FileChange::Delete { content } => content.clone(),
FileChange::Update { unified_diff, move_path } => ...
(codex-rs/app-server-protocol/src/protocol/item_builders.rs, rust-v0.145.0)
Recording that as a diff mislabels every line of an added or deleted file as
context, and actively inverts any line whose content begins with '+' or '-' —
so an added file containing "-minus lead" rendered as a deletion. The payload
is now routed by `kind` rather than by field name, and the "Moved to:"
sentence is stripped since move_path already carries the destination.
The previous v2 tests hid this by using a fixture the real protocol never
emits (an `add` carrying "@@ ... +package main"). They now use upstream's
shape, plus cases for delete, an empty add, and an add whose contents look
like diff headers.
Empty bodies are also kept on both paths: presence of the field, not its
non-emptiness, decides whether a body was reported, so an empty added file
renders as an empty body instead of "no content reported".
Verified: the new assertions fail against the previous normalizer — where an
`add` came through as {"diff": "package main\n"} — and pass now.
Co-authored-by: multica-agent <github@multica.ai>
* fix(transcript): stop treating header-like content lines as file headers
Addresses review on #6158.
parseUnifiedDiff matched "---" / "+++" / "diff --git" / "index " at any
position, so a changed line whose *content* starts with a dash or plus was
silently discarded:
parseUnifiedDiff("@@ -1 +1 @@\n--- old markdown\n+++ new markdown\n")
// before: [{ kind: "gap", ... }] — both changed lines gone
A removal of "-- old markdown" is spelled "--- old markdown" on the wire, so
this hit Markdown rules, embedded patches, and comment banners.
File headers only exist ahead of the first hunk, so they are only recognised
there; once inside a hunk every line is parsed strictly by its first
character.
Also localizes the multi-file summary count, which was hardcoded English and
so leaked into the zh-Hans / ja / ko transcript rows. The presenter owns no
React and no i18n by design, so the phrasing is injected by the caller rather
than imported here, keeping the module unit-testable in isolation; the English
form remains the fallback. The three Chinese/Japanese/Korean truncation
strings now use "..." to match the English source they translate.
Verified: both new parser assertions fail against the previous
strip-anywhere behaviour and pass now; 47 presenter tests green, repo
typecheck and lint clean.
Co-authored-by: multica-agent <github@multica.ai>
* fix(daemon): redact nested tool input before it leaves the daemon
Addresses review on #6158.
Recursive redaction ran only in the server's ingest handler. The daemon built
the new nested edit payload and sent msg.Input verbatim, so a daemon that
self-updated ahead of the server — or one talking to a server mid-rollout —
would ship whole-file edit contents to a peer that does not scrub nested
values yet. The legacy protocol reports a deletion as the whole outgoing file,
so that window covered a deleted .env in cleartext.
Ordering three commits inside one PR is not a deployment barrier, and daemon
and server upgrade independently. Deployment order is not a control we have,
so the sending side is now safe on its own; the server keeps redacting on
ingest as the second line of defence.
Scoped to Input, which is the field this PR newly fills with file contents.
Content and Output are plain strings already redacted server-side, and
changing their daemon-side handling would be unrelated to this fix.
Verified: the new daemon test asserts the nested token is masked in the
reported batch while the change metadata survives. It fails without this
change, reporting the full GITHUB_TOKEN= line on the wire, and passes with
it; ./internal/daemon green.
Co-authored-by: multica-agent <github@multica.ai>
* fix(transcript): correct the Chinese multi-file patch count semantics
Addresses review on #6158.
The summary is handed the number of files *beyond* the named one, but the
Chinese phrasing stated a total: "a.go 等 2 个文件" reads as two files including
a.go, so a three-file patch under-reported by one. English hides the
distinction ("+2 more"), which is why it survived the first pass.
Rewords zh-Hans to "另有 N 个文件". Japanese (他) and Korean (외) already read
as "besides", so their wording is unchanged.
Also renames the interpolation variable from `count` to `extra`, for two
reasons. i18next treats `count` as the plural selector — this very namespace
relies on that for events_one/events_other — so a plain number had no business
borrowing it. And the name is what a translator reads: `extra` cannot be
mistaken for a total the way `count` was.
Guards the whole bug class rather than just this string: a locale test asserts
every locale interpolates {{path}} and {{extra}} and never the reserved
{{count}}, and a presenter test pins that the injected number is the count of
additional files, not the total.
Verified: both new locale assertions fail against the reverted string and pass
now; rendering the real locale strings for a three-file patch yields "+2 more",
"另有 2 个文件", "他 2 件", "외 2개". 53 target tests pass, repo typecheck clean,
views lint back to its pre-existing 16 warnings and 0 errors.
Co-authored-by: multica-agent <github@multica.ai>
* fix(transcript): put the patch surface on the type scale
Addresses review on #6158.
The patch surface wrote text-[10px] / text-[11px] / text-[10px], copied from
the sibling transcript surfaces as they looked when this branch started. Since
then MUL-5451 (#6136) introduced a role-named type scale and migrated those
same siblings to text-micro, so these three call sites were the only remaining
arbitrary sizes — and the type-scale guard reports them precisely.
All three become text-micro. That matches the analogues they were copied from
now that those have moved: the FileWriteSurface line-count row, the
DiffDetailSurface header row, and the ToolDetailSurface body. It is also the
only correct target, since micro (11px) is the smallest step the scale defines
— there is nothing at 10px to map to.
Merges origin/main so the guard runs here rather than only in CI.
Verified: apps/web app/type-scale.test.ts 13/13 (it listed exactly these three
lines before), no `text-[` left in the file, repo typecheck clean, views lint
unchanged at 16 pre-existing warnings and 0 errors.
Co-authored-by: multica-agent <github@multica.ai>
* fix(codex): redact patch bodies before applying the size budget
Addresses review on #6158.
The adapter sized and truncated the normalized changes, and redaction only ran
later — in the daemon before sending, then again in the server on ingest. That
order loses secrets that straddle the budget.
The PEM rule needs both markers to match:
-----BEGIN[A-Z\s]*PRIVATE KEY-----.*?-----END[A-Z\s]*PRIVATE KEY-----
So a 70 KB private key whose BEGIN sits inside the first 64 KiB and whose END
falls past the cut stops matching once truncated. Neither later pass can
recognise what truncation already broke, so the marker and 64 KiB of key
material reach the database and the WebSocket broadcast. Measured on the
previous code:
stored bytes 65536 | BEGIN marker present | key body present | placeholder absent
Redaction now runs first, and the budget measures the redacted bodies — which
is also the honest measurement, since those are what actually gets stored and
redaction usually shrinks them (that key collapses to 23 bytes, so no trimming
is needed at all). `original_bytes` still reports the pre-redaction size so the
reader sees how large the real patch was. The daemon and server passes stay as
defence in depth; redaction is idempotent, so running three times is safe and
that is now asserted.
Note for callers: codexPatchInput no longer trims its argument in place, because
redaction copies first. Two existing tests were asserting on the caller's
original slice and had silently become vacuous; they now read the returned
payload, and one pins the no-mutation contract. The delete fixture in the
diff-vs-content routing test was also a credential-shaped string, which now
redacts — it is plain text so that test keeps testing routing.
Verified: the new boundary test fails on the previous order, reporting the
surviving BEGIN marker and key material, and passes now. go test ./pkg/agent
./pkg/redact ./internal/daemon green; execenv ByteIdentical green; go vet and
gofmt clean.
Co-authored-by: multica-agent <github@multica.ai>
---------
Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
|
||
|
|
d6bd4cf7d5 |
fix(agent): recover from Hermes resumed sessions it refuses without running (MUL-5509) (#6164)
* fix(agent): recover from Hermes resumed sessions it refuses without running Hermes never reports an unknown ACP session as a JSON-RPC error. Its adapter answers session/prompt for a session it cannot load with an ordinary success frame carrying stopReason=refusal, and session/resume echoes nothing back -- ACP's ResumeSessionResponse has no sessionId field at all, unlike NewSessionResponse -- so resolveResumedSessionID keeps the id we asked for. Nothing in the exchange is an error, so the isACPSessionNotFound branches at set_model and prompt time never fire for this backend. The result was a Result carrying the dead session id with ResumeRejected=false, which shouldRetryWithFreshSession reads as "not a rejection". GetLastTaskSession then handed the same dead id to every later dispatch on that (agent, issue) pair, so a Hermes agent was usable exactly once per issue: the first task worked, and every following comment, approval or @mention failed the same way until a manual rerun bought one more turn. When no provider error reached stderr the turn was worse than a failure -- it reported completed with empty output, an agent that silently did nothing. Treat a refusal on a resumed session with zero agent activity as the runtime telling us the session is gone: clear the session id, set ResumeRejected so the existing fresh-session retry and session-retirement path take over, and fail the turn instead of reporting an empty success. Both conditions are required -- stopReason=refusal alone is a legitimate model refusal, and a refusal after real work is not a lost session. The reason is applied after promoteACPResultOnProviderError so a captured provider error stays the user-visible message; the generic fallback only fills in when nothing more specific was seen. Verified against the real hermes CLI (upstream main ba7d214b6) on an isolated HERMES_HOME with a local mock endpoint: a fresh task completes, and resuming it now yields status=failed, SessionID="", ResumeRejected=true with the provider error preserved -- previously the dead id came back with ResumeRejected=false. Refs GH #6150, MUL-5509 Co-authored-by: multica-agent <github@multica.ai> * fix(agent): decide Hermes resume loss after the pipe drain, not at quiescence The notification quiet window closing does not end a turn — stdin EOF and the pipe drain do, and Hermes legitimately delivers a turn's final chunk in that gap (TestHermesBackendDrainsLateFinalNotificationAfterPromptResponse exists because of it). Freezing the resume-loss decision at the quiescence boundary therefore read turnActivity == 0 while a real answer was still in flight: a runtime that merely answered slowly after stopReason=refusal had its healthy session id cleared and ResumeRejected set, discarding a live conversation pointer and, with no tool use recorded, triggering an unnecessary fresh-session retry that re-ran the turn. Move the evaluation to the point where the turn has settled — after waitForHermesPipeDrain and streamingCurrentTurn.Store(false), where every accepted update has been counted. This also drops the resumeLost variable and the ordering hazard that came with it. The new regression sends the late chunk 600ms after the refusal, past the 250ms quiet window and inside the 2s drain grace, and asserts the session id survives with ResumeRejected=false. Confirmed to fail against the previous ordering with exactly the reported symptoms (session id emptied, ResumeRejected=true). Refs GH #6150, MUL-5509 Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
280fa28e0d |
fix(agent): deny CodeBuddy's interactive plan-mode tools in headless runs (MUL-5383) (#6104)
CodeBuddy exempts AskUserQuestion and ExitPlanMode from permission-mode finalization, so `--permission-mode bypassPermissions` never auto-approves them. Once the model entered plan mode the session mode is Plan, not BypassPermissions, so ExitPlanMode also missed the daemon's bypass fast-path and went to the SDK permission bridge — which waits with no timeout for a confirmation the headless runtime cannot render. The task sat in-flight until the 2h tool watchdog, and users killed it by hand. Deny EnterPlanMode/ExitPlanMode alongside the AskUserQuestion we already deny. Each tool is passed as its own argv value because CodeBuddy matches disallowedTools entries exactly and does not split on commas. Also send `allowed: true` on control_response: CodeBuddy's SdkPermissionClient reads `allowed`, not Claude Code's `behavior`, so the daemon's "auto-approve" was being read as a denial. Fixes #6012 Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
de3e0ae556 |
fix(agent): make cursor stream protocol drift loud instead of silent (MUL-5434) (#6089)
* fix(agent): make cursor stream protocol drift loud instead of silent (MUL-5434) #6071 reports Cursor tasks that demonstrably read files, ran commands and called the Multica CLI, yet showed a single blob of agent text with no reasoning and no tool rows. The run still reported success and tools=0. `switch evt.Type` in cursor.go had no `default` branch, so any top-level event type we do not handle was dropped with no counter and no warning. Renaming only the top-level types of a healthy stream (`thinking`->`reasoning`, `tool_call`->`tool_calls`), leaving every nested field untouched, reproduces the report exactly: status=completed, output = the result text alone, tool_use=0, thinking=0, zero diagnostics. The existing unknown-subtype warning cannot catch this — it only increments once the type has already matched — so "no unknown-subtype warning" does not rule out protocol drift. Two diagnostic gaps are closed: - Add the `default` branch with a bounded, content-free tally of unhandled top-level types, reported once per run as a warning and alongside tool_use_count in the protocol summary. Type names are normalized through observedCursorEventType and distinct names are capped at 16 plus an overflow bucket, so a hostile or noisy stream cannot grow the map or leak payload into logs. `user` (the CLI echoing our prompt, present in every recorded run) is explicitly benign so the warning does not fire always. - Count assistant text separately from the terminal result text. The result event writes into the same builder, so last_assistant_bytes equalled result_bytes even when the assistant streamed nothing — erasing the signal that says "only the final answer arrived". This also aligns cursor with claude, which already reports assistant bytes only. Unrecognized events are still never coerced into tool or reasoning messages; guessing at upstream additions is the failure mode MUL-5231 already fixed once. This is diagnosis only and does not itself restore the missing tool rows — identifying which upstream shape changed requires a captured 2026.07.23 stream, and this warning is what makes that identifiable from a single production log line. Co-authored-by: multica-agent <github@multica.ai> * fix(agent): report cursor unhandled events as evidence, not as a verdict Addresses the review's must-fix on MUL-5434. The diagnostic was described as deciding WHY a transcript is empty, which it cannot do: - "tools=0 with unhandled types" does not establish that a rename ate the tool rows. Cursor 2026.07.23 also emits transport/control frames, so an unhandled type only proves the stream carried events we do not parse. - "tools=0 with no unhandled types" does not establish the agent used no tools. The CLI may execute tools without handing the updates to its stream serializer at all — the main branch #6071 has NOT ruled out — a new shape may be nested inside an event type we already recognize, or events may be lost to invalid framing or a scanner boundary. Changes, no behaviour change to messages, status or output: - Reword the tally doc, the switch default, the warning and the shared observation struct to state what a non-zero and a zero count each do and do not establish, and point at the branches that stay open. - Rename unknown* to unhandled* (fields, log keys, warning text, the pre-existing subtype counter) so the diagnostic never implies the type is unrecognized upstream — only that this parser does not handle it. - Classify `connection` and `retry` explicitly. They join `user` in cursorNonTranscriptEventTypes, with per-entry provenance recorded: `user` is confirmed in the recorded 2026.07.20 stream, the control frames are reported on newer builds and listed defensively. A build that does not emit them makes the entry inert; one that does must not have a known control frame reported as an unhandled protocol event. Tests assert signal presence and that nothing is fabricated, not causality: the healthy-stream test now carries `connection` / `retry` and additionally asserts the real thinking/tool rows still arrive, so suppressing the warning cannot silently suppress the transcript. TestCursorNonTranscriptEventType also pins that no type the parser handles can enter the suppression list. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
b3847b4172 |
MUL-5433: fix(agent/hermes): pin session id mid-flight for daemon resume (#6070)
* fix(agent/hermes): pin session id mid-flight for daemon resume Hermes was the only ACP-backed agent that created a session without emitting a running status carrying the session id. When the daemon restarted or a task was cancelled mid-flight, PinTaskSession had nothing to key on, so the resume pointer was lost and the task could not resume. Emit MessageStatus+SessionID immediately after session create, matching claude, codebuddy, codex, grok and qwen. The daemon already listens for this (daemon.go MessageStatus -> PinTaskSession), so this small pin is all Hermes needs. See #4969 for the companion daemon/SQL cancel-salvage work. Co-authored-by: multica-agent <github@multica.ai> * test(agent): cover Hermes mid-flight session pin Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: multica-agent <github@multica.ai> Co-authored-by: Eve <eve@multica-ai.local> |
||
|
|
0066ab259e |
refactor(agent): drop unreachable inline system-prompt branches (MUL-5392) (#6050)
* refactor(agent): drop unreachable inline system-prompt branches (MUL-5392) The daemon only populates ExecOptions.SystemPrompt for openclaw, kimi and traecli (providerNeedsInlineSystemPrompt); every other backend receives the runtime brief as a per-task context file in the workdir. The inline branches in claude, codex, opencode and pi were therefore dead, and read as if they were the live delivery path. Probed each backend over its real launch path with a canary in the context file and no inline delivery — claude 2.1.220 (CLAUDE.md), codex 0.144.6 via the app-server, opencode 1.17.7, pi 0.67.2, hermes 0.18.2 via ACP — and all of them picked the brief up from disk, so the branches were removable with no behaviour change. An empty workdir returned no canary, confirming the probe could fail. opencode's branch was worse than dead: `opencode run` has no --prompt flag, so enabling inline delivery there would have made every opencode task exit 1 with a usage dump. The DevEco backend, forked from opencode, already documents this constraint; opencode itself never got the fix. Regression tests pin all three arg builders against re-adding the flag, and providerNeedsInlineSystemPrompt now documents what was verified and what is still unprobed (grok, qoder, codebuddy). Hermes and kiro are untouched: their exclusion is deliberate and already tested. Co-authored-by: multica-agent <github@multica.ai> * test(agent): pin codex developerInstructions contract, drop stale pi flag doc Review follow-up on MUL-5392. buildPiArgs' doc comment still advertised --append-system-prompt after the branch that emitted it was removed — exactly the stale-signal this PR set out to delete. The two codex sites fixed to a literal nil had no regression test, so restoring nilIfEmpty(opts.SystemPrompt) would still have gone green. Both thread/start and thread/resume now run with a canary SystemPrompt and assert developerInstructions comes through as an explicit null. Mutation-checked: reverting either site fails its test. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
962d376fda |
refactor(agent): share acpDeliverableTracker across ACP backends (MUL-5405) (#6044)
#6022 stopped qoder from delivering interim narration as Result.Output, but hermes / kimi / kiro / traecli / grok still accumulate every MessageText into one builder and hand the whole thing to Result.Output — the same leak (#6006), since Result.Output becomes the channel reply and the auto-generated issue comment. Extract the boundary rule into acpDeliverableTracker (observe / result) and share it across all six backends instead of copying qoder's block five times: Result.Output keeps only the text after the latest tool call, a turn that ends on a tool call falls back to the latest non-empty text block so the reply is never empty, and provider-error detection keeps reading the full text stream. Covered by tracker unit tests and a cross-backend regression test that pins both scenarios on all six backends, over both tool-use emission paths (emitted at the tool call and deferred to tool completion). Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
54469766fd |
fix(qoder): deliver final answer only (MUL-5394) (#6022)
* fix(qoder): deliver final answer only * fix(qoder): preserve tool-terminated replies |
||
|
|
ba108978ac |
fix(channel): deliver the final answer only to Slack and Lark (MUL-5378) (#6016)
* fix(channel): deliver the final answer only to Slack and Lark (MUL-5378) Channel replies could carry the agent's interim narration alongside its answer (GH #6006). Two independent causes, both fixed at the layer that owns the deliverable. Prompt regression. #4776 told every channel-backed chat to reply with the answer only and not narrate. The MUL-4899 split (#5557) moved that rule into the Slack branch along with the `chat history` / `chat thread` commands its wording happened to mention, so Feishu/Lark silently lost it on 2026-07-17. The rule is a third axis — it keys off "is there a channel at all", like the attachment-upload axis — so it now sits outside the Slack gate, generalized from "these history reads" to any progress note. The old two-layer matrix could not catch the regression because it only asserted the rule on the Slack case; the new test pins all three states. Runtime contract. Result.Output is documented as "final user-facing output selected by the backend" (agent.go), but Codex concatenated every agent_message and Copilot joined every assistant turn with "\n\n", so a tool-using run handed the daemon narration + answer as one string. Codex now takes the message its app-server labels phase="final_answer", falling back to the most recent agent message on the legacy protocol; Copilot keeps the latest complete turn, with the streaming deltas retained as the process-died-mid-turn fallback. This narrows delivery only — every message still streams as MessageText, so the Multica transcript is unchanged. Claude Code, CodeBuddy and qwen already selected a terminal result and are untouched; opencode/deveco/openclaw share the accumulating shape and want the same audit (pi was already fixed this way in #4894). Verified: go build ./..., go vet, full ./pkg/agent and ./internal/daemon/... suites pass locally. Co-authored-by: multica-agent <github@multica.ai> * fix(channel): scope the no-narration rule to process, not results Review follow-ups on the channel delivery rule and the Copilot turn boundary. The prompt said a reply "must not say what you are about to do or just did", which literally forbids the deliverable itself: asked to create an issue, the correct reply IS "created issue X". Rewritten to ban planned and in-progress narration while explicitly protecting the completion confirmation. The test now pins both halves — a future edit that drops the carve-out, or restores the blanket past-tense ban, fails. The example also referenced "check the code", which is not a thing an agent does inside a Slack or Lark conversation. Replaced with a generic "let me look into that first". Copilot cleared pendingDelta only when the authoritative assistant.message carried content. A tool-only turn reports content:"" with the requests as the whole turn, so its streamed deltas stayed buffered and were stitched onto the next turn's partial text if the process then died mid-stream — verified: the new test yields "Checking the logs now.The retry loop" before the fix. The reset now happens on every assistant.message, since that event is the turn boundary regardless of whether it carries text. Verified: go build ./..., go vet, full ./pkg/agent and ./internal/daemon/... suites pass locally. Co-authored-by: multica-agent <github@multica.ai> * fix(channel): tighten the no-narration rule to one sentence Same contract, fewer tokens: the four-line rationale comment collapses to one, and the delivery rule drops the restatement, the second example and the completion examples. What survives is exactly the semantic boundary the tests pin — no planned/in-progress narration, completed actions still count as the outcome — plus the one narration example actually observed in the report. Verified: go build ./..., go vet, ./internal/daemon/... suite pass locally. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
2294f450e8 |
fix(kiro): recover from oversized-history-image session resume failures
Merges MUL-5338 / fixes GH #5975. |
||
|
|
286ed61bb0 |
fix(agent): reap claude process group on cancellation (#5918) [MUL-5288] (#5957)
* fix(agent): reap claude process group on cancellation (#5918) The Claude backend spawned its child with a bare exec.CommandContext, so cancellation SIGKILLed only the leader. On a resumed stream-json session (no wall-clock timeout) the MCP servers and tool subprocesses it spawned were orphaned and kept running — 64+ min in #5918 — while, under --max-concurrent-tasks 1, holding the only slot and starving the queue. Put claude in its own process group and drive a group-wide SIGTERM->grace->SIGKILL on cancel/timeout before closing stdout, mirroring the fix already made for codex (#4520) and opencode (#4533). Add claude_cancel_unix_test.go covering the graceful and SIGKILL-escalation paths. MUL-5288 Co-authored-by: multica-agent <github@multica.ai> * fix(agent): gate claude SIGKILL escalation on whole process group Review found the grace-window escalation keyed off procDone (leader exit), not the process group. A SIGTERM-ignoring descendant that does not hold claude's stdout lets the leader exit, closes procDone, and skips the group SIGKILL — leaking exactly the orphan #5918 targets. Escalate to a group SIGKILL unless waitProcessGroupGone confirms the whole group has exited within the grace window (matching codex). It returns as soon as the group empties, so the graceful path adds no latency. Add a mixed-signal regression: TERM-respecting leader + TERM-ignoring, stdio-detached descendant, which fails against the leader-keyed escalation. MUL-5288 Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
e30776dd9b |
feat(agent): add Claude Opus 5 to the Claude runtime catalog (MUL-5282) (#5910)
Opus 5 is the current flagship in Claude Code's bundled catalog (verified against claude-code 2.1.219: id `claude-opus-5`, display name "Opus 5", pricing tier_5_25, capabilities include xhigh_effort/max_effort). Without a catalog entry the model picker never offered it, and — more damaging — ModelKnownIncompatibleWithProvider treats any unlisted `claude-*` id as a known mismatch, so an agent manually pinned to `claude-opus-5` had the value erased on save. - Add `claude-opus-5` to claudeStaticModels(). Sonnet 4.6 stays the sole badged default; Opus remains a deliberate opt-in. - Allow the full low/medium/high/xhigh/max effort range in claudeModelEffortAllow, matching the rest of the Opus family. - Price it on the standard 5/25 Opus tier in both the server table and the frontend estimator, so Opus 5 usage lands in cost totals instead of the unmapped-model diagnostic. The existing `[1m]` and `<provider>/` tolerances cover the other spellings runtimes report. Verified: go test ./pkg/agent/... ./internal/metrics/..., vitest runtimes/utils.test.ts, tsc --noEmit on packages/views. Co-authored-by: Lambda <lambda@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
2691110867 |
fix(agent): scope Hermes provider-error sniffer to real error boundaries (#5864)
Fixes #5862. The Hermes ACP provider-error sniffer over-matched conversation/tool JSON that Hermes echoes to stderr as `[INFO] root:` records, flipping already-completed runs to failed. Error matching is now scoped to real provider-error boundaries: - Skip INFO/DEBUG root-logger records (single- and multi-line JSON) via a structural state machine, so echoed payloads never participate in matching. - Keep genuine bare provider errors (⚠️/❌/📝 Error:) and non-root [ERROR] records, including after a possibly-truncated INFO record. - Length only bounds the persisted error summary, never whether a line is classified as an error (fixes the earlier #1952-style regression). Reported and initial fix by @YikaJ; remaining boundary fixes added directly by the maintainers to land this quickly as a production bug. |
||
|
|
bce0c05856 |
refactor(agent): drop dead Cursor cost fields (CLI reports no cost) (#5852)
MUL-5240 cursor-agent's stream-json never populates total_cost_usd or a per-step cost — the result event's usage object carries token counts only. Both fields were speculatively copied from Claude Code's schema in the original Cursor runtime PR (#1057) and were never read. Verified against the real CLI (2026.07.20, stream-json and json) plus Cursor's CLI docs. Remove the two dead fields and document that Cursor spend stays estimated from the static rate table (no authoritative per-turn cost to carry, unlike Grok). No behavior change. Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
a92bd5387e |
MUL-5231 fix(agent): parse Cursor top-level thinking and tool_call events (#5846)
Cursor stream-json emits reasoning and tool calls as top-level `thinking` / `tool_call` events; the parser only looked for them inside assistant messages, so transcripts showed a single step. Match the top-level events, taking the tool name from the nested `<name>ToolCall` key and normalizing the packed call_id. Subtypes are matched explicitly — only `started` opens a tool, only `completed` closes it, only `delta` carries reasoning — so an unknown/missing subtype is ignored rather than synthesizing a fake result (which would decrement the daemon in-flight tool count early and misfire the watchdog) or polluting reasoning. Covered by a recorded-stream fixture test, an unknown-subtype regression test, and an opt-in real-CLI smoke test. |
||
|
|
e5a48eb59d |
fix(agent): accept ACP configOptions model catalog (kimi-code 0.29) (#5851)
kimi-code 0.29 dropped the top-level `models.availableModels` / `currentModelId` block from its ACP `session/new` response and moved the same catalog into a `configOptions` entry with `id`/`category` of "model". parseACPSessionNewModels only understood the old shape, so discovery silently returned an empty catalog and the model picker showed "no available models" for an online, correctly-detected kimi runtime. Parse `configOptions` as a fallback: the `models` block still wins when present, so no existing ACP provider changes behaviour. `options[].value` becomes the model id, `options[].name` the label, and `currentValue` marks the default. Non-model options (thinking level) are deliberately skipped — they are a separate product surface and would offer values `session/set_model` cannot honour. Also log a debug line with the top-level response keys (keys only, never values) when session/new succeeds but advertises no catalog, so the next round of upstream schema drift is visible in daemon.log instead of looking like a PATH or install failure. Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
ffa8e16369 |
MUL-5228 fix(usage): bill Grok at xAI's reported cost, fix $0 resumed sessions (#5841)
* fix(agent): attribute Grok usage from the turn's own model id A resumed Grok session with no configured model recorded its entire spend under the model id "unknown", which matches no pricing row — so the task reported $0 cost instead of its real spend. grok.go only learned the model from the session handshake, and ACP's `session/load` carries no model id (only `session/new` does). When neither the agent nor MULTICA_GROK_MODEL pins a model, `daemon.go` legitimately passes an empty model, leaving nothing to attribute the usage to. Every Grok turn stamps `result._meta.modelId` with what it actually billed against. Parse it in the shared ACP result parser and use it as the fallback in grok.go. Other ACP backends are untouched — they keep whatever the handshake gave them. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> * fix(metrics): price the Grok catalog in server-side cost metrics server/internal/metrics/pricing.go carried no Grok rows at all, so RecordLLMUsage took the unpriced branch for every Grok turn: llm_cost_usd reported zero Grok spend while the tokens accumulated in llm_unpriced_tokens. Internal cost monitoring simply could not see Grok. Add the six SKUs xAI publishes rates for, mirroring the frontend table in packages/views/runtimes/utils.ts. Aliases are anchored exact matches like the gpt-5.6 rows, so `grok-composer-*` (in the catalog, absent from the price sheet) stays unmapped instead of inheriting a guessed rate. Short-context tier on purpose: xAI bills a request at 2x once its prompt reaches 200K tokens, but a usage record aggregates every model call in a turn and cannot say which tier an individual request hit. A regression test re-derives the cost of a real grok 0.2.106 turn from the table and checks it against the costUsdTicks xAI returned for that turn. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> * docs(changelog): scope the Grok cost claim to what was actually fixed The v0.4.9 entry promised "accurate cost" in all four languages, but the fix corrected catalog pricing and cached-input double-counting — it did not implement xAI's 2x long-context tier, so a turn whose requests reach 200K prompt tokens still under-reports by up to 50%. Say what was fixed instead. Also correct two stale claims in the pricing comment: the daemon tags usage rows with the runtime provider `grok`, not `xai` (the bare `grok-*` keys are what make them resolve), and record why thresholding the long-context tier on an aggregated row would be worse than not pricing it at all. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> * feat(usage): carry the provider's own cost through to the usage record Cost has always been derived client-side as tokens x a static rate, which cannot express request-level pricing rules. xAI bills a Grok request at 2x once its prompt reaches 200K tokens, and a task_usage row aggregates every model call in a turn — so the stored token counts genuinely cannot say which tier any individual request hit. Thresholding on the aggregate would be worse than the status quo: it turns a bounded 50% under-estimate into an unbounded over-estimate for turns made of many short requests. Grok already reports what it charged, per turn, in `_meta.usage.costUsdTicks`. Parse it, carry it through agent -> daemon -> API, and store it on task_usage as a nullable BIGINT of 1e-10 USD ticks (integer, so sub-cent turns stay exact end to end). NULL means the provider reported no cost — every pre-existing row and every provider that doesn't return one. No backfill: there is no authoritative figure to recover for those, and inventing one is the guess this removes. A single hourly bucket can mix rows that carry a cost with rows that don't, so task_usage_hourly gains both halves: `cost_usd_ticks` sums the authoritative side, and `uncosted_*_tokens` carry exactly the tokens that still need a rate-table estimate. Consumers report authoritative + estimate(uncosted), which degrades to today's behaviour when nothing in the bucket is authoritative. The existing token columns keep covering every row, so token displays are untouched. The new columns are additive with defaults, so the unique key, the dirty-queue shape, and migration 102's triggers are unaffected. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> * feat(usage): prefer the provider's own cost over the rate table With the authoritative figure now stored, both cost consumers use it: the usage dashboard (estimateCost / estimateCostBreakdown) and the server-side llm_cost_usd metric. Each reports `authoritative + estimate(uncosted tokens)`, so a row or bucket that mixes priced and unpriced sources stays whole. The static rate tables remain, but for Grok they are now a fallback — they still price usage recorded by a daemon too old to report cost, and every provider that reports none. Custom pricing overrides likewise apply only to the estimated half: they are a user's guess at a rate, and the authoritative half is not a guess. A model with no rate-table row but a provider-reported cost now also drops out of the "unmapped models" banner, since asking the user to supply a rate for it would invite overriding a real bill. llm_cost_usd is labelled by token_type and the provider reports one number per turn, so the charge is distributed across the buckets in the rate table's own proportions. Only the total is authoritative; the split stays an estimate, which is why this scales the existing buckets rather than inventing a label. estimateCostBreakdown does the same, keeping the stacked chart summing to the headline figure instead of silently under-drawing every Grok row. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> * docs(changelog): say Grok cost now follows xAI's actual charge The earlier wording scoped the claim down to catalog pricing and cached input because the long-context tier was still unhandled. It is handled now — the cost comes from what xAI charged for the turn — so the entry can say so. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> * fix(usage): keep the provider's cost when the model has no rate row Both cost consumers bailed out before reading the authoritative figure when the rate table had no row for the model. A `grok-composer-*` turn — in the Grok Build catalog, absent from xAI's price sheet — was therefore reported as $0 spend even though xAI told us exactly what it charged. Worse on the client: estimateCost returned the real cost while estimateCostBreakdown returned zeros, so the headline and the stacked chart disagreed on precisely the rows whose cost is exact — and the unmapped-models banner was (correctly) hidden, so nothing explained the discrepancy. Handle the charge before the rate lookup in both places. Without rates there is nothing to split a total by, so it lands whole in the `input` bucket, the same fallback distributeAuthoritativeCost already uses when it has no shape to scale. Tokens with no rate keep going to llm_unpriced_tokens: "unpriced" describes the rate table, not the money. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> * perf(usage): drop the historical rewrite from the cost-split migration Migration 213 rewrote every existing task_usage_hourly row to seed the uncosted counters. That is a full-table UPDATE inside a schema migration — lock time, WAL and bloat all scaling with table size — for rows this issue explicitly does not care about. Deleting the UPDATE alone would have zeroed historical cost: with `NOT NULL DEFAULT 0`, an untouched row asserts "nothing here needs estimating", so every pre-split bucket would report $0 until the rollup happened to touch it. Make the uncosted columns nullable with no default instead. NULL means "never recomputed since the split existed", readers COALESCE it to the row's own token total ("estimate all of it"), and the pre-split behaviour is preserved exactly — with nothing to seed, so no rewrite. A bare ADD COLUMN is metadata-only, so this is now fast DDL. Rows heal into the split naturally as the rollup recomputes their buckets. Verified on a fresh database: a legacy-shaped row reads back as its full tokens to estimate, and a group mixing legacy and post-split buckets sums to the authoritative cost plus both rows' estimable tokens. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
ecbdbda09e |
fix(agent): stop double-counting Grok cached input tokens (#5838)
Grok Build reports cachedReadTokens inside inputTokens (totalTokens == input + output on a real 0.2.106 turn, and that turn's costUsdTicks matches xAI's rates only when the cached prefix is billed once). The shared ACP parser persisted both counters raw, so the usage dashboard charged the cached prefix at the full input rate *and* the cache-read rate — ~4x the real spend on a cache-heavy turn. Re-bucket cached reads out of input when totalTokens proves the overlap, the same normalization codex.go already applies. Backends that report mutually-exclusive buckets or omit totalTokens are untouched. Co-authored-by: J <agent@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
6992c58de3 |
MUL-5185: add Codex Fast mode (#5821)
* feat(agents): add Codex fast mode (MUL-5185) Co-authored-by: multica-agent <github@multica.ai> * fix(agents): make Codex Fast override authoritative Co-authored-by: multica-agent <github@multica.ai> * fix(agents): remove Codex Fast config conflicts Co-authored-by: multica-agent <github@multica.ai> * chore: refresh checks after conflict resolution Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
09dce598df |
MUL-5155: KAP-1051: diagnose Codex thread/start timeouts fail-closed (#5759)
* fix(agent/codex): diagnose thread start timeouts (KAP-1051) Co-authored-by: multica-agent <github@multica.ai> * fix(agent/codex): validate thread ID before success lifecycle Co-authored-by: multica-agent <github@multica.ai> * fix(agent/codex): confirm process tree cleanup Co-authored-by: multica-agent <github@multica.ai> * fix(agent/codex): bound Windows pipe cleanup Co-authored-by: multica-agent <github@multica.ai> * ci: run bounded Codex cleanup test on Windows Co-authored-by: multica-agent <github@multica.ai> * fix(agent): reap initialize timeout process groups * fix(agent): redact initialize timeout stderr * fix(agent): redact initialize context failures --------- Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
36533bbc2b | fix(test): prevent agent CLI execution in default tests (#5789) | ||
|
|
5d9295ac65 |
feat(agents): add per-agent runtime skill controls (#5686)
* feat(agents): add per-agent runtime skill controls Co-authored-by: multica-agent <github@multica.ai> * fix(agents): renumber runtime-skill migration and broadcast agent:status on toggle Address the MUL-5101 review blockers on PR #5686: - Rebase onto main and renumber the runtime-skill-disable migration 202 -> 203. main added 202_runtime_profile_add_qwen, so the pair collided on prefix 202 and migrations_lint_test would reject the duplicate. 203 is the next free prefix. - Publish an "agent:status" event after persisting a disabled_runtime_skills override, mirroring the workspace-skill toggle in writeUpdatedAgentSkills. The realtime layer keys off this event to invalidate workspaceKeys.agents, so other open web/desktop/mobile clients now drop their stale toggle state instead of only the initiating tab refreshing. Reload junction-table skills before the broadcast so it doesn't signal cleared skills (#3459). - Add a handler regression test proving the broadcast fires on both disable and enable. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: multica-agent <github@multica.ai> Co-authored-by: Walt <walt@multica.ai> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3171e6607f |
MUL-5141: fix(agent): inject --yolo in Qwen headless runs so shell/edit/write tools are available (#5752)
* fix(agent): inject --yolo in Qwen headless runs so shell/edit/write tools are available (Fixes #5743) Qwen Code's non-interactive mode (`-p … --output-format stream-json`) uses a fail-closed approval policy: it silently drops `run_shell_command`, `edit`, `write_file`, and `monitor` from the tool registry unless bypass mode is active. Every other Multica-supported coding adapter already injects its equivalent permission flag as a daemon-owned argument (e.g. Claude uses `--permission-mode bypassPermissions`, Grok uses `--always-approve`, Qoder uses `--yolo --acp`). Qwen was the only exception. Changes: - `buildQwenArgs`: append `--yolo` after the protocol flags and before any custom args so headless daemon runs always receive the full tool set. - `qwenBlockedArgs`: add `--yolo`, `-y`, `--approval-mode`, and `--allowed-tools` as daemon-owned flags that are stripped from custom_args. This prevents users from accidentally or intentionally disabling bypass mode or narrowing the allowed tool set via per-agent settings. `--exclude-tools` is intentionally left unblocked so users can still hard-deny specific tools. - `TestBuildQwenArgsKeepsProtocolManaged`: extend with the new blocked flags and assert daemon-owned `--yolo` appears exactly once. - `TestBuildQwenArgsYoloAlwaysPresent`: new test asserting `--yolo` is present even when `ExecOptions` carries no custom args. * fix(agent): correct Qwen permission args and docs Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
d804dedcc6 |
MUL-5137: fix(agent): parse Grok ACP token usage from session/prompt _meta (#5748)
* fix(agent): parse Grok ACP token usage from session/prompt _meta Grok Build places per-turn metering under result._meta (and _meta.usage), not the top-level ACP usage field. Multica's shared parser only read result.usage, so Grok tasks recorded empty token usage and cost dashboards stayed at zero. Fall back to _meta.usage (then flat _meta counters) when top-level usage is absent. Prefer standard top-level usage when both are present. Update the Grok fake ACP fixture and add regression tests against the live 0.2.x payload shape. * test(agent): cover zero ACP usage meta fallback Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
6ced54b44a |
fix(agent/codex): retry once when the model catalog refresh blocks the first turn (MUL-5110) (#5734)
Retries once when Codex's model catalog refresh blocks the first turn, with a narrow safety gate: first-turn no-progress timeout, zero semantic progress, catalog-refresh evidence in stderr, and a confirmed-reaped process tree. Clears ResumeSessionID on retry so a stalled thread is never resumed, and buffers the leading session pin so a discarded attempt cannot pin the resume pointer. |
||
|
|
a5a42846e6 |
fix(daemon): retry with a fresh session only when the resume was actually rejected (MUL-4966) (#5715)
* fix(daemon): gate fresh-session retry on tools executed, not session id (MUL-4966) Switching provider accounts leaves the stored session id pointing at a conversation the new account does not own. The daemon still passes it to --resume, the provider rejects it, and the task dies before doing any work. The existing fresh-session fallback was supposed to catch this but was gated on `result.SessionID == ""`, which is not a lifecycle fact: - Too narrow: a backend that echoes the requested id back when it rejects a resume keeps SessionID non-empty, so the fallback never fired — the reported bug. - Too broad: a provider 401 before the first stream message also leaves SessionID empty, so an unrecoverable auth failure burned a second full run. Gate on `tools == 0` instead. That states the property that actually makes a retry safe — the agent executed no tool, so it mutated nothing, so re-running cannot double-post a comment (comment creation has no idempotency key and a duplicate re-fires its @mention triggers), reopen a PR, or re-plan on top of its own half-finished work in the reused workdir. Auth failures are excluded, mirroring retryableReasons in service/task.go. The predicate is extracted to shouldRetryWithFreshSession so the tests exercise production logic; both existing fallback tests re-implemented the condition inline and would not have caught a regression in it. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): gate fresh-session retry on a positive resume-rejected signal (MUL-4966) Review of the previous commit was right: `tools == 0` plus "not an auth error" answers whether re-running is *safe*, not whether a new session can *fix* the failure. Those are orthogonal, and answering the second by exclusion inverts the burden of proof — the failures a fresh session cures are a small enumerable set, while the ones it cannot are open-ended. Concretely, the previous predicate fresh-retried on provider_network, 429/529, quota, 5xx and unclassified startup failures. provider_network is the sharpest conflict: internal/service/task.go marks it resume-safe (MUL-4910) specifically so the platform retry inherits the session and continues the truncated conversation. Resetting the session first made that contract unsatisfiable, silently discarding conversation context on a transient blip — and rate limits got an immediate no-backoff re-run. Replace the inference with positive evidence: agent.Result gains an explicit ResumeRejected field, set only when a backend has proof the resume itself was refused. claude/codebuddy/qwen derive it from resumeWasRejected, which promotes the predicate resolveSessionID was already computing and encoding as the side effect of blanking SessionID — using an empty string to carry that meaning is what made the original bug possible. SessionID keeps being dropped for a rejected resume (a dead pointer must not be persisted), but it is no longer the signal the daemon reads to decide *why* a run failed. The six ACP backends that recover from "session not found" set the flag at the same points they already clear the id, so their existing recovery is not caught by the narrower gate. codex needs nothing: thread/resume already falls back to thread/start in-process, and deliberately does not on transport errors. Matching now includes the account-switch guardrail reported in #5704 (Claude Code 2.1.207, zh-CN): "400 此 session 已绑定另外的ai账号,请执行 /new 开启新 session". The en-US wording of the same guardrail has not been captured yet, so those variants are marked inferred in the source; a miss degrades to a terminal failure carrying the provider's raw text rather than a mis-routed run. Tests: backend-level fixtures drive ResumeRejected from real stream-json for both the account-binding 400 and a network drop, and the predicate now covers network/rate-limit/quota/5xx/auth/unclassified as explicit non-retries. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): restore fresh-session recovery for backends with no rejection signal (MUL-4966) Final review caught qwen regressing: its verified rejection string ("No saved session found with ID ...", already captured in testdata/qwen-code-0.20.0-resume-not-found.stderr.txt) was not in the phrase list, and qwen reports no session id on that path, so the new inclusion gate turned a working auto-recovery into a terminal failure. Auditing the other 17 resume-capable backends showed qwen was not alone. antigravity, copilot, cursor, deveco and opencode all recovered from a refused resume purely by reporting an empty SessionID, and none of them has any rejection detection to convert into ResumeRejected — copilot's own comment documents the hole (session.error before session.start), and antigravity's helper returns "" when "the CLI exited before dispatching". Making ResumeRejected the sole gate silently removed recovery from all five. Fixing that by guessing rejection phrases for five more CLIs is the wrong trade: no real output has been captured for any of them, and a false positive discards a recoverable session pointer. So the gate is now two tiers. Positive evidence (ResumeRejected) decides on its own where a backend can produce it. Where none is available, an empty SessionID still gates the retry — it proves no session was established, which is exactly what the pre-change behaviour relied on — minus the classes a fresh session provably cannot cure (network, rate limit, quota, provider 5xx, auth). That keeps the resume-safe contract in internal/service/task.go intact while restoring what these five backends had. Also renames claudeResumeRejectedPhrases to resumeRejectedPhrases: it is matched by claude, codebuddy and qwen, so a qwen-only string living under a claude-prefixed name would be actively misleading. Tests: qwen's existing missing-resume fixture now asserts ResumeRejected (verified failing without the phrase), and the predicate covers the no-signal tiers — retry when nothing was established, no retry once a session exists or the failure classifies as uncurable. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): scope the no-signal fallback to backends that cannot detect rejections (MUL-4966) Final review caught the compatibility path applying to every backend, not just the five it was justified for. shouldRetryWithFreshSession only saw (Result, priorSessionID, tools), so a false ResumeRejected could not be told apart from a backend that has no way to answer — and claude/codebuddy/qwen/ACP startup failures with no session id still fell through to the exclusion branch. That contradicted both the stated intent and the function's own doc comment ("where a backend can produce it, it is the whole answer"). Make the capability explicit. agent.ResumeRejectionUndetectable names the five backends that scrape SessionID out of stream output and have no rejection detection at all; the daemon takes provider and consults it, so a capable backend reporting false is now taken at its word. Membership is opt-in, so a new backend fails closed instead of silently inheriting a guess-based retry. Also completes the exclusion set: missing config, unavailable model, missing executable, unsupported runtime version and (defensively) agent timeout all have defined non-session remedies and were reaching `default: true`. What is left through stays narrow — unknown, process failure, unparseable output, context overflow — because a real rejection from these five most likely surfaces as a non-zero exit or unparseable output, none of them reporting one explicitly. Tests: one identical result asserted across all five undetectable backends (retries), twelve capable ones (no retry), and an unregistered provider (fails closed), plus table cases for each newly excluded reason. Classifier inputs were verified to map to the intended reasons rather than passing by accident. Also updates the ResumeRejected doc comment, which still said the daemon gates on it alone. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
e5e278a1a2 |
fix(agent/cursor): send prompt on stdin so CLI-like flags cannot be re-tokenised (MUL-4992) (#5711)
* fix(agent/cursor): send prompt on stdin so CLI-like flags cannot be re-tokenised (MUL-4992) On Windows a Cursor task whose prompt contains CLI-like flags failed in ~2s with `error: unknown option '-X'` and no agent output (#5649). A pasted build log such as go build -ldflags "-X main.version=foo" -o bin/server ./cmd/server was enough to kill the run. buildCursorArgs put the whole prompt in argv as the positional after -p. The official Windows launcher chain ends in `& node.exe index.js $args` inside cursor-agent.ps1, where PowerShell re-serialises $args onto node's command line. Under Windows PowerShell 5.1 and pwsh <= 7.2 (Legacy native argument passing) an argument holding embedded double quotes is not re-escaped, so the quoted region closes early, node's argv parser re-splits at the interior spaces, and `-X` reaches commander.js as a standalone flag. #1709 removed the cmd.exe `%*` re-tokenisation but stopped at the Go -> PowerShell boundary, one hop before this. Its Windows tests only compare the argv slice Go builds and never execute a shim, which is why the gap stayed invisible. Fix: keep the prompt off every command line. cursor-agent's -p is a boolean print-mode switch and the prompt is positional; with no positional prompt and a non-TTY stdin the CLI reads stdin to EOF and uses that as the prompt. So drop the prompt from argv on all platforms and write it to stdin, leaving only fixed, content-free flags in argv. No shell or launcher on any platform can re-tokenise what is not on a command line. The write runs in its own goroutine: a prompt larger than the pipe buffer (~64 KiB) blocks mid-write until the child drains it, and the child cannot drain while nothing reads its stdout. Closing stdin signals end-of-prompt, so it is closed on both success and error paths, and on cancellation to release a blocked write. Write failures surface in the result diagnostic, ranked below explicit agent errors so an early child exit (bad auth, bad flag) is not masked by the resulting EPIPE. Tests: a prompt carrying the exact `-ldflags "-X ..."` shape must arrive byte-for-byte on stdin and appear nowhere in argv; a 512 KiB prompt must not deadlock; the prompt is written verbatim (the CLI trims it, we do not). Both new unix tests fail against the pre-fix code. A Windows-tagged test drives a real PowerShell host through the same .cmd -> -File rewrite. Closes #5649 Co-authored-by: multica-agent <github@multica.ai> * test(agent/cursor): prove the Windows launcher fix on both PowerShell hosts in CI The stdin fix has to hold on the host that actually exhibits the bug. powershell.exe (5.1) and pwsh <= 7.2 default to Legacy native argument passing; pwsh >= 7.3 defaults to Standard. A fix verified only on the newer host would not be a fix for the reporter. Run the shim probe against every PowerShell host on PATH rather than only the one defaultPowerShellLookup would select, and hook the windows-tagged launcher tests into the existing windows-execenv CI job. These tests are windows-tagged and the backend job runs on ubuntu, so until now they ran nowhere. Co-authored-by: multica-agent <github@multica.ai> * test(ci): run Windows launcher tests verbosely so skips are visible A skipped or unmatched test still reports "ok", which would make the Windows job look like coverage it is not providing. Co-authored-by: multica-agent <github@multica.ai> * test(agent/cursor): close both coverage gaps in the #5649 regression tests Two tests claimed guarantees they did not actually establish. Windows: the fake cursor-agent.ps1 called [Console]::In.ReadToEnd() and wrote the result itself, so it never launched a native child. The official shim ends in `& node.exe index.js $args`, and that last hop is precisely where the bug lives -- PowerShell re-serialises $args onto the child command line, and whether the child inherits stdin was left unproven. The fake ps1 now re-executes the test binary as a real native child (helper-process idiom), which records the argv it actually received and drains stdin. Both PowerShell hosts still run. Interlock: the large-prompt fake drained stdin before writing any stdout, so the child always unblocked the parent immediately and the test passed even against a synchronous write -- it could not fail for the reason it existed. The fake now floods stdout past pipe capacity *before* reading stdin, creating the real mutual block. Verified: with the writer made synchronous the test deadlocks to its 30s timeout, and passes only with the concurrent writer. Also adds the missing cancellation case: a child that never reads stdin leaves the writer blocked forever, so cancelling the context must close stdin, release the writer and settle the run as aborted. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
fbe00ca164 |
fix(daemon): run Codex unsandboxed on Windows to stop reject-by-policy (MUL-4957) (#5672)
* fix(daemon): run Codex unsandboxed on Windows to stop reject-by-policy (MUL-4957) Windows has no Landlock/Seatbelt-equivalent filesystem sandbox that the daemon configures, so the per-task `sandbox_mode = "workspace-write"` it wrote was unenforceable. Worse than having no sandbox, it pushed Codex into rejecting non-safe mutation commands "by policy": `multica issue create` fails with "was rejected by policy" because Codex can neither sandbox the command nor (under approval_policy = "never") escalate it to the daemon's auto-approver, so the request never reaches the approver. Mirror the existing macOS fallback and give Windows danger-full-access so those commands run. Also generalize the danger-full-access warn log so it no longer hardcodes "on macOS" and only surfaces the macOS-specific upgrade hint on macOS (new codexSandboxPolicy.Hint field). Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): correct Windows sandbox rationale and respect user windows.sandbox (MUL-4957) Addresses two review must-fixes on #5672: 1. Correct a false security fact. The comments and log Reason claimed Windows has no filesystem sandbox backend. Codex 0.144.5 does ship a native Windows sandbox (windows.sandbox = "unelevated"/"elevated"); it is experimental with open upstream reliability bugs, so the daemon defaults to danger-full-access as a deliberate compatibility choice. Enabling the native sandbox is tracked as separate follow-up work. 2. Stop silently downgrading users who opted into isolation. The fallback was unconditional. Add codexSandboxPolicyForConfig: on Windows an explicit windows.sandbox = unelevated|elevated keeps workspace-write so Codex enforces task isolation with the user's chosen backend; danger-full-access applies only when windows.sandbox is absent, disabled, or unparseable. This is also the branch point for a future native-sandbox rollout (flip the default; callers unchanged). Adds fixture tests locking the priority (user opt-in kept vs. unconfigured fallback) plus predicate coverage for codexSandboxPolicyForConfig. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): fail closed on undecidable Windows sandbox config, honor -c windows.sandbox (MUL-4957) Second review round on #5672. Two must-fixes. 1. Undecidable config no longer fails open. The old bool detector collapsed "unparseable / invalid value / failed copy" into "unconfigured" and then loosened to danger-full-access. Replaced with a tri-state (absent/native/undecidable): only exact-lowercase unelevated|elevated (the sole values Codex accepts — verified: any other value makes Codex refuse to load the config) counts as native; any other present value, unparseable TOML, a read error, or a missing per-task config when a shared ~/.codex/config.toml exists (i.e. the copy failed) is undecidable and fails closed to workspace-write — it never loosens — logged at error level. 2. windows.sandbox set via `-c`/`--config` custom args is now honored. Such args never land in config.toml, so config-only detection silently downgraded those users' isolation. The effective Codex args (daemon defaults + profile-fixed + per-agent custom_args) are threaded through PrepareParams/ReuseParams/CodexHomeOptions into the sandbox decision and scanned for a windows.sandbox override (inline, two-token, quoted, spaced; last-wins). Also drops issue-status-bound source comments (openai/codex#24098 has since closed). Adds unit coverage for config/args classification, the fold precedence (undecidable > native > absent), and the copy-failed fail-closed path. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): fail closed on config-sync errors and honor shell-quoted -c windows.sandbox (MUL-4957) Round-3 review must-fixes: 1. resolveWindowsSandboxState now takes the config.toml sync error and a tri-state shared-config presence instead of re-stat-ing inside. A failed sync (stale/absent per-task copy) or an un-stat-able shared source is undecidable and keeps workspace-write, closing the fail-open where a failed sync was read as "unconfigured". Splits IO from the decision so the paths are unit-testable without faulting the filesystem. 2. The Windows sandbox decision consumes agent.NormalizeCodexLaunchArgs (the shared helper buildCodexArgs now uses) so a shell-quoted -c windows.sandbox opt-in is normalized identically to launch, instead of being missed by a raw-token scan and silently downgraded. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): abort when the Codex sandbox block cannot be written (MUL-4957) Round-4 review must-fix: ensureCodexSandboxConfig failures were warn-and-continue, so a computed fail-closed workspace-write policy could stay only in memory while config.toml kept a stale danger-full-access from a prior run — the decision failed closed but the effective config failed open. prepareCodexHomeWithOpts now returns the error, which blocks startup on both paths: fresh Prepare fails the task, and Reuse leaves env.CodexHome unset, which configureCodexTaskShellEnvironment already refuses to start. Regression covers the full reuse scenario (stale danger-full-access + failed config sync + failed managed-block write); it fails with "got nil" without the fix. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> Co-authored-by: J <j@multica.ai> |
||
|
|
f8bf6cd8b9 |
feat(runtime): add Qwen Code runtime (MUL-5015)
Merge approved PR #5666. |
||
|
|
e2f4f28462 |
fix(agent): skip redundant Hermes set_model when session already on the requested model (#5690)
When agent.model equals the model the Hermes ACP session already reports as current, Multica was still replaying session/set_model. Hermes' set_model re-runs provider auto-detection on the model id, which for a provider:model id whose parsed provider matches the session's current provider (e.g. a named custom endpoint exposed as "custom") can mis-route to a different provider — custom:deepseek-v4-pro resolves into the OpenRouter catalog and the turn fails with an auth error. Only the retry via session resume, where the corrupted provider no longer matches the custom: prefix, succeeds. Capture the runtime's reported currentModelId from session/new and session/resume and skip the redundant switch when it already equals the requested model. An empty/unparsable current model falls through and still sends set_model, preserving prior behaviour. The explicit provider:model id is never rewritten. Upstream: NousResearch/hermes-agent#59089. MUL-5029 Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
f51828f568 |
MUL-4936: Fix empty replies in resumed Codex and Hermes chats (#5675)
* fix(agent): isolate resumed Codex turns (MUL-4936) Co-authored-by: multica-agent <github@multica.ai> * fix(agent): drain late Hermes replies (MUL-4936) Co-authored-by: multica-agent <github@multica.ai> * fix(agent): harden resumed chat cleanup (MUL-4936) Co-authored-by: multica-agent <github@multica.ai> * fix(agent): filter Codex subagents before turn gate (MUL-4936) Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
adc7636674 |
MUL-4921: fix(codex): report resumed usage as task delta (#5569)
* fix(codex): report resumed usage as task delta Codex session logs expose cumulative total_token_usage values. Subtract the pre-task baseline so issue usage does not count earlier turns again. Refs #5568 * fix(codex): harden resumed usage accounting * fix(codex): scope fallback usage to current thread Co-authored-by: multica-agent <github@multica.ai> * fix(codex): harden rollout usage edge cases Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
910671b185 |
Fix Codex task MULTICA_TOKEN passthrough (ZIC-82)
Merge approved PR after independent final review; all required CI checks passed. |
||
|
|
ea8ccf3123 |
MUL-4903 fix(agent): harden Cursor terminal failures (#5559)
* fix(agent): harden Cursor terminal failures Co-authored-by: multica-agent <github@multica.ai> * fix(agent): preserve UTF-8 stderr tails Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
6c84664089 | test(agent): serialize thinking cache tests (#5551) | ||
|
|
ed9adc2bbe |
feat: improve create issue field controls (#5532)
Co-authored-by: Lambda <lambda@multica.ai> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
983545b34d |
KAP-867: retry safe Codex initialize timeouts
Merge reviewed changes after fresh approval and green CI. |
||
|
|
cf1fc64671 |
MUL-4824: fail closed on incomplete Claude-style result streams (#5522)
* fix(agent): fail closed on incomplete stream results Co-authored-by: multica-agent <github@multica.ai> * fix(agent): tighten empty result fallback Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
5251a0958d |
fix(daemon): bound silent OpenCode streams (#5523)
Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
f0d7a444ac |
MUL-4860 fix(acp/kiro): GPT-5.6 Sol completion — object rawOutput + positive-proof guard (follow-up to #5511) (#5515)
* fix(kiro): require positive proof for -32603 completion rescue (MUL-4860, #5509) Addresses the #5511 review (Elon): Must-fix 1 — a missing result is not proof of success. A mid-command crash / internal session failure produces the same 'tool use, no result, -32603' shape as a genuine completion, so the previous 'saw the use but never saw a result' rescue could mark an unfinished task completed. The guard now rescues only on POSITIVE proof: a terminal ToolResult with status=='completed' for a finishing tool (goal_complete / comment-add). Must-fix 2 — the previous sticky booleans only ever wrote true, so a later failed delivery could not override an earlier success, and call ordering was ignored. Completion state is now the status of the MOST RECENT finishing-tool result (keyed per CallID): 'progress completed → final failed → -32603' stays failed; 'first failed → retry completed → -32603' is a real completion. Comment-add is still recognized by its command payload regardless of the tool's display title (the #5509 core ask), so a completed comment-add carrying a non-terminal title is preserved through the close handshake. Tests: result-less running tool stays failed (both goal_complete and comment-add), completed non-terminal-title comment-add is preserved, failed-final-overrides-earlier-success, and completed-retry-after-failure. Co-authored-by: multica-agent <github@multica.ai> * fix(acp): don't drop completed tool_call_update when rawOutput is an object (MUL-4860, #5509) Captured a live kiro-cli 2.12.3 + gpt-5.6-sol ACP trace and found the true root cause: the completed tool_call_update sends rawOutput as an OBJECT ({"items":[{"Json":{...}}]}), but handleToolCallUpdate typed rawOutput/output as Go strings. json.Unmarshal then failed with 'cannot unmarshal object into Go struct field .rawOutput of type string' and the handler returned early — silently dropping the ENTIRE update, status:"completed" included. So no completion signal ever reached the kiro completion-preservation guard, and the durable-but-closed task was marked failed. This also explains the issue's 'shell tool recorded as running': the tool_call title is literally 'Running: <cmd>' (kind execute), which normalizes to 'running', not 'terminal'. Fix: type rawOutput/output as json.RawMessage and render them via a new acpRawText helper that accepts both a JSON string and a structured value. This is a shared ACP-layer fix (hermes/kimi/kiro). Combined with the earlier payload-based comment-add recognition and the positive-proof completion guard, the completed result now flows through and the -32603 close handshake correctly preserves completion. Adds TestKiroBackendPreservesCompletionOnRealGPT56SolFrames (the exact captured wire shape end-to-end) and TestACPRawText. Verified the end-to-end test fails (status=failed) when rawOutput is typed as string and passes after the fix. Co-authored-by: multica-agent <github@multica.ai> * fix(kiro): remove duplicate stale-session branch that faked completion (MUL-4860, #5509) Addresses Elon's 3rd-round review of #5515. The -32603 guard had two consecutive, identically-conditioned else-if branches for a stale resumed session (opts.ResumeSessionID != "" && isACPSessionNotFound(err)). The first wrongly forced finalStatus="completed" and kept the stale SessionID; the second — the correct handler that clears SessionID so the daemon's fresh-session retry fires — was unreachable. So a recovery whose resumed session was gone at session/prompt time was reported as a fake success AND skipped the retry. Remove the bogus first branch, keeping the positive-proof completion branch and the original stale-session (SessionID="") branch. Adds TestKiroBackendClearsSessionIDWhenPromptSessionNotFound (the prompt path; existing coverage only exercised session/set_model). Verified it returns completed+ses_stale on the pre-fix code and failed+empty SessionID after. Also gofmt on the touched files. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
dd692058d7 |
fix(kiro): preserve completed status on GPT-5.6 Sol -32603 close handshake (#5509) (#5511)
The -32603 completion-preservation guard coupled to the Claude/Kiro ACP event shape. GPT-5.6 Sol leaves the finishing tool (goal_complete / 'multica issue comment add') parked at 'running' and titles the shell tool with a name that does not normalize to 'terminal', so: - isKiroIssueCommentAddTool's msg.Tool=='terminal' gate dropped the tool use entirely, and - with no completed/failed ToolCallUpdate, no ToolResult is emitted, so the saw*Completed flags never flipped. Completed tasks were therefore reversed to failed with agent_error.provider_server_error ~17s after finishing their work. Fix: - Recognize comment-add by its command payload, independent of the normalized tool title. - Track each finishing tool along use / result / completed axes and preserve completed when the -32603 close handshake follows a tool use that never produced a terminal result. A genuinely failed ToolResult still trips saw*Result and keeps the task failed. Adds regression fixtures for the running-tool path (comment-add and goal_complete) plus a failed-result safety case. Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
6455d390e2 |
fix(agent): fail opencode runs whose stream ends without a terminal signal (#5238)
* fix(agent): fail opencode runs whose stream ends without a terminal signal * fix(agent): treat step_finish(reason "tool-calls") as non-terminal in the EOF guard Review follow-up: step bracketing alone left a false-green window. A tool step closes with step_finish(reason "tool-calls") before its continuation step_start, so a stream that dies in that gap had no open step and still reported "completed". Parse part.reason (live-probed on opencode 1.17.16: tool loop emits reasons "tool-calls" then "stop") and keep the run non-terminal after a "tool-calls" finish until the next step_start. Any other reason — including absent, for older opencode versions that predate the field — stays terminal so healthy runs on old protocols are not mass-failed. Tests: tool-calls finish then EOF now fails; the multi-step happy path uses the real wire shape ending in reason "stop"; a reason-less step_finish (legacy protocol) still completes. * fix(agent): keep stop tool steps pending continuation Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |
||
|
|
66316c2614 |
fix(grok): harden ACP authentication and capabilities (#5440)
Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai> |