234 Commits

Author SHA1 Message Date
Bohan Jiang
cd9b956269 fix(agent): spawn Copilot's native binary on Windows so the prompt survives (MUL-5586) (#6236)
On Windows the daemon passes the full multi-line prompt as `-p <prompt>` but
spawns npm's `copilot.cmd`, which we already rewrite to
`powershell -File copilot.ps1`. Neither launcher can carry that argument:

- `copilot.cmd` forwards with `%*`, which cmd.exe expands by re-tokenising
  the raw command line.
- `copilot.ps1` ends in `& node.exe npm-loader.js $args`, and PowerShell
  re-serialises `$args` onto node's command line. Under Windows PowerShell
  5.1 (and pwsh <= 7.2, which default to Legacy native argument passing)
  embedded double quotes are not re-escaped, so the prompt is re-tokenised.

Copilot then sees several argv tokens where one was intended and refuses the
run with "It looks like your prompt was not quoted, so the extra words were
treated as separate arguments" — the same defect class already fixed for
cursor-agent in #5649, except Copilot has no stdin prompt channel to escape
through, so the prompt must stay on the command line and the launchers have
to go.

Copilot CLI ships a native per-platform binary and `npm-loader.js` does
nothing but `spawnSync` it with argv untouched, so resolve
`copilot-win32-{x64,arm64}\copilot.exe` out of the npm layout and spawn it
directly. That leaves exactly one hop, Go -> native binary, and Go's
syscall.EscapeArg is the exact inverse of the CRT parsing that binary uses.
This mirrors resolveOpenCodeNativeFromShim / resolveDevecoNativeFromShim.

Both the nested (current npm) and hoisted (older npm) platform-package
locations are probed; when neither resolves, we keep falling back to the
PowerShell launcher, which is still better than cmd.exe.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 16:09:36 +08:00
Multica Eve
e4b6f7a31b MUL-5581: add Qoder CN CLI runtime (#6232)
* feat(agent): add Qoder CN CLI runtime

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): address Qoder CN review nits

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): defer Qoder CN version gate

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-31 16:01:53 +08:00
Otis Cui
cee985592e fix(agent): drain trailing ACP notifications in kimi, kiro, qoder, and traecli (#5951) 2026-07-31 15:22:47 +08:00
Bohan Jiang
75c11db048 MUL-5549: feat(agent): discover codebuddy models over ACP instead of scraping --help (#6203)
* feat(agent): discover codebuddy models over ACP instead of scraping --help (MUL-5549)

CodeBuddy speaks ACP, and `session/new` answers with a structured catalog under
models.availableModels plus a currentModelId — exactly the shape the shared
parseACPSessionNewModels already reads for Copilot / Kimi / Kiro / Qoder / Grok /
TRAE. Scraping the `--model` line out of `codebuddy --help` was never necessary.

The help text carried IDs and nothing else, which cost us three things:

- Labels were guessed from the ID and were simply wrong. `kimi-k3-1` rendered as
  "Kimi K3 1" where the CLI says Kimi-K3; `deepseek-v3-2-volc` as
  "Deepseek V3 2 Volc" where the CLI says DeepSeek-V3.2.
- The default model was a "first entry wins" guess rather than the advertised
  currentModelId.
- The effort catalog needed a second regex over the same output.

All three come from the handshake now. The effort catalog rides along in the
same session/new response as the `thought_level` config option, so it costs no
extra process — which also retires the "at most one --help per request"
constraint added in #6196, because --help is no longer run at all.

One trap worth naming: thought_level advertises `enabled` ("On (default)")
alongside the six real levels, but `--effort enabled` is not a valid command
line — the daemon passes the selected level straight to the flag. Advertised
levels are filtered against the flag's accepted set, and a currentValue outside
that set (the default `enabled`) becomes an empty DefaultLevel, which the UI
renders as a generic "Default" instead of a value we cannot pass through.

Two adjacent inaccuracies surfaced while confirming the real level set against
CodeBuddy 2.130.0, both fixed here: the static effort fallback omitted `minimal`
and `max`, and so did the server-side IsKnownThinkingValue gate — so the server
rejected two levels the CLI genuinely accepts.

Discovery keeps its fallback, still marked Fallback so it can never be cached as
authoritative (#6196). That covers the not-logged-in case, which is deliberately
NOT special-cased with an auth step: the catalog came back without calling
authenticate on a logged-in CLI, and inventing an auth branch we cannot exercise
would be speculation.

Removes codebuddyModelRe, parseCodebuddyModels, codebuddyModelLabel,
codebuddyModelProvider, codebuddyEffortRe, parseCodebuddyEffortHelp,
codebuddyEffortSuperset, codebuddyHelpOutput and its 60s help cache.

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): keep codebuddy's vendor grouping after the ACP migration (MUL-5549)

Review nit, and a real regression in the previous commit. Dropping
codebuddyModelProvider looked like removing dead code, but it was the only thing
populating Model.Provider for CodeBuddy — and the picker groups on that field.

acpModelEntry can only recover a vendor from a `vendor:model` id. CodeBuddy's
are bare (`glm-5.2`, `kimi-k3-1`), so every model came back with an empty
Provider, and model-dropdown renders the empty group with no header at all: all
16 models would have collapsed into one unlabelled list where main shows Zhipu /
Kimi / MiniMax / DeepSeek / Hunyuan sections.

Restores the prefix inference as a post-pass over the ACP catalog, exactly the
shape discoverCopilotModels already uses for the same reason.

Verified against the real CLI: all 16 models land in five vendor groups with none
ungrouped. Tests assert the vendor for every id CodeBuddy 2.130.0 advertises plus
the static fallback ids, and that the fallback entries' hardcoded providers agree
with the inference. Removing the post-pass fails them.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 21:40:23 +08:00
Bohan Jiang
44ce16d9b8 MUL-5549: fix(agent): stop reporting a failed model discovery as a real catalog (#6196)
* fix(agent): stop reporting a failed model discovery as a real catalog (MUL-5549)

Selecting the CodeBuddy runtime showed a model list that shares no IDs with
what the CLI actually supports, so every pick was an ID codebuddy rejects
(GH #6180). The list in the report is codebuddyStaticModels() verbatim: the
daemon had fallen back, but nothing downstream could tell.

discoverCodebuddyModels returned (staticModels, nil) on all three failure
paths, and copilot/cursor/grok do the same. A failed discovery therefore
arrived as a successful one, which defeated every guard built to catch it:
the daemon reported status "completed", the picker's discovery_failed hint
only renders on isError, and cacheableModelCatalog — whose own comment says
an empty list means transient failure — waves through a non-empty stand-in
and stores it as last-known-good for the full 24h serve window. One blip got
pinned as the answer for a day.

Discovery now returns a Catalog carrying a Fallback marker, which the daemon
forwards as an additive `fallback` field (older servers ignore it; an older
daemon omitting it keeps the previous behaviour). A fallback catalog is still
rendered — the picker stays populated and manual entry still works — but it
is kept out of both the daemon's 60s discovery cache and the server's catalog
cache. On the server it maps to Keep rather than Drop: a stand-in is no
grounds to evict a real catalog, matching how a `failed` report is treated.

Also stop codebuddyHelpOutput swallowing the exec error. CombinedOutput folds
in stderr, so a codebuddy whose `#!/usr/bin/env node` interpreter is missing
from a GUI-launched daemon's PATH had `env: node: No such file or directory`
parsed as help text — and cached as such for 60s.

Verified against CodeBuddy CLI v2.130.0: the parser itself is fine (16 models
from real --help), so this fixes the reporting of the failure, not the parse.

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): run codebuddy --help at most once per model-list request (MUL-5549)

Review catch on the previous commit. Model discovery and effort discovery both
read `codebuddy --help`, and the effort pass called it independently. That was
free while a failed --help was (wrongly) memoised, but once failures correctly
stopped being cached, the failure path ran the 35s command twice in a single
request — past the server's 60s running timeout, so the request timed out and
the late report was then discarded as stale. The user got nothing, not even the
fallback list the previous commit exists to preserve.

discoverCodebuddyModels now owns the thinking annotation, so the one help
capture feeds both catalogs, and the failure path uses codebuddyFallbackCatalog
to apply the static effort levels without exec'ing at all: whatever broke
--help for the model catalog breaks it for the effort catalog too.

Also strengthen the handler tests. They decoded into a struct declared in the
test rather than calling ReportModelListResult, so a wrong JSON tag or a
mis-wired cache branch would have passed. They now drive the real endpoint with
daemon auth and chi params, covering: a fallback report leaving a previously
discovered catalog intact, an older daemon omitting the field still warming the
cache, and an authoritative empty catalog still dropping the snapshot.

Both fixes are mutation-tested — reverting either makes the new tests fail.

Co-authored-by: multica-agent <github@multica.ai>

* docs(agent): correct codebuddy --help comments after the single-capture refactor (MUL-5549)

Review nit. The comments still described the pre-refactor call graph, where
both discoverCodebuddyModels and codebuddyEffortSuperset called
codebuddyHelpOutput and the cache was what stopped the duplicate run. The
effort parser now takes an already-captured string, and the single-invocation
guarantee is structural rather than cache-dependent — which matters, because a
failed --help is deliberately not cached, so a second caller would re-run the
full 35s timeout.

Also note on codebuddyHelpOutput that it has exactly one caller and why a new
one would reintroduce the bug.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 20:34:01 +08:00
Bohan Jiang
99d8f29bde fix(codex): raise the first-turn no-progress ceiling to 60s (MUL-5542) (#6192)
The first-turn watchdog killed healthy turns. Two independent field
reports measured the first progress event landing just past 30s on
gpt-5.5, and ~39s for a WSL app-server, all inside the window the
watchdog treats as "stuck". Raise the ceiling to 60s.

The window's only job is to fail fast instead of waiting out the 10
minute semantic inactivity backstop, so 60s keeps that value while
clearing the observed evidence with margin.

Add a regression test for the clamp, which had none: the configured
semantic inactivity timeout can only shrink this window, never raise it.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 18:24:45 +08:00
Bohan Jiang
1d2a3499c9 test(agent): stop the cursor fixtures racing their own prompt write (MUL-5536) (#6174)
TestCursorExecuteFailsOnCleanEOFWithoutResult failed on main with
"cursor-agent prompt write failed: write |1: broken pipe" where it
expects "stream ended without terminal result".

The fake cursor-agent exits without reading stdin, so the prompt write
races the child's exit: win and the pipe buffer swallows it, lose and the
read end is gone and the write returns EPIPE. writeErr outranks both
exitErr and the generic no-terminal-result error in cursor.go, so a lost
race replaces the asserted failure with the EPIPE one.

A real cursor-agent reads stdin to EOF, so the fixtures now do too. That
removes the race rather than reordering production error precedence,
which is deliberate. Only the two fixtures whose expected error ranks
below writeErr need the drain.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 16:15:36 +08:00
Multica Eve
999e9f93c7 fix(codex): record file edit payload for patch_apply events (#6158)
* fix(redact): scrub secrets nested inside tool input maps and slices

InputMap only passed top-level string values through Text and documented
non-string values as "preserved as-is". Any secret one level down reached
the database and the WebSocket broadcast untouched:

    flat    -> [REDACTED ...]                       (scrubbed)
    nested  -> [map[diff:token=ghp_... path:a.go]]  (leaked verbatim)

This is a prerequisite for recording structured file-edit payloads. Codex
reports an edit as changes[]{path, diff, content}, and the legacy protocol
reports a deletion as the whole outgoing file — so without this, deleting a
.env would persist its full contents in cleartext.

redactValue now walks the composite shapes json.Unmarshal produces, plus
[]string and map[string]string for argv-style inputs. Composites are copied
rather than scrubbed in place, because the caller keeps using the map it
passed in.

Nesting depth comes from daemon-supplied JSON, so the walk is bounded at 32
levels; a pathologically nested payload would otherwise recurse until the
stack blows. Hitting the bound yields a placeholder rather than the raw
value, keeping the fail-safe direction.

Verified: the five new tests each fail against the previous top-level-only
implementation and pass now; full ./pkg/redact suite green.

Co-authored-by: multica-agent <github@multica.ai>

* fix(codex): record file edit payload for patch_apply events

Both Codex protocol paths recorded a file edit as a bare call ID and set no
payload, so a run that edited six files left six blank, unexpandable rows in
the transcript. The same task on Claude or Grok showed a readable diff, and
the branch Codex pushed was the only surviving record of what it changed
(GH #6157).

The omission was specific to this one tool, not to the adapter: the
exec_command handlers directly above already captured command and output.

Both paths are fixed, since the protocol is sniffed at runtime. Their wire
shapes differ more than they appear, and the normalizer reconciles that:

  - legacy patch_apply_begin/end carry map[path]FileChange, internally tagged
    on `type`, where add/delete hold whole-file `content` and only update
    holds a `unified_diff` plus `move_path`. There is no diff for every case,
    so the normalized form keeps diff and content as alternatives.
  - v2 fileChange items carry an ordered array of {path, kind, diff} where
    `kind` is an object, not a string — reading it as a string silently
    yields "" and loses the add/delete/update distinction.
  - status spellings differ too: legacy is snake_case, v2 is camelCase and
    adds inProgress. Both normalize onto one vocabulary, and a legacy event
    predating `status` falls back to its `success` bool.

Legacy map iteration is sorted by path so a replayed event does not reshuffle
the file list.

Completion events now also produce a non-empty output (status, file count,
and any apply_patch stdout/stderr), because an empty output renders as an
unexpandable blank row just like a missing input.

Anything unrecognised — absent, wrongly typed, or malformed changes — returns
no payload, preserving exactly the previous degradation rather than risking
the transcript.

Total diff/content bytes are bounded at 64 KiB with UTF-8-safe truncation,
recording `truncated` and `original_bytes`; paths and kinds always survive,
since they are what a reviewer needs when the body is gone. The bound is
deliberately scoped to this new payload: other providers stream tool inputs
through unbounded, and clamping them here would silently truncate
transcripts that render correctly today. Unifying the limit at the
persistence boundary is left as a follow-up.

Verified: the new tests reproduce the reported symptom (Input:map[],
Output:"") against the previous call sites and pass now; ./pkg/agent and
./pkg/redact green, go vet and gofmt clean.

Co-authored-by: multica-agent <github@multica.ai>

* feat(transcript): render Codex multi-file patch payloads as diffs

The presenter identified an edit by input shape — a top-level file_path plus
old_string/new_string or content — which is Claude's and Grok's shape. Codex
records one patch_apply covering several files as changes[], so even with the
payload now populated it fell through to pretty JSON instead of a diff.

A new `patch` detail kind carries one entry per file, since collapsing them
into a single body would lose which change belongs where. Each file reuses the
existing single-file surfaces, so all bodies behave alike inside the
virtualized list.

Codex hands over a ready-made unified diff, so parseUnifiedDiff maps it onto
diff rows rather than recomputing one — there is no before/after pair to
compare, and reconstructing both sides from the diff just to diff them again
would be circular. Hunk headers become `gap` rows, which is what they denote:
skipped unchanged content.

A deletion renders as all-removals rather than as a whole-file write, because
the legacy protocol reports it as the outgoing file's content and a green
"+N" gutter would state the opposite of what happened.

The collapsed row needed its own fix: with no single path field, the summary
fell through the preference chain and came back empty. It now reads as the
first path plus "+N more".

Anything that is not this shape still falls back to pretty JSON, so a payload
this presenter does not understand stays readable.

Verified: 17 new tests (43 in the presenter suite) pass; repo typecheck and
lint clean. The one failing views test, layout/sidebar-resize, fails
identically on an untouched checkout.

Co-authored-by: multica-agent <github@multica.ai>

* fix(codex): route v2 add/delete payloads as content, not diff

Addresses review on #6158.

Upstream's format_file_change_diff only produces a unified diff for `update`.
For `add` and `delete` it returns the whole file's contents under the same
`diff` field, and for a moved `update` it appends a trailing
"\n\nMoved to: <path>" line:

    FileChange::Add    { content } => content.clone(),
    FileChange::Delete { content } => content.clone(),
    FileChange::Update { unified_diff, move_path } => ...

(codex-rs/app-server-protocol/src/protocol/item_builders.rs, rust-v0.145.0)

Recording that as a diff mislabels every line of an added or deleted file as
context, and actively inverts any line whose content begins with '+' or '-' —
so an added file containing "-minus lead" rendered as a deletion. The payload
is now routed by `kind` rather than by field name, and the "Moved to:"
sentence is stripped since move_path already carries the destination.

The previous v2 tests hid this by using a fixture the real protocol never
emits (an `add` carrying "@@ ... +package main"). They now use upstream's
shape, plus cases for delete, an empty add, and an add whose contents look
like diff headers.

Empty bodies are also kept on both paths: presence of the field, not its
non-emptiness, decides whether a body was reported, so an empty added file
renders as an empty body instead of "no content reported".

Verified: the new assertions fail against the previous normalizer — where an
`add` came through as {"diff": "package main\n"} — and pass now.

Co-authored-by: multica-agent <github@multica.ai>

* fix(transcript): stop treating header-like content lines as file headers

Addresses review on #6158.

parseUnifiedDiff matched "---" / "+++" / "diff --git" / "index " at any
position, so a changed line whose *content* starts with a dash or plus was
silently discarded:

    parseUnifiedDiff("@@ -1 +1 @@\n--- old markdown\n+++ new markdown\n")
    // before: [{ kind: "gap", ... }] — both changed lines gone

A removal of "-- old markdown" is spelled "--- old markdown" on the wire, so
this hit Markdown rules, embedded patches, and comment banners.

File headers only exist ahead of the first hunk, so they are only recognised
there; once inside a hunk every line is parsed strictly by its first
character.

Also localizes the multi-file summary count, which was hardcoded English and
so leaked into the zh-Hans / ja / ko transcript rows. The presenter owns no
React and no i18n by design, so the phrasing is injected by the caller rather
than imported here, keeping the module unit-testable in isolation; the English
form remains the fallback. The three Chinese/Japanese/Korean truncation
strings now use "..." to match the English source they translate.

Verified: both new parser assertions fail against the previous
strip-anywhere behaviour and pass now; 47 presenter tests green, repo
typecheck and lint clean.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): redact nested tool input before it leaves the daemon

Addresses review on #6158.

Recursive redaction ran only in the server's ingest handler. The daemon built
the new nested edit payload and sent msg.Input verbatim, so a daemon that
self-updated ahead of the server — or one talking to a server mid-rollout —
would ship whole-file edit contents to a peer that does not scrub nested
values yet. The legacy protocol reports a deletion as the whole outgoing file,
so that window covered a deleted .env in cleartext.

Ordering three commits inside one PR is not a deployment barrier, and daemon
and server upgrade independently. Deployment order is not a control we have,
so the sending side is now safe on its own; the server keeps redacting on
ingest as the second line of defence.

Scoped to Input, which is the field this PR newly fills with file contents.
Content and Output are plain strings already redacted server-side, and
changing their daemon-side handling would be unrelated to this fix.

Verified: the new daemon test asserts the nested token is masked in the
reported batch while the change metadata survives. It fails without this
change, reporting the full GITHUB_TOKEN= line on the wire, and passes with
it; ./internal/daemon green.

Co-authored-by: multica-agent <github@multica.ai>

* fix(transcript): correct the Chinese multi-file patch count semantics

Addresses review on #6158.

The summary is handed the number of files *beyond* the named one, but the
Chinese phrasing stated a total: "a.go 等 2 个文件" reads as two files including
a.go, so a three-file patch under-reported by one. English hides the
distinction ("+2 more"), which is why it survived the first pass.

Rewords zh-Hans to "另有 N 个文件". Japanese (他) and Korean (외) already read
as "besides", so their wording is unchanged.

Also renames the interpolation variable from `count` to `extra`, for two
reasons. i18next treats `count` as the plural selector — this very namespace
relies on that for events_one/events_other — so a plain number had no business
borrowing it. And the name is what a translator reads: `extra` cannot be
mistaken for a total the way `count` was.

Guards the whole bug class rather than just this string: a locale test asserts
every locale interpolates {{path}} and {{extra}} and never the reserved
{{count}}, and a presenter test pins that the injected number is the count of
additional files, not the total.

Verified: both new locale assertions fail against the reverted string and pass
now; rendering the real locale strings for a three-file patch yields "+2 more",
"另有 2 个文件", "他 2 件", "외 2개". 53 target tests pass, repo typecheck clean,
views lint back to its pre-existing 16 warnings and 0 errors.

Co-authored-by: multica-agent <github@multica.ai>

* fix(transcript): put the patch surface on the type scale

Addresses review on #6158.

The patch surface wrote text-[10px] / text-[11px] / text-[10px], copied from
the sibling transcript surfaces as they looked when this branch started. Since
then MUL-5451 (#6136) introduced a role-named type scale and migrated those
same siblings to text-micro, so these three call sites were the only remaining
arbitrary sizes — and the type-scale guard reports them precisely.

All three become text-micro. That matches the analogues they were copied from
now that those have moved: the FileWriteSurface line-count row, the
DiffDetailSurface header row, and the ToolDetailSurface body. It is also the
only correct target, since micro (11px) is the smallest step the scale defines
— there is nothing at 10px to map to.

Merges origin/main so the guard runs here rather than only in CI.

Verified: apps/web app/type-scale.test.ts 13/13 (it listed exactly these three
lines before), no `text-[` left in the file, repo typecheck clean, views lint
unchanged at 16 pre-existing warnings and 0 errors.

Co-authored-by: multica-agent <github@multica.ai>

* fix(codex): redact patch bodies before applying the size budget

Addresses review on #6158.

The adapter sized and truncated the normalized changes, and redaction only ran
later — in the daemon before sending, then again in the server on ingest. That
order loses secrets that straddle the budget.

The PEM rule needs both markers to match:

    -----BEGIN[A-Z\s]*PRIVATE KEY-----.*?-----END[A-Z\s]*PRIVATE KEY-----

So a 70 KB private key whose BEGIN sits inside the first 64 KiB and whose END
falls past the cut stops matching once truncated. Neither later pass can
recognise what truncation already broke, so the marker and 64 KiB of key
material reach the database and the WebSocket broadcast. Measured on the
previous code:

    stored bytes 65536 | BEGIN marker present | key body present | placeholder absent

Redaction now runs first, and the budget measures the redacted bodies — which
is also the honest measurement, since those are what actually gets stored and
redaction usually shrinks them (that key collapses to 23 bytes, so no trimming
is needed at all). `original_bytes` still reports the pre-redaction size so the
reader sees how large the real patch was. The daemon and server passes stay as
defence in depth; redaction is idempotent, so running three times is safe and
that is now asserted.

Note for callers: codexPatchInput no longer trims its argument in place, because
redaction copies first. Two existing tests were asserting on the caller's
original slice and had silently become vacuous; they now read the returned
payload, and one pins the no-mutation contract. The delete fixture in the
diff-vs-content routing test was also a credential-shaped string, which now
redacts — it is plain text so that test keeps testing routing.

Verified: the new boundary test fails on the previous order, reporting the
surviving BEGIN marker and key material, and passes now. go test ./pkg/agent
./pkg/redact ./internal/daemon green; execenv ByteIdentical green; go vet and
gofmt clean.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 16:01:10 +08:00
Bohan Jiang
d6bd4cf7d5 fix(agent): recover from Hermes resumed sessions it refuses without running (MUL-5509) (#6164)
* fix(agent): recover from Hermes resumed sessions it refuses without running

Hermes never reports an unknown ACP session as a JSON-RPC error. Its adapter
answers session/prompt for a session it cannot load with an ordinary success
frame carrying stopReason=refusal, and session/resume echoes nothing back --
ACP's ResumeSessionResponse has no sessionId field at all, unlike
NewSessionResponse -- so resolveResumedSessionID keeps the id we asked for.
Nothing in the exchange is an error, so the isACPSessionNotFound branches at
set_model and prompt time never fire for this backend.

The result was a Result carrying the dead session id with ResumeRejected=false,
which shouldRetryWithFreshSession reads as "not a rejection". GetLastTaskSession
then handed the same dead id to every later dispatch on that (agent, issue)
pair, so a Hermes agent was usable exactly once per issue: the first task
worked, and every following comment, approval or @mention failed the same way
until a manual rerun bought one more turn. When no provider error reached
stderr the turn was worse than a failure -- it reported completed with empty
output, an agent that silently did nothing.

Treat a refusal on a resumed session with zero agent activity as the runtime
telling us the session is gone: clear the session id, set ResumeRejected so the
existing fresh-session retry and session-retirement path take over, and fail
the turn instead of reporting an empty success. Both conditions are required --
stopReason=refusal alone is a legitimate model refusal, and a refusal after
real work is not a lost session.

The reason is applied after promoteACPResultOnProviderError so a captured
provider error stays the user-visible message; the generic fallback only fills
in when nothing more specific was seen.

Verified against the real hermes CLI (upstream main ba7d214b6) on an isolated
HERMES_HOME with a local mock endpoint: a fresh task completes, and resuming it
now yields status=failed, SessionID="", ResumeRejected=true with the provider
error preserved -- previously the dead id came back with ResumeRejected=false.

Refs GH #6150, MUL-5509

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): decide Hermes resume loss after the pipe drain, not at quiescence

The notification quiet window closing does not end a turn — stdin EOF and the
pipe drain do, and Hermes legitimately delivers a turn's final chunk in that
gap (TestHermesBackendDrainsLateFinalNotificationAfterPromptResponse exists
because of it). Freezing the resume-loss decision at the quiescence boundary
therefore read turnActivity == 0 while a real answer was still in flight: a
runtime that merely answered slowly after stopReason=refusal had its healthy
session id cleared and ResumeRejected set, discarding a live conversation
pointer and, with no tool use recorded, triggering an unnecessary fresh-session
retry that re-ran the turn.

Move the evaluation to the point where the turn has settled — after
waitForHermesPipeDrain and streamingCurrentTurn.Store(false), where every
accepted update has been counted. This also drops the resumeLost variable and
the ordering hazard that came with it.

The new regression sends the late chunk 600ms after the refusal, past the 250ms
quiet window and inside the 2s drain grace, and asserts the session id survives
with ResumeRejected=false. Confirmed to fail against the previous ordering with
exactly the reported symptoms (session id emptied, ResumeRejected=true).

Refs GH #6150, MUL-5509

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-30 15:23:36 +08:00
Bohan Jiang
280fa28e0d fix(agent): deny CodeBuddy's interactive plan-mode tools in headless runs (MUL-5383) (#6104)
CodeBuddy exempts AskUserQuestion and ExitPlanMode from permission-mode
finalization, so `--permission-mode bypassPermissions` never auto-approves
them. Once the model entered plan mode the session mode is Plan, not
BypassPermissions, so ExitPlanMode also missed the daemon's bypass
fast-path and went to the SDK permission bridge — which waits with no
timeout for a confirmation the headless runtime cannot render. The task
sat in-flight until the 2h tool watchdog, and users killed it by hand.

Deny EnterPlanMode/ExitPlanMode alongside the AskUserQuestion we already
deny. Each tool is passed as its own argv value because CodeBuddy matches
disallowedTools entries exactly and does not split on commas.

Also send `allowed: true` on control_response: CodeBuddy's
SdkPermissionClient reads `allowed`, not Claude Code's `behavior`, so the
daemon's "auto-approve" was being read as a denial.

Fixes #6012

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 15:55:02 +08:00
Multica Eve
de3e0ae556 fix(agent): make cursor stream protocol drift loud instead of silent (MUL-5434) (#6089)
* fix(agent): make cursor stream protocol drift loud instead of silent (MUL-5434)

#6071 reports Cursor tasks that demonstrably read files, ran commands and
called the Multica CLI, yet showed a single blob of agent text with no
reasoning and no tool rows. The run still reported success and tools=0.

`switch evt.Type` in cursor.go had no `default` branch, so any top-level
event type we do not handle was dropped with no counter and no warning.
Renaming only the top-level types of a healthy stream
(`thinking`->`reasoning`, `tool_call`->`tool_calls`), leaving every nested
field untouched, reproduces the report exactly: status=completed, output =
the result text alone, tool_use=0, thinking=0, zero diagnostics. The
existing unknown-subtype warning cannot catch this — it only increments
once the type has already matched — so "no unknown-subtype warning" does
not rule out protocol drift.

Two diagnostic gaps are closed:

- Add the `default` branch with a bounded, content-free tally of unhandled
  top-level types, reported once per run as a warning and alongside
  tool_use_count in the protocol summary. Type names are normalized through
  observedCursorEventType and distinct names are capped at 16 plus an
  overflow bucket, so a hostile or noisy stream cannot grow the map or leak
  payload into logs. `user` (the CLI echoing our prompt, present in every
  recorded run) is explicitly benign so the warning does not fire always.
- Count assistant text separately from the terminal result text. The result
  event writes into the same builder, so last_assistant_bytes equalled
  result_bytes even when the assistant streamed nothing — erasing the
  signal that says "only the final answer arrived". This also aligns cursor
  with claude, which already reports assistant bytes only.

Unrecognized events are still never coerced into tool or reasoning
messages; guessing at upstream additions is the failure mode MUL-5231
already fixed once. This is diagnosis only and does not itself restore the
missing tool rows — identifying which upstream shape changed requires a
captured 2026.07.23 stream, and this warning is what makes that
identifiable from a single production log line.

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): report cursor unhandled events as evidence, not as a verdict

Addresses the review's must-fix on MUL-5434. The diagnostic was described
as deciding WHY a transcript is empty, which it cannot do:

- "tools=0 with unhandled types" does not establish that a rename ate the
  tool rows. Cursor 2026.07.23 also emits transport/control frames, so an
  unhandled type only proves the stream carried events we do not parse.
- "tools=0 with no unhandled types" does not establish the agent used no
  tools. The CLI may execute tools without handing the updates to its
  stream serializer at all — the main branch #6071 has NOT ruled out — a
  new shape may be nested inside an event type we already recognize, or
  events may be lost to invalid framing or a scanner boundary.

Changes, no behaviour change to messages, status or output:

- Reword the tally doc, the switch default, the warning and the shared
  observation struct to state what a non-zero and a zero count each do and
  do not establish, and point at the branches that stay open.
- Rename unknown* to unhandled* (fields, log keys, warning text, the
  pre-existing subtype counter) so the diagnostic never implies the type is
  unrecognized upstream — only that this parser does not handle it.
- Classify `connection` and `retry` explicitly. They join `user` in
  cursorNonTranscriptEventTypes, with per-entry provenance recorded: `user`
  is confirmed in the recorded 2026.07.20 stream, the control frames are
  reported on newer builds and listed defensively. A build that does not
  emit them makes the entry inert; one that does must not have a known
  control frame reported as an unhandled protocol event.

Tests assert signal presence and that nothing is fabricated, not causality:
the healthy-stream test now carries `connection` / `retry` and additionally
asserts the real thinking/tool rows still arrive, so suppressing the warning
cannot silently suppress the transcript. TestCursorNonTranscriptEventType
also pins that no type the parser handles can enter the suppression list.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-29 15:15:41 +08:00
David Zhang
b3847b4172 MUL-5433: fix(agent/hermes): pin session id mid-flight for daemon resume (#6070)
* fix(agent/hermes): pin session id mid-flight for daemon resume

Hermes was the only ACP-backed agent that created a session without
emitting a running status carrying the session id. When the daemon
restarted or a task was cancelled mid-flight, PinTaskSession had nothing
to key on, so the resume pointer was lost and the task could not resume.

Emit MessageStatus+SessionID immediately after session create, matching
claude, codebuddy, codex, grok and qwen. The daemon already listens for
this (daemon.go MessageStatus -> PinTaskSession), so this small pin is
all Hermes needs.

See #4969 for the companion daemon/SQL cancel-salvage work.

Co-authored-by: multica-agent <github@multica.ai>

* test(agent): cover Hermes mid-flight session pin

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Eve <eve@multica-ai.local>
2026-07-29 14:58:01 +08:00
Bohan Jiang
0066ab259e refactor(agent): drop unreachable inline system-prompt branches (MUL-5392) (#6050)
* refactor(agent): drop unreachable inline system-prompt branches (MUL-5392)

The daemon only populates ExecOptions.SystemPrompt for openclaw, kimi and
traecli (providerNeedsInlineSystemPrompt); every other backend receives the
runtime brief as a per-task context file in the workdir. The inline branches in
claude, codex, opencode and pi were therefore dead, and read as if they were the
live delivery path.

Probed each backend over its real launch path with a canary in the context file
and no inline delivery — claude 2.1.220 (CLAUDE.md), codex 0.144.6 via the
app-server, opencode 1.17.7, pi 0.67.2, hermes 0.18.2 via ACP — and all of them
picked the brief up from disk, so the branches were removable with no behaviour
change. An empty workdir returned no canary, confirming the probe could fail.

opencode's branch was worse than dead: `opencode run` has no --prompt flag, so
enabling inline delivery there would have made every opencode task exit 1 with a
usage dump. The DevEco backend, forked from opencode, already documents this
constraint; opencode itself never got the fix.

Regression tests pin all three arg builders against re-adding the flag, and
providerNeedsInlineSystemPrompt now documents what was verified and what is
still unprobed (grok, qoder, codebuddy).

Hermes and kiro are untouched: their exclusion is deliberate and already tested.

Co-authored-by: multica-agent <github@multica.ai>

* test(agent): pin codex developerInstructions contract, drop stale pi flag doc

Review follow-up on MUL-5392.

buildPiArgs' doc comment still advertised --append-system-prompt after the
branch that emitted it was removed — exactly the stale-signal this PR set out
to delete.

The two codex sites fixed to a literal nil had no regression test, so restoring
nilIfEmpty(opts.SystemPrompt) would still have gone green. Both thread/start and
thread/resume now run with a canary SystemPrompt and assert developerInstructions
comes through as an explicit null. Mutation-checked: reverting either site fails
its test.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-28 22:02:55 +08:00
Bohan Jiang
962d376fda refactor(agent): share acpDeliverableTracker across ACP backends (MUL-5405) (#6044)
#6022 stopped qoder from delivering interim narration as Result.Output, but
hermes / kimi / kiro / traecli / grok still accumulate every MessageText into
one builder and hand the whole thing to Result.Output — the same leak (#6006),
since Result.Output becomes the channel reply and the auto-generated issue
comment.

Extract the boundary rule into acpDeliverableTracker (observe / result) and
share it across all six backends instead of copying qoder's block five times:
Result.Output keeps only the text after the latest tool call, a turn that ends
on a tool call falls back to the latest non-empty text block so the reply is
never empty, and provider-error detection keeps reading the full text stream.

Covered by tracker unit tests and a cross-backend regression test that pins
both scenarios on all six backends, over both tool-use emission paths
(emitted at the tool call and deferred to tool completion).

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-28 21:18:12 +08:00
YYClaw
54469766fd fix(qoder): deliver final answer only (MUL-5394) (#6022)
* fix(qoder): deliver final answer only

* fix(qoder): preserve tool-terminated replies
2026-07-28 17:28:22 +08:00
Bohan Jiang
ba108978ac fix(channel): deliver the final answer only to Slack and Lark (MUL-5378) (#6016)
* fix(channel): deliver the final answer only to Slack and Lark (MUL-5378)

Channel replies could carry the agent's interim narration alongside its
answer (GH #6006). Two independent causes, both fixed at the layer that
owns the deliverable.

Prompt regression. #4776 told every channel-backed chat to reply with the
answer only and not narrate. The MUL-4899 split (#5557) moved that rule
into the Slack branch along with the `chat history` / `chat thread`
commands its wording happened to mention, so Feishu/Lark silently lost it
on 2026-07-17. The rule is a third axis — it keys off "is there a channel
at all", like the attachment-upload axis — so it now sits outside the
Slack gate, generalized from "these history reads" to any progress note.
The old two-layer matrix could not catch the regression because it only
asserted the rule on the Slack case; the new test pins all three states.

Runtime contract. Result.Output is documented as "final user-facing output
selected by the backend" (agent.go), but Codex concatenated every
agent_message and Copilot joined every assistant turn with "\n\n", so a
tool-using run handed the daemon narration + answer as one string. Codex
now takes the message its app-server labels phase="final_answer", falling
back to the most recent agent message on the legacy protocol; Copilot
keeps the latest complete turn, with the streaming deltas retained as the
process-died-mid-turn fallback. This narrows delivery only — every message
still streams as MessageText, so the Multica transcript is unchanged.

Claude Code, CodeBuddy and qwen already selected a terminal result and are
untouched; opencode/deveco/openclaw share the accumulating shape and want
the same audit (pi was already fixed this way in #4894).

Verified: go build ./..., go vet, full ./pkg/agent and ./internal/daemon/...
suites pass locally.

Co-authored-by: multica-agent <github@multica.ai>

* fix(channel): scope the no-narration rule to process, not results

Review follow-ups on the channel delivery rule and the Copilot turn
boundary.

The prompt said a reply "must not say what you are about to do or just
did", which literally forbids the deliverable itself: asked to create an
issue, the correct reply IS "created issue X". Rewritten to ban planned
and in-progress narration while explicitly protecting the completion
confirmation. The test now pins both halves — a future edit that drops
the carve-out, or restores the blanket past-tense ban, fails.

The example also referenced "check the code", which is not a thing an
agent does inside a Slack or Lark conversation. Replaced with a generic
"let me look into that first".

Copilot cleared pendingDelta only when the authoritative assistant.message
carried content. A tool-only turn reports content:"" with the requests as
the whole turn, so its streamed deltas stayed buffered and were stitched
onto the next turn's partial text if the process then died mid-stream —
verified: the new test yields "Checking the logs now.The retry loop"
before the fix. The reset now happens on every assistant.message, since
that event is the turn boundary regardless of whether it carries text.

Verified: go build ./..., go vet, full ./pkg/agent and ./internal/daemon/...
suites pass locally.

Co-authored-by: multica-agent <github@multica.ai>

* fix(channel): tighten the no-narration rule to one sentence

Same contract, fewer tokens: the four-line rationale comment collapses to
one, and the delivery rule drops the restatement, the second example and
the completion examples. What survives is exactly the semantic boundary
the tests pin — no planned/in-progress narration, completed actions still
count as the outcome — plus the one narration example actually observed in
the report.

Verified: go build ./..., go vet, ./internal/daemon/... suite pass locally.
Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-28 14:28:38 +08:00
Multica Eve
2294f450e8 fix(kiro): recover from oversized-history-image session resume failures
Merges MUL-5338 / fixes GH #5975.
2026-07-27 13:53:58 +08:00
Bohan Jiang
286ed61bb0 fix(agent): reap claude process group on cancellation (#5918) [MUL-5288] (#5957)
* fix(agent): reap claude process group on cancellation (#5918)

The Claude backend spawned its child with a bare exec.CommandContext, so
cancellation SIGKILLed only the leader. On a resumed stream-json session
(no wall-clock timeout) the MCP servers and tool subprocesses it spawned
were orphaned and kept running — 64+ min in #5918 — while, under
--max-concurrent-tasks 1, holding the only slot and starving the queue.

Put claude in its own process group and drive a group-wide
SIGTERM->grace->SIGKILL on cancel/timeout before closing stdout, mirroring
the fix already made for codex (#4520) and opencode (#4533). Add
claude_cancel_unix_test.go covering the graceful and SIGKILL-escalation
paths.

MUL-5288

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): gate claude SIGKILL escalation on whole process group

Review found the grace-window escalation keyed off procDone (leader exit),
not the process group. A SIGTERM-ignoring descendant that does not hold
claude's stdout lets the leader exit, closes procDone, and skips the group
SIGKILL — leaking exactly the orphan #5918 targets.

Escalate to a group SIGKILL unless waitProcessGroupGone confirms the whole
group has exited within the grace window (matching codex). It returns as
soon as the group empties, so the graceful path adds no latency. Add a
mixed-signal regression: TERM-respecting leader + TERM-ignoring,
stdio-detached descendant, which fails against the leader-keyed escalation.

MUL-5288

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-27 02:01:41 +08:00
Jiayuan Zhang
e30776dd9b feat(agent): add Claude Opus 5 to the Claude runtime catalog (MUL-5282) (#5910)
Opus 5 is the current flagship in Claude Code's bundled catalog (verified
against claude-code 2.1.219: id `claude-opus-5`, display name "Opus 5",
pricing tier_5_25, capabilities include xhigh_effort/max_effort). Without a
catalog entry the model picker never offered it, and — more damaging —
ModelKnownIncompatibleWithProvider treats any unlisted `claude-*` id as a
known mismatch, so an agent manually pinned to `claude-opus-5` had the value
erased on save.

- Add `claude-opus-5` to claudeStaticModels(). Sonnet 4.6 stays the sole
  badged default; Opus remains a deliberate opt-in.
- Allow the full low/medium/high/xhigh/max effort range in
  claudeModelEffortAllow, matching the rest of the Opus family.
- Price it on the standard 5/25 Opus tier in both the server table and the
  frontend estimator, so Opus 5 usage lands in cost totals instead of the
  unmapped-model diagnostic. The existing `[1m]` and `<provider>/` tolerances
  cover the other spellings runtimes report.

Verified: go test ./pkg/agent/... ./internal/metrics/..., vitest
runtimes/utils.test.ts, tsc --noEmit on packages/views.

Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-25 02:48:03 +08:00
YikaJ
2691110867 fix(agent): scope Hermes provider-error sniffer to real error boundaries (#5864)
Fixes #5862.

The Hermes ACP provider-error sniffer over-matched conversation/tool JSON that
Hermes echoes to stderr as `[INFO] root:` records, flipping already-completed
runs to failed. Error matching is now scoped to real provider-error boundaries:

- Skip INFO/DEBUG root-logger records (single- and multi-line JSON) via a
  structural state machine, so echoed payloads never participate in matching.
- Keep genuine bare provider errors (⚠️//📝 Error:) and non-root [ERROR]
  records, including after a possibly-truncated INFO record.
- Length only bounds the persisted error summary, never whether a line is
  classified as an error (fixes the earlier #1952-style regression).

Reported and initial fix by @YikaJ; remaining boundary fixes added directly by
the maintainers to land this quickly as a production bug.
2026-07-24 16:31:00 +08:00
Bohan Jiang
bce0c05856 refactor(agent): drop dead Cursor cost fields (CLI reports no cost) (#5852)
MUL-5240

cursor-agent's stream-json never populates total_cost_usd or a per-step
cost — the result event's usage object carries token counts only. Both
fields were speculatively copied from Claude Code's schema in the original
Cursor runtime PR (#1057) and were never read. Verified against the real
CLI (2026.07.20, stream-json and json) plus Cursor's CLI docs.

Remove the two dead fields and document that Cursor spend stays estimated
from the static rate table (no authoritative per-turn cost to carry, unlike
Grok). No behavior change.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-24 01:55:07 +08:00
Bohan Jiang
a92bd5387e MUL-5231 fix(agent): parse Cursor top-level thinking and tool_call events (#5846)
Cursor stream-json emits reasoning and tool calls as top-level `thinking` /
`tool_call` events; the parser only looked for them inside assistant messages,
so transcripts showed a single step. Match the top-level events, taking the
tool name from the nested `<name>ToolCall` key and normalizing the packed
call_id. Subtypes are matched explicitly — only `started` opens a tool, only
`completed` closes it, only `delta` carries reasoning — so an unknown/missing
subtype is ignored rather than synthesizing a fake result (which would decrement
the daemon in-flight tool count early and misfire the watchdog) or polluting
reasoning. Covered by a recorded-stream fixture test, an unknown-subtype
regression test, and an opt-in real-CLI smoke test.
2026-07-24 01:52:15 +08:00
Bohan Jiang
e5a48eb59d fix(agent): accept ACP configOptions model catalog (kimi-code 0.29) (#5851)
kimi-code 0.29 dropped the top-level `models.availableModels` /
`currentModelId` block from its ACP `session/new` response and moved the
same catalog into a `configOptions` entry with `id`/`category` of
"model". parseACPSessionNewModels only understood the old shape, so
discovery silently returned an empty catalog and the model picker showed
"no available models" for an online, correctly-detected kimi runtime.

Parse `configOptions` as a fallback: the `models` block still wins when
present, so no existing ACP provider changes behaviour. `options[].value`
becomes the model id, `options[].name` the label, and `currentValue`
marks the default. Non-model options (thinking level) are deliberately
skipped — they are a separate product surface and would offer values
`session/set_model` cannot honour.

Also log a debug line with the top-level response keys (keys only, never
values) when session/new succeeds but advertises no catalog, so the next
round of upstream schema drift is visible in daemon.log instead of
looking like a PATH or install failure.

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-24 01:45:24 +08:00
Bohan Jiang
ffa8e16369 MUL-5228 fix(usage): bill Grok at xAI's reported cost, fix $0 resumed sessions (#5841)
* fix(agent): attribute Grok usage from the turn's own model id

A resumed Grok session with no configured model recorded its entire spend
under the model id "unknown", which matches no pricing row — so the task
reported $0 cost instead of its real spend.

grok.go only learned the model from the session handshake, and ACP's
`session/load` carries no model id (only `session/new` does). When neither
the agent nor MULTICA_GROK_MODEL pins a model, `daemon.go` legitimately
passes an empty model, leaving nothing to attribute the usage to.

Every Grok turn stamps `result._meta.modelId` with what it actually billed
against. Parse it in the shared ACP result parser and use it as the fallback
in grok.go. Other ACP backends are untouched — they keep whatever the
handshake gave them.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>

* fix(metrics): price the Grok catalog in server-side cost metrics

server/internal/metrics/pricing.go carried no Grok rows at all, so
RecordLLMUsage took the unpriced branch for every Grok turn: llm_cost_usd
reported zero Grok spend while the tokens accumulated in
llm_unpriced_tokens. Internal cost monitoring simply could not see Grok.

Add the six SKUs xAI publishes rates for, mirroring the frontend table in
packages/views/runtimes/utils.ts. Aliases are anchored exact matches like
the gpt-5.6 rows, so `grok-composer-*` (in the catalog, absent from the
price sheet) stays unmapped instead of inheriting a guessed rate.

Short-context tier on purpose: xAI bills a request at 2x once its prompt
reaches 200K tokens, but a usage record aggregates every model call in a
turn and cannot say which tier an individual request hit.

A regression test re-derives the cost of a real grok 0.2.106 turn from the
table and checks it against the costUsdTicks xAI returned for that turn.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>

* docs(changelog): scope the Grok cost claim to what was actually fixed

The v0.4.9 entry promised "accurate cost" in all four languages, but the
fix corrected catalog pricing and cached-input double-counting — it did not
implement xAI's 2x long-context tier, so a turn whose requests reach 200K
prompt tokens still under-reports by up to 50%. Say what was fixed instead.

Also correct two stale claims in the pricing comment: the daemon tags usage
rows with the runtime provider `grok`, not `xai` (the bare `grok-*` keys are
what make them resolve), and record why thresholding the long-context tier
on an aggregated row would be worse than not pricing it at all.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>

* feat(usage): carry the provider's own cost through to the usage record

Cost has always been derived client-side as tokens x a static rate, which
cannot express request-level pricing rules. xAI bills a Grok request at 2x
once its prompt reaches 200K tokens, and a task_usage row aggregates every
model call in a turn — so the stored token counts genuinely cannot say which
tier any individual request hit. Thresholding on the aggregate would be worse
than the status quo: it turns a bounded 50% under-estimate into an unbounded
over-estimate for turns made of many short requests.

Grok already reports what it charged, per turn, in `_meta.usage.costUsdTicks`.
Parse it, carry it through agent -> daemon -> API, and store it on task_usage
as a nullable BIGINT of 1e-10 USD ticks (integer, so sub-cent turns stay exact
end to end). NULL means the provider reported no cost — every pre-existing row
and every provider that doesn't return one. No backfill: there is no
authoritative figure to recover for those, and inventing one is the guess this
removes.

A single hourly bucket can mix rows that carry a cost with rows that don't, so
task_usage_hourly gains both halves: `cost_usd_ticks` sums the authoritative
side, and `uncosted_*_tokens` carry exactly the tokens that still need a
rate-table estimate. Consumers report authoritative + estimate(uncosted),
which degrades to today's behaviour when nothing in the bucket is
authoritative. The existing token columns keep covering every row, so token
displays are untouched. The new columns are additive with defaults, so the
unique key, the dirty-queue shape, and migration 102's triggers are unaffected.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>

* feat(usage): prefer the provider's own cost over the rate table

With the authoritative figure now stored, both cost consumers use it: the
usage dashboard (estimateCost / estimateCostBreakdown) and the server-side
llm_cost_usd metric. Each reports `authoritative + estimate(uncosted tokens)`,
so a row or bucket that mixes priced and unpriced sources stays whole.

The static rate tables remain, but for Grok they are now a fallback — they
still price usage recorded by a daemon too old to report cost, and every
provider that reports none. Custom pricing overrides likewise apply only to
the estimated half: they are a user's guess at a rate, and the authoritative
half is not a guess. A model with no rate-table row but a provider-reported
cost now also drops out of the "unmapped models" banner, since asking the user
to supply a rate for it would invite overriding a real bill.

llm_cost_usd is labelled by token_type and the provider reports one number per
turn, so the charge is distributed across the buckets in the rate table's own
proportions. Only the total is authoritative; the split stays an estimate,
which is why this scales the existing buckets rather than inventing a label.
estimateCostBreakdown does the same, keeping the stacked chart summing to the
headline figure instead of silently under-drawing every Grok row.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>

* docs(changelog): say Grok cost now follows xAI's actual charge

The earlier wording scoped the claim down to catalog pricing and cached input
because the long-context tier was still unhandled. It is handled now — the
cost comes from what xAI charged for the turn — so the entry can say so.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>

* fix(usage): keep the provider's cost when the model has no rate row

Both cost consumers bailed out before reading the authoritative figure when
the rate table had no row for the model. A `grok-composer-*` turn — in the
Grok Build catalog, absent from xAI's price sheet — was therefore reported as
$0 spend even though xAI told us exactly what it charged.

Worse on the client: estimateCost returned the real cost while
estimateCostBreakdown returned zeros, so the headline and the stacked chart
disagreed on precisely the rows whose cost is exact — and the unmapped-models
banner was (correctly) hidden, so nothing explained the discrepancy.

Handle the charge before the rate lookup in both places. Without rates there
is nothing to split a total by, so it lands whole in the `input` bucket, the
same fallback distributeAuthoritativeCost already uses when it has no shape to
scale. Tokens with no rate keep going to llm_unpriced_tokens: "unpriced"
describes the rate table, not the money.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>

* perf(usage): drop the historical rewrite from the cost-split migration

Migration 213 rewrote every existing task_usage_hourly row to seed the
uncosted counters. That is a full-table UPDATE inside a schema migration —
lock time, WAL and bloat all scaling with table size — for rows this issue
explicitly does not care about.

Deleting the UPDATE alone would have zeroed historical cost: with
`NOT NULL DEFAULT 0`, an untouched row asserts "nothing here needs
estimating", so every pre-split bucket would report $0 until the rollup
happened to touch it. Make the uncosted columns nullable with no default
instead. NULL means "never recomputed since the split existed", readers
COALESCE it to the row's own token total ("estimate all of it"), and the
pre-split behaviour is preserved exactly — with nothing to seed, so no
rewrite. A bare ADD COLUMN is metadata-only, so this is now fast DDL.

Rows heal into the split naturally as the rollup recomputes their buckets.

Verified on a fresh database: a legacy-shaped row reads back as its full
tokens to estimate, and a group mixing legacy and post-split buckets sums to
the authoritative cost plus both rows' estimable tokens.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-24 01:42:08 +08:00
Bohan Jiang
ecbdbda09e fix(agent): stop double-counting Grok cached input tokens (#5838)
Grok Build reports cachedReadTokens inside inputTokens (totalTokens ==
input + output on a real 0.2.106 turn, and that turn's costUsdTicks
matches xAI's rates only when the cached prefix is billed once). The
shared ACP parser persisted both counters raw, so the usage dashboard
charged the cached prefix at the full input rate *and* the cache-read
rate — ~4x the real spend on a cache-heavy turn.

Re-bucket cached reads out of input when totalTokens proves the overlap,
the same normalization codex.go already applies. Backends that report
mutually-exclusive buckets or omit totalTokens are untouched.

Co-authored-by: J <agent@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-23 17:55:33 +08:00
Multica Eve
6992c58de3 MUL-5185: add Codex Fast mode (#5821)
* feat(agents): add Codex fast mode (MUL-5185)

Co-authored-by: multica-agent <github@multica.ai>

* fix(agents): make Codex Fast override authoritative

Co-authored-by: multica-agent <github@multica.ai>

* fix(agents): remove Codex Fast config conflicts

Co-authored-by: multica-agent <github@multica.ai>

* chore: refresh checks after conflict resolution

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-23 15:40:18 +08:00
milymarkovic
09dce598df MUL-5155: KAP-1051: diagnose Codex thread/start timeouts fail-closed (#5759)
* fix(agent/codex): diagnose thread start timeouts (KAP-1051)

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent/codex): validate thread ID before success lifecycle

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent/codex): confirm process tree cleanup

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent/codex): bound Windows pipe cleanup

Co-authored-by: multica-agent <github@multica.ai>

* ci: run bounded Codex cleanup test on Windows

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): reap initialize timeout process groups

* fix(agent): redact initialize timeout stderr

* fix(agent): redact initialize context failures

---------

Co-authored-by: multica-agent <github@multica.ai>
2026-07-23 15:09:50 +08:00
YYClaw
36533bbc2b fix(test): prevent agent CLI execution in default tests (#5789) 2026-07-23 01:59:06 +08:00
Jiayuan Zhang
5d9295ac65 feat(agents): add per-agent runtime skill controls (#5686)
* feat(agents): add per-agent runtime skill controls

Co-authored-by: multica-agent <github@multica.ai>

* fix(agents): renumber runtime-skill migration and broadcast agent:status on toggle

Address the MUL-5101 review blockers on PR #5686:

- Rebase onto main and renumber the runtime-skill-disable migration
  202 -> 203. main added 202_runtime_profile_add_qwen, so the pair
  collided on prefix 202 and migrations_lint_test would reject the
  duplicate. 203 is the next free prefix.
- Publish an "agent:status" event after persisting a
  disabled_runtime_skills override, mirroring the workspace-skill toggle
  in writeUpdatedAgentSkills. The realtime layer keys off this event to
  invalidate workspaceKeys.agents, so other open web/desktop/mobile
  clients now drop their stale toggle state instead of only the
  initiating tab refreshing. Reload junction-table skills before the
  broadcast so it doesn't signal cleared skills (#3459).
- Add a handler regression test proving the broadcast fires on both
  disable and enable.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Walt <walt@multica.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-22 15:38:38 +08:00
Rusty Raven
3171e6607f MUL-5141: fix(agent): inject --yolo in Qwen headless runs so shell/edit/write tools are available (#5752)
* fix(agent): inject --yolo in Qwen headless runs so shell/edit/write tools are available (Fixes #5743)

Qwen Code's non-interactive mode (`-p … --output-format stream-json`) uses a
fail-closed approval policy: it silently drops `run_shell_command`, `edit`,
`write_file`, and `monitor` from the tool registry unless bypass mode is
active.  Every other Multica-supported coding adapter already injects its
equivalent permission flag as a daemon-owned argument (e.g. Claude uses
`--permission-mode bypassPermissions`, Grok uses `--always-approve`, Qoder
uses `--yolo --acp`).  Qwen was the only exception.

Changes:
- `buildQwenArgs`: append `--yolo` after the protocol flags and before any
  custom args so headless daemon runs always receive the full tool set.
- `qwenBlockedArgs`: add `--yolo`, `-y`, `--approval-mode`, and
  `--allowed-tools` as daemon-owned flags that are stripped from custom_args.
  This prevents users from accidentally or intentionally disabling bypass mode
  or narrowing the allowed tool set via per-agent settings.  `--exclude-tools`
  is intentionally left unblocked so users can still hard-deny specific tools.
- `TestBuildQwenArgsKeepsProtocolManaged`: extend with the new blocked flags
  and assert daemon-owned `--yolo` appears exactly once.
- `TestBuildQwenArgsYoloAlwaysPresent`: new test asserting `--yolo` is present
  even when `ExecOptions` carries no custom args.

* fix(agent): correct Qwen permission args and docs

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-22 13:35:59 +08:00
chenow9
d804dedcc6 MUL-5137: fix(agent): parse Grok ACP token usage from session/prompt _meta (#5748)
* fix(agent): parse Grok ACP token usage from session/prompt _meta

Grok Build places per-turn metering under result._meta (and
_meta.usage), not the top-level ACP usage field. Multica's shared
parser only read result.usage, so Grok tasks recorded empty token
usage and cost dashboards stayed at zero.

Fall back to _meta.usage (then flat _meta counters) when top-level
usage is absent. Prefer standard top-level usage when both are
present. Update the Grok fake ACP fixture and add regression tests
against the live 0.2.x payload shape.

* test(agent): cover zero ACP usage meta fallback

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-22 12:53:49 +08:00
Bohan Jiang
6ced54b44a fix(agent/codex): retry once when the model catalog refresh blocks the first turn (MUL-5110) (#5734)
Retries once when Codex's model catalog refresh blocks the first turn, with a narrow safety gate: first-turn no-progress timeout, zero semantic progress, catalog-refresh evidence in stderr, and a confirmed-reaped process tree. Clears ResumeSessionID on retry so a stalled thread is never resumed, and buffers the leading session pin so a discarded attempt cannot pin the resume pointer.
2026-07-21 21:26:11 +08:00
Bohan Jiang
a5a42846e6 fix(daemon): retry with a fresh session only when the resume was actually rejected (MUL-4966) (#5715)
* fix(daemon): gate fresh-session retry on tools executed, not session id (MUL-4966)

Switching provider accounts leaves the stored session id pointing at a
conversation the new account does not own. The daemon still passes it to
--resume, the provider rejects it, and the task dies before doing any work.

The existing fresh-session fallback was supposed to catch this but was
gated on `result.SessionID == ""`, which is not a lifecycle fact:

- Too narrow: a backend that echoes the requested id back when it rejects
  a resume keeps SessionID non-empty, so the fallback never fired — the
  reported bug.
- Too broad: a provider 401 before the first stream message also leaves
  SessionID empty, so an unrecoverable auth failure burned a second full
  run.

Gate on `tools == 0` instead. That states the property that actually makes
a retry safe — the agent executed no tool, so it mutated nothing, so
re-running cannot double-post a comment (comment creation has no
idempotency key and a duplicate re-fires its @mention triggers), reopen a
PR, or re-plan on top of its own half-finished work in the reused workdir.
Auth failures are excluded, mirroring retryableReasons in service/task.go.

The predicate is extracted to shouldRetryWithFreshSession so the tests
exercise production logic; both existing fallback tests re-implemented the
condition inline and would not have caught a regression in it.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): gate fresh-session retry on a positive resume-rejected signal (MUL-4966)

Review of the previous commit was right: `tools == 0` plus "not an auth
error" answers whether re-running is *safe*, not whether a new session can
*fix* the failure. Those are orthogonal, and answering the second by
exclusion inverts the burden of proof — the failures a fresh session cures
are a small enumerable set, while the ones it cannot are open-ended.

Concretely, the previous predicate fresh-retried on provider_network,
429/529, quota, 5xx and unclassified startup failures. provider_network is
the sharpest conflict: internal/service/task.go marks it resume-safe
(MUL-4910) specifically so the platform retry inherits the session and
continues the truncated conversation. Resetting the session first made that
contract unsatisfiable, silently discarding conversation context on a
transient blip — and rate limits got an immediate no-backoff re-run.

Replace the inference with positive evidence: agent.Result gains an explicit
ResumeRejected field, set only when a backend has proof the resume itself
was refused. claude/codebuddy/qwen derive it from resumeWasRejected, which
promotes the predicate resolveSessionID was already computing and encoding
as the side effect of blanking SessionID — using an empty string to carry
that meaning is what made the original bug possible. SessionID keeps being
dropped for a rejected resume (a dead pointer must not be persisted), but it
is no longer the signal the daemon reads to decide *why* a run failed.

The six ACP backends that recover from "session not found" set the flag at
the same points they already clear the id, so their existing recovery is not
caught by the narrower gate. codex needs nothing: thread/resume already falls
back to thread/start in-process, and deliberately does not on transport
errors.

Matching now includes the account-switch guardrail reported in #5704
(Claude Code 2.1.207, zh-CN): "400 此 session 已绑定另外的ai账号,请执行
/new 开启新 session". The en-US wording of the same guardrail has not been
captured yet, so those variants are marked inferred in the source; a miss
degrades to a terminal failure carrying the provider's raw text rather than
a mis-routed run.

Tests: backend-level fixtures drive ResumeRejected from real stream-json for
both the account-binding 400 and a network drop, and the predicate now
covers network/rate-limit/quota/5xx/auth/unclassified as explicit
non-retries.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): restore fresh-session recovery for backends with no rejection signal (MUL-4966)

Final review caught qwen regressing: its verified rejection string
("No saved session found with ID ...", already captured in
testdata/qwen-code-0.20.0-resume-not-found.stderr.txt) was not in the phrase
list, and qwen reports no session id on that path, so the new inclusion gate
turned a working auto-recovery into a terminal failure.

Auditing the other 17 resume-capable backends showed qwen was not alone.
antigravity, copilot, cursor, deveco and opencode all recovered from a
refused resume purely by reporting an empty SessionID, and none of them has
any rejection detection to convert into ResumeRejected — copilot's own
comment documents the hole (session.error before session.start), and
antigravity's helper returns "" when "the CLI exited before dispatching".
Making ResumeRejected the sole gate silently removed recovery from all five.

Fixing that by guessing rejection phrases for five more CLIs is the wrong
trade: no real output has been captured for any of them, and a false
positive discards a recoverable session pointer. So the gate is now two
tiers. Positive evidence (ResumeRejected) decides on its own where a backend
can produce it. Where none is available, an empty SessionID still gates the
retry — it proves no session was established, which is exactly what the
pre-change behaviour relied on — minus the classes a fresh session provably
cannot cure (network, rate limit, quota, provider 5xx, auth). That keeps the
resume-safe contract in internal/service/task.go intact while restoring what
these five backends had.

Also renames claudeResumeRejectedPhrases to resumeRejectedPhrases: it is
matched by claude, codebuddy and qwen, so a qwen-only string living under a
claude-prefixed name would be actively misleading.

Tests: qwen's existing missing-resume fixture now asserts ResumeRejected
(verified failing without the phrase), and the predicate covers the
no-signal tiers — retry when nothing was established, no retry once a
session exists or the failure classifies as uncurable.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): scope the no-signal fallback to backends that cannot detect rejections (MUL-4966)

Final review caught the compatibility path applying to every backend, not
just the five it was justified for. shouldRetryWithFreshSession only saw
(Result, priorSessionID, tools), so a false ResumeRejected could not be told
apart from a backend that has no way to answer — and claude/codebuddy/qwen/ACP
startup failures with no session id still fell through to the exclusion
branch. That contradicted both the stated intent and the function's own doc
comment ("where a backend can produce it, it is the whole answer").

Make the capability explicit. agent.ResumeRejectionUndetectable names the
five backends that scrape SessionID out of stream output and have no
rejection detection at all; the daemon takes provider and consults it, so a
capable backend reporting false is now taken at its word. Membership is
opt-in, so a new backend fails closed instead of silently inheriting a
guess-based retry.

Also completes the exclusion set: missing config, unavailable model, missing
executable, unsupported runtime version and (defensively) agent timeout all
have defined non-session remedies and were reaching `default: true`. What is
left through stays narrow — unknown, process failure, unparseable output,
context overflow — because a real rejection from these five most likely
surfaces as a non-zero exit or unparseable output, none of them reporting one
explicitly.

Tests: one identical result asserted across all five undetectable backends
(retries), twelve capable ones (no retry), and an unregistered provider
(fails closed), plus table cases for each newly excluded reason. Classifier
inputs were verified to map to the intended reasons rather than passing by
accident.

Also updates the ResumeRejected doc comment, which still said the daemon
gates on it alone.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-21 19:07:46 +08:00
Bohan Jiang
e5e278a1a2 fix(agent/cursor): send prompt on stdin so CLI-like flags cannot be re-tokenised (MUL-4992) (#5711)
* fix(agent/cursor): send prompt on stdin so CLI-like flags cannot be re-tokenised (MUL-4992)

On Windows a Cursor task whose prompt contains CLI-like flags failed in ~2s
with `error: unknown option '-X'` and no agent output (#5649). A pasted build
log such as

    go build -ldflags "-X main.version=foo" -o bin/server ./cmd/server

was enough to kill the run.

buildCursorArgs put the whole prompt in argv as the positional after -p.
The official Windows launcher chain ends in `& node.exe index.js $args`
inside cursor-agent.ps1, where PowerShell re-serialises $args onto node's
command line. Under Windows PowerShell 5.1 and pwsh <= 7.2 (Legacy native
argument passing) an argument holding embedded double quotes is not
re-escaped, so the quoted region closes early, node's argv parser re-splits
at the interior spaces, and `-X` reaches commander.js as a standalone flag.

#1709 removed the cmd.exe `%*` re-tokenisation but stopped at the
Go -> PowerShell boundary, one hop before this. Its Windows tests only
compare the argv slice Go builds and never execute a shim, which is why the
gap stayed invisible.

Fix: keep the prompt off every command line. cursor-agent's -p is a boolean
print-mode switch and the prompt is positional; with no positional prompt and
a non-TTY stdin the CLI reads stdin to EOF and uses that as the prompt. So
drop the prompt from argv on all platforms and write it to stdin, leaving
only fixed, content-free flags in argv. No shell or launcher on any platform
can re-tokenise what is not on a command line.

The write runs in its own goroutine: a prompt larger than the pipe buffer
(~64 KiB) blocks mid-write until the child drains it, and the child cannot
drain while nothing reads its stdout. Closing stdin signals end-of-prompt, so
it is closed on both success and error paths, and on cancellation to release
a blocked write. Write failures surface in the result diagnostic, ranked
below explicit agent errors so an early child exit (bad auth, bad flag) is
not masked by the resulting EPIPE.

Tests: a prompt carrying the exact `-ldflags "-X ..."` shape must arrive
byte-for-byte on stdin and appear nowhere in argv; a 512 KiB prompt must not
deadlock; the prompt is written verbatim (the CLI trims it, we do not). Both
new unix tests fail against the pre-fix code. A Windows-tagged test drives a
real PowerShell host through the same .cmd -> -File rewrite.

Closes #5649

Co-authored-by: multica-agent <github@multica.ai>

* test(agent/cursor): prove the Windows launcher fix on both PowerShell hosts in CI

The stdin fix has to hold on the host that actually exhibits the bug.
powershell.exe (5.1) and pwsh <= 7.2 default to Legacy native argument
passing; pwsh >= 7.3 defaults to Standard. A fix verified only on the newer
host would not be a fix for the reporter.

Run the shim probe against every PowerShell host on PATH rather than only the
one defaultPowerShellLookup would select, and hook the windows-tagged launcher
tests into the existing windows-execenv CI job. These tests are windows-tagged
and the backend job runs on ubuntu, so until now they ran nowhere.

Co-authored-by: multica-agent <github@multica.ai>

* test(ci): run Windows launcher tests verbosely so skips are visible

A skipped or unmatched test still reports "ok", which would make the
Windows job look like coverage it is not providing.

Co-authored-by: multica-agent <github@multica.ai>

* test(agent/cursor): close both coverage gaps in the #5649 regression tests

Two tests claimed guarantees they did not actually establish.

Windows: the fake cursor-agent.ps1 called [Console]::In.ReadToEnd() and wrote
the result itself, so it never launched a native child. The official shim ends
in `& node.exe index.js $args`, and that last hop is precisely where the bug
lives -- PowerShell re-serialises $args onto the child command line, and
whether the child inherits stdin was left unproven. The fake ps1 now
re-executes the test binary as a real native child (helper-process idiom),
which records the argv it actually received and drains stdin. Both PowerShell
hosts still run.

Interlock: the large-prompt fake drained stdin before writing any stdout, so
the child always unblocked the parent immediately and the test passed even
against a synchronous write -- it could not fail for the reason it existed.
The fake now floods stdout past pipe capacity *before* reading stdin, creating
the real mutual block. Verified: with the writer made synchronous the test
deadlocks to its 30s timeout, and passes only with the concurrent writer.

Also adds the missing cancellation case: a child that never reads stdin leaves
the writer blocked forever, so cancelling the context must close stdin,
release the writer and settle the run as aborted.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-21 16:01:04 +08:00
Bohan Jiang
fbe00ca164 fix(daemon): run Codex unsandboxed on Windows to stop reject-by-policy (MUL-4957) (#5672)
* fix(daemon): run Codex unsandboxed on Windows to stop reject-by-policy (MUL-4957)

Windows has no Landlock/Seatbelt-equivalent filesystem sandbox that the
daemon configures, so the per-task `sandbox_mode = "workspace-write"` it
wrote was unenforceable. Worse than having no sandbox, it pushed Codex
into rejecting non-safe mutation commands "by policy": `multica issue
create` fails with "was rejected by policy" because Codex can neither
sandbox the command nor (under approval_policy = "never") escalate it to
the daemon's auto-approver, so the request never reaches the approver.

Mirror the existing macOS fallback and give Windows danger-full-access so
those commands run. Also generalize the danger-full-access warn log so it
no longer hardcodes "on macOS" and only surfaces the macOS-specific
upgrade hint on macOS (new codexSandboxPolicy.Hint field).

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): correct Windows sandbox rationale and respect user windows.sandbox (MUL-4957)

Addresses two review must-fixes on #5672:

1. Correct a false security fact. The comments and log Reason claimed
   Windows has no filesystem sandbox backend. Codex 0.144.5 does ship a
   native Windows sandbox (windows.sandbox = "unelevated"/"elevated"); it
   is experimental with open upstream reliability bugs, so the daemon
   defaults to danger-full-access as a deliberate compatibility choice.
   Enabling the native sandbox is tracked as separate follow-up work.

2. Stop silently downgrading users who opted into isolation. The fallback
   was unconditional. Add codexSandboxPolicyForConfig: on Windows an
   explicit windows.sandbox = unelevated|elevated keeps workspace-write so
   Codex enforces task isolation with the user's chosen backend;
   danger-full-access applies only when windows.sandbox is absent,
   disabled, or unparseable. This is also the branch point for a future
   native-sandbox rollout (flip the default; callers unchanged).

Adds fixture tests locking the priority (user opt-in kept vs. unconfigured
fallback) plus predicate coverage for codexSandboxPolicyForConfig.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): fail closed on undecidable Windows sandbox config, honor -c windows.sandbox (MUL-4957)

Second review round on #5672. Two must-fixes.

1. Undecidable config no longer fails open. The old bool detector collapsed
   "unparseable / invalid value / failed copy" into "unconfigured" and then
   loosened to danger-full-access. Replaced with a tri-state
   (absent/native/undecidable): only exact-lowercase unelevated|elevated (the
   sole values Codex accepts — verified: any other value makes Codex refuse to
   load the config) counts as native; any other present value, unparseable
   TOML, a read error, or a missing per-task config when a shared
   ~/.codex/config.toml exists (i.e. the copy failed) is undecidable and fails
   closed to workspace-write — it never loosens — logged at error level.

2. windows.sandbox set via `-c`/`--config` custom args is now honored. Such
   args never land in config.toml, so config-only detection silently
   downgraded those users' isolation. The effective Codex args (daemon
   defaults + profile-fixed + per-agent custom_args) are threaded through
   PrepareParams/ReuseParams/CodexHomeOptions into the sandbox decision and
   scanned for a windows.sandbox override (inline, two-token, quoted, spaced;
   last-wins).

Also drops issue-status-bound source comments (openai/codex#24098 has since
closed). Adds unit coverage for config/args classification, the fold
precedence (undecidable > native > absent), and the copy-failed fail-closed
path.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): fail closed on config-sync errors and honor shell-quoted -c windows.sandbox (MUL-4957)

Round-3 review must-fixes:

1. resolveWindowsSandboxState now takes the config.toml sync error and a
   tri-state shared-config presence instead of re-stat-ing inside. A failed
   sync (stale/absent per-task copy) or an un-stat-able shared source is
   undecidable and keeps workspace-write, closing the fail-open where a failed
   sync was read as "unconfigured". Splits IO from the decision so the paths
   are unit-testable without faulting the filesystem.

2. The Windows sandbox decision consumes agent.NormalizeCodexLaunchArgs (the
   shared helper buildCodexArgs now uses) so a shell-quoted -c windows.sandbox
   opt-in is normalized identically to launch, instead of being missed by a
   raw-token scan and silently downgraded.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): abort when the Codex sandbox block cannot be written (MUL-4957)

Round-4 review must-fix: ensureCodexSandboxConfig failures were warn-and-continue,
so a computed fail-closed workspace-write policy could stay only in memory while
config.toml kept a stale danger-full-access from a prior run — the decision
failed closed but the effective config failed open.

prepareCodexHomeWithOpts now returns the error, which blocks startup on both
paths: fresh Prepare fails the task, and Reuse leaves env.CodexHome unset, which
configureCodexTaskShellEnvironment already refuses to start.

Regression covers the full reuse scenario (stale danger-full-access + failed
config sync + failed managed-block write); it fails with "got nil" without the
fix.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: J <j@multica.ai>
2026-07-21 15:18:44 +08:00
Bowser
f8bf6cd8b9 feat(runtime): add Qwen Code runtime (MUL-5015)
Merge approved PR #5666.
2026-07-21 14:55:08 +08:00
Bohan Jiang
e2f4f28462 fix(agent): skip redundant Hermes set_model when session already on the requested model (#5690)
When agent.model equals the model the Hermes ACP session already reports
as current, Multica was still replaying session/set_model. Hermes'
set_model re-runs provider auto-detection on the model id, which for a
provider:model id whose parsed provider matches the session's current
provider (e.g. a named custom endpoint exposed as "custom") can mis-route
to a different provider — custom:deepseek-v4-pro resolves into the
OpenRouter catalog and the turn fails with an auth error. Only the retry
via session resume, where the corrupted provider no longer matches the
custom: prefix, succeeds.

Capture the runtime's reported currentModelId from session/new and
session/resume and skip the redundant switch when it already equals the
requested model. An empty/unparsable current model falls through and
still sends set_model, preserving prior behaviour. The explicit
provider:model id is never rewritten. Upstream: NousResearch/hermes-agent#59089.

MUL-5029

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-20 23:01:48 +08:00
Multica Eve
f51828f568 MUL-4936: Fix empty replies in resumed Codex and Hermes chats (#5675)
* fix(agent): isolate resumed Codex turns (MUL-4936)

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): drain late Hermes replies (MUL-4936)

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): harden resumed chat cleanup (MUL-4936)

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): filter Codex subagents before turn gate (MUL-4936)

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-20 16:54:30 +08:00
Wangjue Yao
adc7636674 MUL-4921: fix(codex): report resumed usage as task delta (#5569)
* fix(codex): report resumed usage as task delta

Codex session logs expose cumulative total_token_usage values. Subtract the pre-task baseline so issue usage does not count earlier turns again.

Refs #5568

* fix(codex): harden resumed usage accounting

* fix(codex): scope fallback usage to current thread

Co-authored-by: multica-agent <github@multica.ai>

* fix(codex): harden rollout usage edge cases

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-20 14:59:04 +08:00
ZIce
910671b185 Fix Codex task MULTICA_TOKEN passthrough (ZIC-82)
Merge approved PR after independent final review; all required CI checks passed.
2026-07-17 14:29:46 +08:00
Multica Eve
ea8ccf3123 MUL-4903 fix(agent): harden Cursor terminal failures (#5559)
* fix(agent): harden Cursor terminal failures

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): preserve UTF-8 stderr tails

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-17 13:09:20 +08:00
lizhongxuan
6c84664089 test(agent): serialize thinking cache tests (#5551) 2026-07-17 11:20:57 +08:00
Jiayuan Zhang
ed9adc2bbe feat: improve create issue field controls (#5532)
Co-authored-by: Lambda <lambda@multica.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-17 11:11:52 +08:00
milymarkovic
983545b34d KAP-867: retry safe Codex initialize timeouts
Merge reviewed changes after fresh approval and green CI.
2026-07-16 16:35:51 +08:00
Multica Eve
cf1fc64671 MUL-4824: fail closed on incomplete Claude-style result streams (#5522)
* fix(agent): fail closed on incomplete stream results

Co-authored-by: multica-agent <github@multica.ai>

* fix(agent): tighten empty result fallback

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-16 15:22:46 +08:00
Multica Eve
5251a0958d fix(daemon): bound silent OpenCode streams (#5523)
Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-16 15:14:36 +08:00
Multica Eve
f0d7a444ac MUL-4860 fix(acp/kiro): GPT-5.6 Sol completion — object rawOutput + positive-proof guard (follow-up to #5511) (#5515)
* fix(kiro): require positive proof for -32603 completion rescue (MUL-4860, #5509)

Addresses the #5511 review (Elon):

Must-fix 1 — a missing result is not proof of success. A mid-command
crash / internal session failure produces the same 'tool use, no result,
-32603' shape as a genuine completion, so the previous 'saw the use but
never saw a result' rescue could mark an unfinished task completed. The
guard now rescues only on POSITIVE proof: a terminal ToolResult with
status=='completed' for a finishing tool (goal_complete / comment-add).

Must-fix 2 — the previous sticky booleans only ever wrote true, so a
later failed delivery could not override an earlier success, and call
ordering was ignored. Completion state is now the status of the MOST
RECENT finishing-tool result (keyed per CallID): 'progress completed →
final failed → -32603' stays failed; 'first failed → retry completed →
-32603' is a real completion.

Comment-add is still recognized by its command payload regardless of the
tool's display title (the #5509 core ask), so a completed comment-add
carrying a non-terminal title is preserved through the close handshake.

Tests: result-less running tool stays failed (both goal_complete and
comment-add), completed non-terminal-title comment-add is preserved,
failed-final-overrides-earlier-success, and completed-retry-after-failure.

Co-authored-by: multica-agent <github@multica.ai>

* fix(acp): don't drop completed tool_call_update when rawOutput is an object (MUL-4860, #5509)

Captured a live kiro-cli 2.12.3 + gpt-5.6-sol ACP trace and found the
true root cause: the completed tool_call_update sends rawOutput as an
OBJECT ({"items":[{"Json":{...}}]}), but handleToolCallUpdate typed
rawOutput/output as Go strings. json.Unmarshal then failed with
'cannot unmarshal object into Go struct field .rawOutput of type string'
and the handler returned early — silently dropping the ENTIRE update,
status:"completed" included. So no completion signal ever reached the
kiro completion-preservation guard, and the durable-but-closed task was
marked failed. This also explains the issue's 'shell tool recorded as
running': the tool_call title is literally 'Running: <cmd>' (kind
execute), which normalizes to 'running', not 'terminal'.

Fix: type rawOutput/output as json.RawMessage and render them via a new
acpRawText helper that accepts both a JSON string and a structured
value. This is a shared ACP-layer fix (hermes/kimi/kiro).

Combined with the earlier payload-based comment-add recognition and the
positive-proof completion guard, the completed result now flows through
and the -32603 close handshake correctly preserves completion.

Adds TestKiroBackendPreservesCompletionOnRealGPT56SolFrames (the exact
captured wire shape end-to-end) and TestACPRawText. Verified the
end-to-end test fails (status=failed) when rawOutput is typed as string
and passes after the fix.

Co-authored-by: multica-agent <github@multica.ai>

* fix(kiro): remove duplicate stale-session branch that faked completion (MUL-4860, #5509)

Addresses Elon's 3rd-round review of #5515. The -32603 guard had two
consecutive, identically-conditioned else-if branches for a stale resumed
session (opts.ResumeSessionID != "" && isACPSessionNotFound(err)). The
first wrongly forced finalStatus="completed" and kept the stale
SessionID; the second — the correct handler that clears SessionID so the
daemon's fresh-session retry fires — was unreachable. So a recovery whose
resumed session was gone at session/prompt time was reported as a fake
success AND skipped the retry.

Remove the bogus first branch, keeping the positive-proof completion
branch and the original stale-session (SessionID="") branch.

Adds TestKiroBackendClearsSessionIDWhenPromptSessionNotFound (the prompt
path; existing coverage only exercised session/set_model). Verified it
returns completed+ses_stale on the pre-fix code and failed+empty
SessionID after. Also gofmt on the touched files.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-16 13:36:45 +08:00
Multica Eve
dd692058d7 fix(kiro): preserve completed status on GPT-5.6 Sol -32603 close handshake (#5509) (#5511)
The -32603 completion-preservation guard coupled to the Claude/Kiro ACP
event shape. GPT-5.6 Sol leaves the finishing tool (goal_complete /
'multica issue comment add') parked at 'running' and titles the shell
tool with a name that does not normalize to 'terminal', so:

- isKiroIssueCommentAddTool's msg.Tool=='terminal' gate dropped the tool
  use entirely, and
- with no completed/failed ToolCallUpdate, no ToolResult is emitted, so
  the saw*Completed flags never flipped.

Completed tasks were therefore reversed to failed with
agent_error.provider_server_error ~17s after finishing their work.

Fix:
- Recognize comment-add by its command payload, independent of the
  normalized tool title.
- Track each finishing tool along use / result / completed axes and
  preserve completed when the -32603 close handshake follows a tool use
  that never produced a terminal result. A genuinely failed ToolResult
  still trips saw*Result and keeps the task failed.

Adds regression fixtures for the running-tool path (comment-add and
goal_complete) plus a failed-result safety case.

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-16 12:33:15 +08:00
Tran Quang
6455d390e2 fix(agent): fail opencode runs whose stream ends without a terminal signal (#5238)
* fix(agent): fail opencode runs whose stream ends without a terminal signal

* fix(agent): treat step_finish(reason "tool-calls") as non-terminal in the EOF guard

Review follow-up: step bracketing alone left a false-green window. A tool
step closes with step_finish(reason "tool-calls") before its continuation
step_start, so a stream that dies in that gap had no open step and still
reported "completed".

Parse part.reason (live-probed on opencode 1.17.16: tool loop emits
reasons "tool-calls" then "stop") and keep the run non-terminal after a
"tool-calls" finish until the next step_start. Any other reason —
including absent, for older opencode versions that predate the field —
stays terminal so healthy runs on old protocols are not mass-failed.

Tests: tool-calls finish then EOF now fails; the multi-step happy path
uses the real wire shape ending in reason "stop"; a reason-less
step_finish (legacy protocol) still completes.

* fix(agent): keep stop tool steps pending continuation

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-15 16:40:59 +08:00
Multica Eve
66316c2614 fix(grok): harden ACP authentication and capabilities (#5440)
Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-15 13:35:55 +08:00