Files
multica/server/pkg/agent/traecli.go
Bohan Jiang a5a42846e6 fix(daemon): retry with a fresh session only when the resume was actually rejected (MUL-4966) (#5715)
* fix(daemon): gate fresh-session retry on tools executed, not session id (MUL-4966)

Switching provider accounts leaves the stored session id pointing at a
conversation the new account does not own. The daemon still passes it to
--resume, the provider rejects it, and the task dies before doing any work.

The existing fresh-session fallback was supposed to catch this but was
gated on `result.SessionID == ""`, which is not a lifecycle fact:

- Too narrow: a backend that echoes the requested id back when it rejects
  a resume keeps SessionID non-empty, so the fallback never fired — the
  reported bug.
- Too broad: a provider 401 before the first stream message also leaves
  SessionID empty, so an unrecoverable auth failure burned a second full
  run.

Gate on `tools == 0` instead. That states the property that actually makes
a retry safe — the agent executed no tool, so it mutated nothing, so
re-running cannot double-post a comment (comment creation has no
idempotency key and a duplicate re-fires its @mention triggers), reopen a
PR, or re-plan on top of its own half-finished work in the reused workdir.
Auth failures are excluded, mirroring retryableReasons in service/task.go.

The predicate is extracted to shouldRetryWithFreshSession so the tests
exercise production logic; both existing fallback tests re-implemented the
condition inline and would not have caught a regression in it.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): gate fresh-session retry on a positive resume-rejected signal (MUL-4966)

Review of the previous commit was right: `tools == 0` plus "not an auth
error" answers whether re-running is *safe*, not whether a new session can
*fix* the failure. Those are orthogonal, and answering the second by
exclusion inverts the burden of proof — the failures a fresh session cures
are a small enumerable set, while the ones it cannot are open-ended.

Concretely, the previous predicate fresh-retried on provider_network,
429/529, quota, 5xx and unclassified startup failures. provider_network is
the sharpest conflict: internal/service/task.go marks it resume-safe
(MUL-4910) specifically so the platform retry inherits the session and
continues the truncated conversation. Resetting the session first made that
contract unsatisfiable, silently discarding conversation context on a
transient blip — and rate limits got an immediate no-backoff re-run.

Replace the inference with positive evidence: agent.Result gains an explicit
ResumeRejected field, set only when a backend has proof the resume itself
was refused. claude/codebuddy/qwen derive it from resumeWasRejected, which
promotes the predicate resolveSessionID was already computing and encoding
as the side effect of blanking SessionID — using an empty string to carry
that meaning is what made the original bug possible. SessionID keeps being
dropped for a rejected resume (a dead pointer must not be persisted), but it
is no longer the signal the daemon reads to decide *why* a run failed.

The six ACP backends that recover from "session not found" set the flag at
the same points they already clear the id, so their existing recovery is not
caught by the narrower gate. codex needs nothing: thread/resume already falls
back to thread/start in-process, and deliberately does not on transport
errors.

Matching now includes the account-switch guardrail reported in #5704
(Claude Code 2.1.207, zh-CN): "400 此 session 已绑定另外的ai账号,请执行
/new 开启新 session". The en-US wording of the same guardrail has not been
captured yet, so those variants are marked inferred in the source; a miss
degrades to a terminal failure carrying the provider's raw text rather than
a mis-routed run.

Tests: backend-level fixtures drive ResumeRejected from real stream-json for
both the account-binding 400 and a network drop, and the predicate now
covers network/rate-limit/quota/5xx/auth/unclassified as explicit
non-retries.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): restore fresh-session recovery for backends with no rejection signal (MUL-4966)

Final review caught qwen regressing: its verified rejection string
("No saved session found with ID ...", already captured in
testdata/qwen-code-0.20.0-resume-not-found.stderr.txt) was not in the phrase
list, and qwen reports no session id on that path, so the new inclusion gate
turned a working auto-recovery into a terminal failure.

Auditing the other 17 resume-capable backends showed qwen was not alone.
antigravity, copilot, cursor, deveco and opencode all recovered from a
refused resume purely by reporting an empty SessionID, and none of them has
any rejection detection to convert into ResumeRejected — copilot's own
comment documents the hole (session.error before session.start), and
antigravity's helper returns "" when "the CLI exited before dispatching".
Making ResumeRejected the sole gate silently removed recovery from all five.

Fixing that by guessing rejection phrases for five more CLIs is the wrong
trade: no real output has been captured for any of them, and a false
positive discards a recoverable session pointer. So the gate is now two
tiers. Positive evidence (ResumeRejected) decides on its own where a backend
can produce it. Where none is available, an empty SessionID still gates the
retry — it proves no session was established, which is exactly what the
pre-change behaviour relied on — minus the classes a fresh session provably
cannot cure (network, rate limit, quota, provider 5xx, auth). That keeps the
resume-safe contract in internal/service/task.go intact while restoring what
these five backends had.

Also renames claudeResumeRejectedPhrases to resumeRejectedPhrases: it is
matched by claude, codebuddy and qwen, so a qwen-only string living under a
claude-prefixed name would be actively misleading.

Tests: qwen's existing missing-resume fixture now asserts ResumeRejected
(verified failing without the phrase), and the predicate covers the
no-signal tiers — retry when nothing was established, no retry once a
session exists or the failure classifies as uncurable.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): scope the no-signal fallback to backends that cannot detect rejections (MUL-4966)

Final review caught the compatibility path applying to every backend, not
just the five it was justified for. shouldRetryWithFreshSession only saw
(Result, priorSessionID, tools), so a false ResumeRejected could not be told
apart from a backend that has no way to answer — and claude/codebuddy/qwen/ACP
startup failures with no session id still fell through to the exclusion
branch. That contradicted both the stated intent and the function's own doc
comment ("where a backend can produce it, it is the whole answer").

Make the capability explicit. agent.ResumeRejectionUndetectable names the
five backends that scrape SessionID out of stream output and have no
rejection detection at all; the daemon takes provider and consults it, so a
capable backend reporting false is now taken at its word. Membership is
opt-in, so a new backend fails closed instead of silently inheriting a
guess-based retry.

Also completes the exclusion set: missing config, unavailable model, missing
executable, unsupported runtime version and (defensively) agent timeout all
have defined non-session remedies and were reaching `default: true`. What is
left through stays narrow — unknown, process failure, unparseable output,
context overflow — because a real rejection from these five most likely
surfaces as a non-zero exit or unparseable output, none of them reporting one
explicitly.

Tests: one identical result asserted across all five undetectable backends
(retries), twelve capable ones (no retry), and an unregistered provider
(fails closed), plus table cases for each newly excluded reason. Classifier
inputs were verified to map to the intended reasons rather than passing by
accident.

Also updates the ResumeRejected doc comment, which still said the daemon
gates on it alone.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-21 19:07:46 +08:00

448 lines
14 KiB
Go

package agent
import (
"bufio"
"context"
"fmt"
"io"
"os/exec"
"strings"
"sync"
"sync/atomic"
"time"
)
// traecliBlockedArgs are flags hardcoded by the daemon that must not be
// overridden by user-configured custom_args. `acp` and `serve` are the
// protocol subcommand/action; `-y`/`--yolo` is daemon-owned so headless ACP
// always runs in bypass-permissions mode (the official traecli gates non-read
// tools behind a permission prompt otherwise — see
// https://docs.trae.cn/cli_permission-mode). `--print`/`-p` and
// `--output-format` would switch the binary out of ACP into print mode and
// break the daemon↔traecli transport, and `--permission-mode` is owned by the
// daemon via --yolo.
var traecliBlockedArgs = map[string]blockedArgMode{
"acp": blockedStandalone,
"serve": blockedStandalone,
"-y": blockedStandalone,
"--yolo": blockedStandalone,
"-p": blockedStandalone,
"--print": blockedStandalone,
"--output-format": blockedWithValue,
"--permission-mode": blockedWithValue,
}
// traecliBackend implements Backend by spawning `traecli acp serve --yolo` and
// communicating via the standard ACP (Agent Client Protocol) JSON-RPC 2.0
// transport over stdin/stdout.
//
// This targets ByteDance's official TRAE CLI (the `traecli` binary documented
// at https://docs.trae.cn/cli — NOT the open-source bytedance/trae-agent
// `trae-cli`, which has no ACP transport). traecli is ACP-native: it advertises
// `loadSession: true` and `mcpCapabilities: {http, sse}` from `initialize`,
// returns its model catalog from `session/new`, and supports `session/load`
// (resume), `session/set_model`, and `session/prompt`. That lets the existing
// Hermes/Kimi/Kiro/Qoder ACP client (hermesClient) drive it with only
// provider-specific launch args and tool-name normalization.
//
// The `initialize` capabilities above were captured from the real traecli
// v0.120.42 binary; the streaming gate mirrors the other ACP backends so
// history replay flushed during session setup does not corrupt the streamed
// output.
type traecliBackend struct {
cfg Config
}
var traecliReaderDrainGrace = 2 * time.Second
// traecliMessageStream serializes sends and the final close so a late stdout
// reader (traecli may keep the process and its pipes open briefly after
// session/prompt returns) cannot send on a closed channel. Mirrors qoder.
type traecliMessageStream struct {
ch chan Message
mu sync.Mutex
closed bool
}
func newTraecliMessageStream(size int) *traecliMessageStream {
return &traecliMessageStream{ch: make(chan Message, size)}
}
func (s *traecliMessageStream) send(msg Message) {
s.mu.Lock()
defer s.mu.Unlock()
if s.closed {
return
}
trySend(s.ch, msg)
}
func (s *traecliMessageStream) close() {
s.mu.Lock()
defer s.mu.Unlock()
if s.closed {
return
}
s.closed = true
close(s.ch)
}
func (b *traecliBackend) Execute(ctx context.Context, prompt string, opts ExecOptions) (*Session, error) {
execPath := b.cfg.ExecutablePath
if execPath == "" {
execPath = "traecli"
}
if _, err := exec.LookPath(execPath); err != nil {
return nil, fmt.Errorf("traecli executable not found at %q: %w", execPath, err)
}
// Translate the agent's mcp_config (Claude-style object of objects) into
// the array shape ACP session/new and session/load expect. Fail closed on
// malformed JSON so the launch surfaces the real error instead of silently
// dropping every MCP server.
mcpServers, err := buildACPMcpServers(opts.McpConfig, b.cfg.Logger)
if err != nil {
return nil, fmt.Errorf("traecli: invalid mcp_config: %w", err)
}
timeout := opts.Timeout
runCtx, cancel := runContext(ctx, timeout)
traecliArgs := append(
[]string{"acp", "serve", "--yolo"},
filterCustomArgs(opts.CustomArgs, traecliBlockedArgs, b.cfg.Logger)...,
)
cmd := exec.CommandContext(runCtx, execPath, traecliArgs...)
hideAgentWindow(cmd)
b.cfg.Logger.Info("agent command", "exec", execPath, "args", traecliArgs)
if opts.Cwd != "" {
cmd.Dir = opts.Cwd
}
cmd.Env = buildEnv(b.cfg.Env)
stdout, err := cmd.StdoutPipe()
if err != nil {
cancel()
return nil, fmt.Errorf("traecli stdout pipe: %w", err)
}
stdin, err := cmd.StdinPipe()
if err != nil {
cancel()
return nil, fmt.Errorf("traecli stdin pipe: %w", err)
}
// StderrPipe + an explicit copier give us a join point (`stderrDone`) that
// fires before the failure-promotion decision; see the matching comment in
// hermes.go for why the io.MultiWriter form races with stopReason=end_turn
// under load.
providerErr := newACPProviderErrorSniffer("traecli")
stderr, err := cmd.StderrPipe()
if err != nil {
cancel()
return nil, fmt.Errorf("traecli stderr pipe: %w", err)
}
if err := cmd.Start(); err != nil {
cancel()
return nil, fmt.Errorf("start traecli: %w", err)
}
stderrSink := io.MultiWriter(newLogWriter(b.cfg.Logger, "[traecli:stderr] "), providerErr)
stderrDone := make(chan struct{})
go func() {
defer close(stderrDone)
_, _ = io.Copy(stderrSink, stderr)
}()
b.cfg.Logger.Info("traecli acp started", "pid", cmd.Process.Pid, "cwd", opts.Cwd)
msgStream := newTraecliMessageStream(256)
resCh := make(chan Result, 1)
var outputMu sync.Mutex
var output strings.Builder
var streamingCurrentTurn atomic.Bool
promptDone := make(chan hermesPromptResult, 1)
c := &hermesClient{
cfg: b.cfg,
stdin: stdin,
pending: make(map[int]*pendingRPC),
pendingTools: make(map[string]*pendingToolCall),
acceptNotification: func(string) bool {
return streamingCurrentTurn.Load()
},
onMessage: func(msg Message) {
if !streamingCurrentTurn.Load() {
return
}
if msg.Type == MessageToolUse {
msg.Tool = kimiToolNameFromTitle(msg.Tool)
}
if msg.Type == MessageText {
outputMu.Lock()
output.WriteString(msg.Content)
outputMu.Unlock()
}
msgStream.send(msg)
},
onPromptDone: func(result hermesPromptResult) {
if !streamingCurrentTurn.Load() {
return
}
select {
case promptDone <- result:
default:
}
},
}
readerDone := make(chan struct{})
go func() {
defer close(readerDone)
scanner := bufio.NewScanner(stdout)
scanner.Buffer(make([]byte, 0, 1024*1024), 10*1024*1024)
for scanner.Scan() {
line := strings.TrimSpace(scanner.Text())
if line == "" {
continue
}
c.handleLine(line)
}
c.closeAllPending(fmt.Errorf("traecli process exited"))
}()
go func() {
defer cancel()
defer msgStream.close()
defer close(resCh)
defer func() {
stdin.Close()
_ = cmd.Wait()
}()
startTime := time.Now()
finalStatus := "completed"
var finalError string
var sessionID string
// Set when the ACP runtime refuses the session we asked to
// resume. Only that is curable by starting a fresh session, so
// handshake/network failures below must leave it false.
var resumeRejected bool
effectiveModel := strings.TrimSpace(opts.Model)
initResult, err := c.request(runCtx, "initialize", map[string]any{
"protocolVersion": 1,
"clientInfo": map[string]any{
"name": "multica-agent-sdk",
"version": "0.2.0",
},
"clientCapabilities": map[string]any{},
})
if err != nil {
finalStatus = "failed"
finalError = fmt.Sprintf("traecli initialize failed: %v", err)
resCh <- Result{Status: finalStatus, Error: finalError, DurationMs: time.Since(startTime).Milliseconds()}
return
}
// Drop MCP entries whose remote transport the runtime didn't advertise
// (traecli advertises mcpCapabilities {http, sse}). See hermes.go for
// why sending an unsupported transport tanks the whole session/new.
mcpServers = filterACPMcpServersByCapability(mcpServers, extractACPMcpCapabilities(initResult), "traecli", b.cfg.Logger)
cwd := opts.Cwd
if cwd == "" {
cwd = "."
}
if opts.ResumeSessionID != "" {
// traecli advertises loadSession:true, so resume goes through the
// standard ACP session/load (same as Kiro).
result, err := c.request(runCtx, "session/load", map[string]any{
"cwd": cwd,
"sessionId": opts.ResumeSessionID,
"mcpServers": mcpServers,
})
if err != nil {
finalStatus = "failed"
finalError = fmt.Sprintf("traecli session/load failed: %v", err)
resCh <- Result{Status: finalStatus, Error: finalError, DurationMs: time.Since(startTime).Milliseconds()}
return
}
var changed bool
sessionID, changed = resolveResumedSessionID(opts.ResumeSessionID, result)
if changed {
b.cfg.Logger.Warn("agent returned a different session id on resume — original was likely lost; continuing with the new id",
"backend", "traecli",
"requested", opts.ResumeSessionID,
"actual", sessionID,
)
}
if effectiveModel == "" {
effectiveModel = extractACPCurrentModelID(result)
}
} else {
result, err := c.request(runCtx, "session/new", map[string]any{
"cwd": cwd,
"mcpServers": mcpServers,
})
if err != nil {
finalStatus = "failed"
finalError = fmt.Sprintf("traecli session/new failed: %v", err)
resCh <- Result{Status: finalStatus, Error: finalError, DurationMs: time.Since(startTime).Milliseconds()}
return
}
sessionID = extractACPSessionID(result)
if sessionID == "" {
finalStatus = "failed"
finalError = "traecli session/new returned no session ID"
resCh <- Result{Status: finalStatus, Error: finalError, DurationMs: time.Since(startTime).Milliseconds()}
return
}
if effectiveModel == "" {
effectiveModel = extractACPCurrentModelID(result)
}
}
c.sessionID = sessionID
b.cfg.Logger.Info("traecli session created", "session_id", sessionID)
if opts.Model != "" {
if _, err := c.request(runCtx, "session/set_model", map[string]any{
"sessionId": sessionID,
"modelId": opts.Model,
}); err != nil {
b.cfg.Logger.Warn("traecli set_session_model failed", "error", err, "requested_model", opts.Model)
finalStatus = "failed"
finalError = fmt.Sprintf("traecli could not switch to model %q: %v", opts.Model, err)
if opts.ResumeSessionID != "" && isACPSessionNotFound(err) {
b.cfg.Logger.Warn("resumed session not found at set_model time; clearing session id so the daemon retries fresh",
"backend", "traecli",
"session_id", sessionID,
)
sessionID = ""
resumeRejected = true
}
resCh <- Result{
Status: finalStatus,
Error: finalError,
DurationMs: time.Since(startTime).Milliseconds(),
SessionID: sessionID,
ResumeRejected: resumeRejected,
}
return
}
b.cfg.Logger.Info("traecli session model set", "model", opts.Model)
}
userText := prompt
if opts.SystemPrompt != "" {
userText = opts.SystemPrompt + "\n\n---\n\n" + prompt
}
streamingCurrentTurn.Store(true)
_, err = c.request(runCtx, "session/prompt", map[string]any{
"sessionId": sessionID,
"prompt": []map[string]any{
{"type": "text", "text": userText},
},
})
if err != nil {
if runCtx.Err() == context.DeadlineExceeded {
finalStatus = "timeout"
finalError = fmt.Sprintf("traecli timed out after %s", timeout)
} else if runCtx.Err() == context.Canceled {
finalStatus = "aborted"
finalError = "execution cancelled"
} else {
finalStatus = "failed"
finalError = fmt.Sprintf("traecli session/prompt failed: %v", err)
if opts.ResumeSessionID != "" && isACPSessionNotFound(err) {
b.cfg.Logger.Warn("resumed session not found at prompt time; clearing session id so the daemon retries fresh",
"backend", "traecli",
"session_id", sessionID,
)
sessionID = ""
resumeRejected = true
}
}
} else {
select {
case pr := <-promptDone:
if pr.stopReason == "cancelled" {
finalStatus = "aborted"
finalError = "traecli cancelled the prompt"
}
c.usageMu.Lock()
c.usage.InputTokens += pr.usage.InputTokens
c.usage.OutputTokens += pr.usage.OutputTokens
c.usage.CacheReadTokens += pr.usage.CacheReadTokens
c.usageMu.Unlock()
default:
}
}
duration := time.Since(startTime)
b.cfg.Logger.Info("traecli finished", "pid", cmd.Process.Pid, "status", finalStatus, "duration", duration.Round(time.Millisecond).String())
stdin.Close()
cancel()
// traecli ACP may keep the process — and the stdout/stderr pipes — open
// briefly after session/prompt returns. The prompt response is already
// terminal, so bound the drain: wait for the stdout reader and the
// stderr copier, but no longer than the grace window (CommandContext
// cancellation tears the process down). Draining stderr is what makes
// the provider-error promotion below see a terminal marker.
drainCtx, drainCancel := context.WithTimeout(context.Background(), traecliReaderDrainGrace)
select {
case <-readerDone:
case <-drainCtx.Done():
}
select {
case <-stderrDone:
case <-drainCtx.Done():
}
drainCancel()
// Flip the gate before the defer closes msgStream; a late reader that
// already passed the gate is serialized by traecliMessageStream so the
// late send is dropped instead of panicking.
streamingCurrentTurn.Store(false)
outputMu.Lock()
finalOutput := output.String()
outputMu.Unlock()
// Promote completed→failed when stderr or the agent text stream show a
// terminal upstream-LLM failure (HTTP 4xx / rate-limit / expired token).
// Mirrors hermes/kimi/kiro/qoder.
finalStatus, finalError = promoteACPResultOnProviderError(finalStatus, finalError, finalOutput, providerErr)
c.usageMu.Lock()
u := c.usage
c.usageMu.Unlock()
var usageMap map[string]TokenUsage
if u.InputTokens > 0 || u.OutputTokens > 0 || u.CacheReadTokens > 0 || u.CacheWriteTokens > 0 {
model := effectiveModel
if model == "" {
model = "unknown"
}
usageMap = map[string]TokenUsage{model: u}
}
resCh <- Result{
Status: finalStatus,
Output: finalOutput,
Error: finalError,
DurationMs: duration.Milliseconds(),
SessionID: sessionID,
ResumeRejected: resumeRejected,
Usage: usageMap,
}
}()
return &Session{Messages: msgStream.ch, Result: resCh}, nil
}