mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-05 09:30:05 +02:00
* fix(agent): make cursor stream protocol drift loud instead of silent (MUL-5434) #6071 reports Cursor tasks that demonstrably read files, ran commands and called the Multica CLI, yet showed a single blob of agent text with no reasoning and no tool rows. The run still reported success and tools=0. `switch evt.Type` in cursor.go had no `default` branch, so any top-level event type we do not handle was dropped with no counter and no warning. Renaming only the top-level types of a healthy stream (`thinking`->`reasoning`, `tool_call`->`tool_calls`), leaving every nested field untouched, reproduces the report exactly: status=completed, output = the result text alone, tool_use=0, thinking=0, zero diagnostics. The existing unknown-subtype warning cannot catch this — it only increments once the type has already matched — so "no unknown-subtype warning" does not rule out protocol drift. Two diagnostic gaps are closed: - Add the `default` branch with a bounded, content-free tally of unhandled top-level types, reported once per run as a warning and alongside tool_use_count in the protocol summary. Type names are normalized through observedCursorEventType and distinct names are capped at 16 plus an overflow bucket, so a hostile or noisy stream cannot grow the map or leak payload into logs. `user` (the CLI echoing our prompt, present in every recorded run) is explicitly benign so the warning does not fire always. - Count assistant text separately from the terminal result text. The result event writes into the same builder, so last_assistant_bytes equalled result_bytes even when the assistant streamed nothing — erasing the signal that says "only the final answer arrived". This also aligns cursor with claude, which already reports assistant bytes only. Unrecognized events are still never coerced into tool or reasoning messages; guessing at upstream additions is the failure mode MUL-5231 already fixed once. This is diagnosis only and does not itself restore the missing tool rows — identifying which upstream shape changed requires a captured 2026.07.23 stream, and this warning is what makes that identifiable from a single production log line. Co-authored-by: multica-agent <github@multica.ai> * fix(agent): report cursor unhandled events as evidence, not as a verdict Addresses the review's must-fix on MUL-5434. The diagnostic was described as deciding WHY a transcript is empty, which it cannot do: - "tools=0 with unhandled types" does not establish that a rename ate the tool rows. Cursor 2026.07.23 also emits transport/control frames, so an unhandled type only proves the stream carried events we do not parse. - "tools=0 with no unhandled types" does not establish the agent used no tools. The CLI may execute tools without handing the updates to its stream serializer at all — the main branch #6071 has NOT ruled out — a new shape may be nested inside an event type we already recognize, or events may be lost to invalid framing or a scanner boundary. Changes, no behaviour change to messages, status or output: - Reword the tally doc, the switch default, the warning and the shared observation struct to state what a non-zero and a zero count each do and do not establish, and point at the branches that stay open. - Rename unknown* to unhandled* (fields, log keys, warning text, the pre-existing subtype counter) so the diagnostic never implies the type is unrecognized upstream — only that this parser does not handle it. - Classify `connection` and `retry` explicitly. They join `user` in cursorNonTranscriptEventTypes, with per-entry provenance recorded: `user` is confirmed in the recorded 2026.07.20 stream, the control frames are reported on newer builds and listed defensively. A build that does not emit them makes the entry inert; one that does must not have a known control frame reported as an unhandled protocol event. Tests assert signal presence and that nothing is fabricated, not causality: the healthy-stream test now carries `connection` / `retry` and additionally asserts the real thinking/tool rows still arrive, so suppressing the warning cannot silently suppress the transcript. TestCursorNonTranscriptEventType also pins that no type the parser handles can enter the suppression list. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Eve <eve@multica-ai.local> Co-authored-by: multica-agent <github@multica.ai>
1012 lines
34 KiB
Go
1012 lines
34 KiB
Go
package agent
|
|
|
|
import (
|
|
"bufio"
|
|
"context"
|
|
"encoding/json"
|
|
"fmt"
|
|
"io"
|
|
"log/slog"
|
|
"os/exec"
|
|
"regexp"
|
|
"sort"
|
|
"strings"
|
|
"sync"
|
|
"time"
|
|
)
|
|
|
|
// cursorBackend implements Backend by spawning the Cursor Agent CLI
|
|
// (cursor-agent) with --output-format stream-json and parsing the JSONL
|
|
// event stream. The protocol is similar to Claude Code's stream-json
|
|
// format: events are newline-delimited JSON objects with a "type" field.
|
|
type cursorBackend struct {
|
|
cfg Config
|
|
}
|
|
|
|
func (b *cursorBackend) Execute(ctx context.Context, prompt string, opts ExecOptions) (*Session, error) {
|
|
execName := b.cfg.ExecutablePath
|
|
if execName == "" {
|
|
execName = "cursor-agent"
|
|
}
|
|
lookedUp, err := exec.LookPath(execName)
|
|
if err != nil {
|
|
return nil, fmt.Errorf("cursor-agent executable not found at %q: %w", execName, err)
|
|
}
|
|
|
|
timeout := opts.Timeout
|
|
runCtx, cancel := runContext(ctx, timeout)
|
|
|
|
args := buildCursorArgs(opts, b.cfg.Logger)
|
|
argv0, cmdArgs := chooseCursorInvocation(execName, lookedUp, args, b.cfg.Logger)
|
|
|
|
cmd := exec.CommandContext(runCtx, argv0, cmdArgs...)
|
|
hideAgentWindow(cmd)
|
|
b.cfg.Logger.Info("agent command", "exec", argv0, "args", cmdArgs)
|
|
cmd.WaitDelay = 500 * time.Millisecond
|
|
if opts.Cwd != "" {
|
|
cmd.Dir = opts.Cwd
|
|
}
|
|
cmd.Env = buildEnv(b.cfg.Env)
|
|
|
|
stdout, err := cmd.StdoutPipe()
|
|
if err != nil {
|
|
cancel()
|
|
return nil, fmt.Errorf("cursor stdout pipe: %w", err)
|
|
}
|
|
stdin, err := cmd.StdinPipe()
|
|
if err != nil {
|
|
cancel()
|
|
return nil, fmt.Errorf("cursor stdin pipe: %w", err)
|
|
}
|
|
var closeStdinOnce sync.Once
|
|
closeStdin := func() { closeStdinOnce.Do(func() { _ = stdin.Close() }) }
|
|
stderrBuf := newStderrTail(newLogWriter(b.cfg.Logger, "[cursor:stderr] "), agentStderrTailBytes)
|
|
cmd.Stderr = stderrBuf
|
|
|
|
if err := cmd.Start(); err != nil {
|
|
closeStdin()
|
|
cancel()
|
|
return nil, fmt.Errorf("start cursor-agent: %w", err)
|
|
}
|
|
|
|
b.cfg.Logger.Info("cursor-agent started", "pid", cmd.Process.Pid, "cwd", opts.Cwd, "model", opts.Model)
|
|
|
|
msgCh := make(chan Message, 256)
|
|
resCh := make(chan Result, 1)
|
|
|
|
// The prompt is delivered on stdin (see buildCursorArgs). Write it from its
|
|
// own goroutine so it cannot deadlock against the stdout reader below: a
|
|
// prompt larger than the OS pipe buffer (~64 KiB) blocks mid-write until the
|
|
// child drains it, and the child cannot drain while we are not reading its
|
|
// stdout. Closing stdin is what signals end-of-prompt — cursor-agent reads
|
|
// to EOF — so we always close, on both the success and error paths.
|
|
writeErrCh := make(chan error, 1)
|
|
go func() {
|
|
_, err := io.WriteString(stdin, prompt)
|
|
closeStdin()
|
|
writeErrCh <- err
|
|
}()
|
|
|
|
go func() {
|
|
defer cancel()
|
|
defer close(msgCh)
|
|
defer close(resCh)
|
|
|
|
// Close stdout when the context is cancelled so scanner.Scan() unblocks.
|
|
// Closing stdin too releases a prompt write still blocked on a full pipe
|
|
// (e.g. the child died before draining it), so that goroutine cannot leak.
|
|
go func() {
|
|
<-runCtx.Done()
|
|
closeStdin()
|
|
_ = stdout.Close()
|
|
}()
|
|
|
|
startTime := time.Now()
|
|
configuredModel := strings.TrimSpace(opts.Model)
|
|
var output strings.Builder
|
|
var sessionID string
|
|
finalStatus := "completed"
|
|
var finalError string
|
|
var protocolError string
|
|
resultSeen := false
|
|
resultIsError := false
|
|
resultBytes := 0
|
|
eventCount := 0
|
|
invalidEventCount := 0
|
|
assistantEventCount := 0
|
|
toolUseCount := 0
|
|
// unhandledSubtypeCount tracks thinking/tool_call events whose subtype we
|
|
// don't recognize. They are ignored (never synthesized into a message)
|
|
// and surfaced once as a bounded, content-free diagnostic so an upstream
|
|
// protocol addition is visible instead of silent.
|
|
unhandledSubtypeCount := 0
|
|
// unhandledTypes tracks top-level event types the switch below does not
|
|
// handle. See cursorUnhandledTypeTally: it makes a dropped event
|
|
// observable, and documents what the count does and does not establish
|
|
// (MUL-5434).
|
|
var unhandledTypes cursorUnhandledTypeTally
|
|
lastEventType := "none"
|
|
// assistantBytes counts only model-authored streamed text. It is kept
|
|
// separate from `output` because the result event also writes into
|
|
// `output`, which made the reported last_assistant_bytes equal to
|
|
// result_bytes on a run where the assistant streamed nothing at all —
|
|
// hiding the exact signal needed to tell those two cases apart.
|
|
assistantBytes := 0
|
|
// stepUsage accumulates per-step token counts from "step_finish" events.
|
|
// resultUsage holds authoritative session totals from "result" events.
|
|
// If the result event includes usage, we use resultUsage exclusively;
|
|
// otherwise we fall back to stepUsage.
|
|
stepUsage := make(map[string]TokenUsage)
|
|
resultUsage := make(map[string]TokenUsage)
|
|
hasResultUsage := false
|
|
var thinking cursorThinkingStream
|
|
|
|
scanner := bufio.NewScanner(stdout)
|
|
scanner.Buffer(make([]byte, 0, 1024*1024), 10*1024*1024)
|
|
|
|
for scanner.Scan() {
|
|
raw := scanner.Text()
|
|
line := normalizeCursorStreamLine(raw)
|
|
if line == "" {
|
|
continue
|
|
}
|
|
|
|
var evt cursorStreamEvent
|
|
if err := json.Unmarshal([]byte(line), &evt); err != nil {
|
|
invalidEventCount++
|
|
continue
|
|
}
|
|
eventCount++
|
|
lastEventType = observedCursorEventType(evt.Type)
|
|
|
|
if sid := evt.readSessionID(); sid != "" {
|
|
sessionID = sid
|
|
}
|
|
|
|
switch evt.Type {
|
|
case "system":
|
|
if evt.Subtype == "init" {
|
|
trySend(msgCh, Message{Type: MessageStatus, Status: "running"})
|
|
}
|
|
if evt.Subtype == "error" {
|
|
errMsg := cursorErrorText(&evt)
|
|
if errMsg != "" {
|
|
protocolError = errMsg
|
|
trySend(msgCh, Message{Type: MessageError, Content: errMsg})
|
|
}
|
|
}
|
|
|
|
case "assistant":
|
|
assistantEventCount++
|
|
assistantBytes += b.handleCursorAssistant(&evt, msgCh, &output)
|
|
|
|
case "thinking":
|
|
// Reasoning is a top-level event streamed as deltas, not a
|
|
// content block inside assistant messages. Match subtypes
|
|
// explicitly: only `delta` carries reasoning content (forwarded
|
|
// as it lands so the daemon's 500ms flush shows it mid-run),
|
|
// `completed` closes the block. An unknown subtype is NOT folded
|
|
// into reasoning — silently absorbing upstream additions is the
|
|
// exact failure mode this fix exists to prevent (MUL-5231).
|
|
switch evt.Subtype {
|
|
case "delta":
|
|
if content := thinking.delta(evt.Text); content != "" {
|
|
trySend(msgCh, Message{Type: MessageThinking, Content: content})
|
|
}
|
|
case "completed":
|
|
thinking.complete()
|
|
default:
|
|
unhandledSubtypeCount++
|
|
}
|
|
|
|
case "tool_call":
|
|
// Only the two subtypes that define a call's boundaries drive the
|
|
// transcript: `started` opens it, `completed` closes it. A
|
|
// non-terminal (e.g. a future `progress`) or missing subtype must
|
|
// NOT synthesize a result — that would decrement the daemon's
|
|
// in-flight tool count early and drop a still-running long tool
|
|
// from the tool watchdog onto the shorter idle watchdog, which
|
|
// can force-stop it as falsely stuck (MUL-5231 review).
|
|
switch evt.Subtype {
|
|
case "started":
|
|
call := parseCursorToolCall(&evt)
|
|
toolUseCount++
|
|
trySend(msgCh, Message{
|
|
Type: MessageToolUse,
|
|
Tool: call.Name,
|
|
CallID: call.CallID,
|
|
Input: call.Input,
|
|
})
|
|
case "completed":
|
|
call := parseCursorToolCall(&evt)
|
|
trySend(msgCh, Message{
|
|
Type: MessageToolResult,
|
|
Tool: call.Name,
|
|
CallID: call.CallID,
|
|
Output: call.Result,
|
|
})
|
|
default:
|
|
unhandledSubtypeCount++
|
|
}
|
|
|
|
case "tool_use":
|
|
toolUseCount++
|
|
var params map[string]any
|
|
if evt.Parameters != nil {
|
|
_ = json.Unmarshal(evt.Parameters, ¶ms)
|
|
}
|
|
trySend(msgCh, Message{
|
|
Type: MessageToolUse,
|
|
Tool: evt.ToolName,
|
|
CallID: evt.ToolID,
|
|
Input: params,
|
|
})
|
|
|
|
case "tool_result":
|
|
trySend(msgCh, Message{
|
|
Type: MessageToolResult,
|
|
CallID: evt.ToolID,
|
|
Output: evt.Output,
|
|
})
|
|
|
|
case "result":
|
|
resultSeen = true
|
|
if evt.IsError || evt.Subtype == "error" {
|
|
finalStatus = "failed"
|
|
finalError = cursorErrorText(&evt)
|
|
resultIsError = true
|
|
}
|
|
resultBytes = len(evt.ResultText)
|
|
if evt.ResultText != "" && output.Len() == 0 {
|
|
output.WriteString(evt.ResultText)
|
|
trySend(msgCh, Message{Type: MessageText, Content: evt.ResultText})
|
|
}
|
|
b.accumulateResultUsage(resultUsage, &evt, configuredModel)
|
|
if evt.hasResultUsage() {
|
|
hasResultUsage = true
|
|
}
|
|
// Current Cursor Agent versions can emit the terminal result
|
|
// event but keep a worker process alive. Treat result as the
|
|
// protocol boundary so the daemon can report completion.
|
|
cancel()
|
|
|
|
case "error":
|
|
errMsg := cursorErrorText(&evt)
|
|
if errMsg != "" {
|
|
protocolError = errMsg
|
|
}
|
|
trySend(msgCh, Message{Type: MessageError, Content: errMsg})
|
|
|
|
case "text":
|
|
if evt.Part != nil {
|
|
var part cursorTextPart
|
|
_ = json.Unmarshal(evt.Part, &part)
|
|
if part.Text != "" {
|
|
output.WriteString(part.Text)
|
|
assistantBytes += len(part.Text)
|
|
trySend(msgCh, Message{Type: MessageText, Content: part.Text})
|
|
}
|
|
}
|
|
|
|
case "step_finish":
|
|
if evt.Part != nil {
|
|
var part cursorStepFinishPart
|
|
_ = json.Unmarshal(evt.Part, &part)
|
|
model := cursorUsageModel(evt.Model, configuredModel)
|
|
u := stepUsage[model]
|
|
u.InputTokens += int64(part.Tokens.Input)
|
|
u.OutputTokens += int64(part.Tokens.Output)
|
|
u.CacheReadTokens += int64(part.Tokens.Cache.Read)
|
|
stepUsage[model] = u
|
|
}
|
|
|
|
default:
|
|
// A top-level type this parser does not handle. Falling through
|
|
// here used to be completely silent, and that silence is the
|
|
// bug: if the CLI renames `tool_call` or `thinking`, every tool
|
|
// and reasoning row vanishes while the run still reports
|
|
// success and tool_use_count=0, with no diagnostic to
|
|
// distinguish that from a tool-free run (MUL-5434).
|
|
//
|
|
// Counting makes the drop observable. It does not identify the
|
|
// cause on its own — see cursorUnhandledTypeTally for what a
|
|
// non-zero and a zero tally each do and do not establish.
|
|
//
|
|
// Counting is also all we do. An unrecognized event is never
|
|
// coerced into a tool or reasoning message: guessing at
|
|
// upstream additions is the failure mode MUL-5231 already
|
|
// fixed once.
|
|
if !cursorNonTranscriptEventType(evt.Type) {
|
|
unhandledTypes.observe(evt.Type)
|
|
}
|
|
}
|
|
}
|
|
scanErr := scanner.Err()
|
|
if scanErr != nil {
|
|
// Scanner stopped consuming stdout. Close the pipe before Wait so a
|
|
// child writing a malformed or oversized event cannot deadlock on a
|
|
// full OS pipe; the scanner error remains the primary failure.
|
|
_ = stdout.Close()
|
|
}
|
|
|
|
// Use result usage if available (session totals); otherwise fall back
|
|
// to accumulated step_finish usage.
|
|
if !hasResultUsage {
|
|
resultUsage = stepUsage
|
|
}
|
|
|
|
exitErr := cmd.Wait()
|
|
duration := time.Since(startTime)
|
|
|
|
// Wait has already closed the stdin pipe, so a prompt write still blocked
|
|
// on a full pipe has returned by now; the writer sends exactly once.
|
|
writeErr := <-writeErrCh
|
|
|
|
if resultSeen {
|
|
// A parsed result is the protocol boundary. Ignore the cancellation and
|
|
// exit error caused by stopping a Cursor worker that lingers afterward.
|
|
if finalStatus == "failed" && finalError == "" {
|
|
finalError = "cursor-agent returned an error result without details"
|
|
}
|
|
} else {
|
|
switch {
|
|
case runCtx.Err() == context.DeadlineExceeded:
|
|
finalStatus = "timeout"
|
|
finalError = fmt.Sprintf("cursor-agent timed out after %s", timeout)
|
|
case runCtx.Err() == context.Canceled:
|
|
finalStatus = "aborted"
|
|
finalError = "execution cancelled"
|
|
case scanErr != nil:
|
|
finalStatus = "failed"
|
|
finalError = fmt.Sprintf("cursor-agent stdout read error: %v", scanErr)
|
|
case protocolError != "":
|
|
finalStatus = "failed"
|
|
finalError = protocolError
|
|
case writeErr != nil:
|
|
// Ranked below explicit agent errors because a child that exits
|
|
// early for its own reason (bad auth, bad flag) makes our write
|
|
// fail with EPIPE as a side effect. The stderr tail is appended
|
|
// to every !resultSeen failure below, so the real cause still
|
|
// surfaces either way.
|
|
finalStatus = "failed"
|
|
finalError = fmt.Sprintf("cursor-agent prompt write failed: %v", writeErr)
|
|
case exitErr != nil:
|
|
finalStatus = "failed"
|
|
finalError = fmt.Sprintf("cursor-agent exited with error: %v", exitErr)
|
|
default:
|
|
finalStatus = "failed"
|
|
finalError = "cursor-agent stream ended without terminal result"
|
|
}
|
|
}
|
|
|
|
if finalError != "" {
|
|
finalError = sanitizeAgentDiagnostic(finalError)
|
|
}
|
|
if finalStatus == "failed" && !resultSeen {
|
|
finalError = cursorFailureDiagnostic(
|
|
finalError,
|
|
exitErr,
|
|
scanErr,
|
|
eventCount,
|
|
invalidEventCount,
|
|
lastEventType,
|
|
)
|
|
finalError = withAgentStderr(finalError, "cursor", sanitizeAgentDiagnostic(stderrBuf.Tail()))
|
|
}
|
|
|
|
logStreamProtocolObservation(b.cfg.Logger, streamProtocolObservation{
|
|
provider: "cursor-agent",
|
|
cliVersion: b.cfg.CLIVersion,
|
|
model: opts.Model,
|
|
exitCode: streamProcessExitCode(exitErr),
|
|
eventCount: eventCount,
|
|
invalidEventCount: invalidEventCount,
|
|
assistantEventCount: assistantEventCount,
|
|
toolUseCount: toolUseCount,
|
|
sawResult: resultSeen,
|
|
resultIsError: resultIsError,
|
|
resultBytes: resultBytes,
|
|
lastAssistantBytes: assistantBytes,
|
|
scannerError: scanErr != nil && !resultSeen,
|
|
lastEventType: lastEventType,
|
|
unhandledEventTypeCount: unhandledTypes.total,
|
|
unhandledEventTypes: unhandledTypes.summary(),
|
|
unhandledSubtypeCount: unhandledSubtypeCount,
|
|
})
|
|
|
|
if unhandledSubtypeCount > 0 {
|
|
// Content-free and emitted once per run: signals that the CLI sent a
|
|
// thinking/tool_call subtype we chose to ignore, so a protocol
|
|
// addition is diagnosable without the parser having guessed at it.
|
|
b.cfg.Logger.Warn("cursor-agent ignored unhandled event subtypes", "count", unhandledSubtypeCount)
|
|
}
|
|
|
|
if unhandledTypes.total > 0 {
|
|
// Same contract one level up, and the broader signal of the two: an
|
|
// unhandled top-level type means whole events never reached the
|
|
// transcript. What was in them is not decided here — the type names
|
|
// are the starting point for that, not the conclusion.
|
|
b.cfg.Logger.Warn("cursor-agent ignored unhandled event types",
|
|
"count", unhandledTypes.total,
|
|
"types", unhandledTypes.summary(),
|
|
)
|
|
}
|
|
|
|
b.cfg.Logger.Info("cursor-agent finished", "pid", cmd.Process.Pid, "status", finalStatus, "duration", duration.Round(time.Millisecond).String())
|
|
|
|
finalOutput := output.String()
|
|
if finalStatus != "completed" {
|
|
// A partial transcript is not a final answer. Keep it in Messages for
|
|
// observability, but never expose it as Result.Output on failure.
|
|
finalOutput = ""
|
|
}
|
|
|
|
resCh <- Result{
|
|
Status: finalStatus,
|
|
Output: finalOutput,
|
|
Error: finalError,
|
|
DurationMs: duration.Milliseconds(),
|
|
SessionID: sessionID,
|
|
Usage: resultUsage,
|
|
}
|
|
}()
|
|
|
|
return &Session{Messages: msgCh, Result: resCh}, nil
|
|
}
|
|
|
|
const cursorIncompleteFinalizationWarning = "actions completed before finalization may already have taken effect"
|
|
|
|
func cursorFailureDiagnostic(message string, exitErr, scanErr error, eventCount, invalidEventCount int, lastEventType string) string {
|
|
return fmt.Sprintf(
|
|
"%s (result_seen=false, exit_code=%d, scanner_error=%t, event_count=%d, invalid_event_count=%d, last_event_type=%s); %s",
|
|
message,
|
|
streamProcessExitCode(exitErr),
|
|
scanErr != nil,
|
|
eventCount,
|
|
invalidEventCount,
|
|
lastEventType,
|
|
cursorIncompleteFinalizationWarning,
|
|
)
|
|
}
|
|
|
|
// observedCursorEventType keeps protocol diagnostics bounded and content-free.
|
|
// Event types are identifiers; arbitrary values are collapsed instead of being
|
|
// copied into daemon logs or task errors.
|
|
func observedCursorEventType(value string) string {
|
|
value = strings.TrimSpace(value)
|
|
if value == "" {
|
|
return "unknown"
|
|
}
|
|
if len(value) > 64 {
|
|
return "invalid"
|
|
}
|
|
for _, r := range value {
|
|
if (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z') || (r >= '0' && r <= '9') || r == '_' || r == '-' {
|
|
continue
|
|
}
|
|
return "invalid"
|
|
}
|
|
return value
|
|
}
|
|
|
|
// cursorNonTranscriptEventTypes are top-level events that are known to exist,
|
|
// carry nothing the transcript needs, and are therefore NOT reported as
|
|
// protocol drift. Dropping them is the behaviour that already shipped; listing
|
|
// them here only keeps them out of the diagnostic, because a warning that fires
|
|
// on every task is a warning nobody reads.
|
|
//
|
|
// Provenance differs per entry and matters when revisiting this list:
|
|
// - `user` — the CLI echoing our own prompt back. Confirmed present in the
|
|
// recorded 2026.07.20 stream, i.e. in every real run.
|
|
// - `connection`, `retry` — transport/control frames reported on newer CLI
|
|
// builds (MUL-5434 review). Not reproduced locally, so they are listed
|
|
// defensively: if a build does not emit them the entry is inert, and if it
|
|
// does we must not call a known control frame an unhandled protocol event.
|
|
//
|
|
// An entry here is a claim that the type carries no transcript content. Do not
|
|
// add a type merely to silence the warning.
|
|
var cursorNonTranscriptEventTypes = map[string]struct{}{
|
|
"user": {},
|
|
"connection": {},
|
|
"retry": {},
|
|
}
|
|
|
|
func cursorNonTranscriptEventType(eventType string) bool {
|
|
_, ok := cursorNonTranscriptEventTypes[strings.TrimSpace(eventType)]
|
|
return ok
|
|
}
|
|
|
|
// cursorUnhandledTypeCardinalityCap bounds how many distinct type names the
|
|
// tally retains. A stream emitting novel type names (or garbage that still
|
|
// parses as JSON) must not grow the map without limit inside a long-running
|
|
// daemon; names past the cap collapse into one overflow bucket.
|
|
const cursorUnhandledTypeCardinalityCap = 16
|
|
|
|
// cursorUnhandledTypeOverflowKey holds the counts that exceeded the cardinality
|
|
// cap. The parentheses cannot collide with a real type name: observedCursorEventType
|
|
// only ever returns identifier characters, "unknown", or "invalid".
|
|
const cursorUnhandledTypeOverflowKey = "(overflow)"
|
|
|
|
// cursorUnhandledTypeTally counts top-level event types this parser does not
|
|
// handle, so that dropping them stops being silent. Before it existed a renamed
|
|
// `tool_call` fell through the type switch and the run reported success with
|
|
// tool_use_count=0 and no diagnostic whatsoever (MUL-5434).
|
|
//
|
|
// Read it as evidence, not as a verdict — in both directions:
|
|
//
|
|
// - A non-zero tally shows the stream carried top-level events we do not
|
|
// handle. Which events, and whether they are what the transcript lost,
|
|
// still has to be established from the type names, invalid_event_count and
|
|
// ultimately a captured stream.
|
|
// - A zero tally shows only that no unhandled top-level type was observed. It
|
|
// does NOT establish that the model used no tools: the CLI may have executed
|
|
// tools without handing the updates to its stream serializer at all, a new
|
|
// shape may be nested inside an event type we already recognize (e.g. an
|
|
// assistant content block), or the events may have been lost to invalid
|
|
// framing or a scanner boundary. Those branches are open on #6071.
|
|
//
|
|
// Only normalized identifiers and counts are retained, never payload content,
|
|
// so the tally is safe to log beside the rest of the protocol summary.
|
|
type cursorUnhandledTypeTally struct {
|
|
total int
|
|
counts map[string]int
|
|
}
|
|
|
|
func (t *cursorUnhandledTypeTally) observe(eventType string) {
|
|
t.total++
|
|
if t.counts == nil {
|
|
t.counts = make(map[string]int, 4)
|
|
}
|
|
key := observedCursorEventType(eventType)
|
|
if _, seen := t.counts[key]; !seen && len(t.counts) >= cursorUnhandledTypeCardinalityCap {
|
|
key = cursorUnhandledTypeOverflowKey
|
|
}
|
|
t.counts[key]++
|
|
}
|
|
|
|
// summary renders the tally as a stable, sorted "name=count" list (e.g.
|
|
// "reasoning=9,tool_calls=6") so it can be grepped and diffed across runs.
|
|
func (t *cursorUnhandledTypeTally) summary() string {
|
|
if len(t.counts) == 0 {
|
|
return ""
|
|
}
|
|
keys := make([]string, 0, len(t.counts))
|
|
for key := range t.counts {
|
|
keys = append(keys, key)
|
|
}
|
|
sort.Strings(keys)
|
|
parts := make([]string, 0, len(keys))
|
|
for _, key := range keys {
|
|
parts = append(parts, fmt.Sprintf("%s=%d", key, t.counts[key]))
|
|
}
|
|
return strings.Join(parts, ",")
|
|
}
|
|
|
|
// handleCursorAssistant forwards one assistant event's content blocks and
|
|
// returns how many bytes of model-authored text it appended to output. The
|
|
// caller tracks that separately from output.Len(), which also absorbs the
|
|
// terminal result text.
|
|
func (b *cursorBackend) handleCursorAssistant(evt *cursorStreamEvent, ch chan<- Message, output *strings.Builder) int {
|
|
if evt.Message == nil {
|
|
return 0
|
|
}
|
|
|
|
var content cursorAssistantMessage
|
|
if err := json.Unmarshal(evt.Message, &content); err != nil {
|
|
return 0
|
|
}
|
|
|
|
// Note: per-message usage in assistant events is intentionally ignored.
|
|
// Token usage is taken exclusively from "result" events (session totals)
|
|
// to avoid double-counting.
|
|
|
|
written := 0
|
|
for _, block := range content.Content {
|
|
switch block.Type {
|
|
case "output_text", "text":
|
|
if block.Text != "" {
|
|
output.WriteString(block.Text)
|
|
written += len(block.Text)
|
|
trySend(ch, Message{Type: MessageText, Content: block.Text})
|
|
}
|
|
case "thinking":
|
|
if block.Text != "" {
|
|
trySend(ch, Message{Type: MessageThinking, Content: block.Text})
|
|
}
|
|
case "tool_use":
|
|
var input map[string]any
|
|
if block.Input != nil {
|
|
_ = json.Unmarshal(block.Input, &input)
|
|
}
|
|
trySend(ch, Message{
|
|
Type: MessageToolUse,
|
|
Tool: block.Name,
|
|
CallID: block.ID,
|
|
Input: input,
|
|
})
|
|
}
|
|
}
|
|
return written
|
|
}
|
|
|
|
// cursorThinkingStream turns the CLI's `thinking` event sequence into the
|
|
// content we forward as MessageThinking. Cursor streams a reasoning block as
|
|
// `subtype:"delta"` events carrying a text fragment each, terminated by a
|
|
// `subtype:"completed"` event that carries no text of its own.
|
|
//
|
|
// Deltas are forwarded immediately (the daemon concatenates and flushes them on
|
|
// a ticker, so reasoning shows up while the task runs) and blocks are separated
|
|
// by a blank line, because consecutive blocks would otherwise be glued together
|
|
// into one paragraph by that same concatenation.
|
|
type cursorThinkingStream struct {
|
|
blockOpen bool
|
|
anySent bool
|
|
}
|
|
|
|
// delta returns the content to forward for one `delta` event, or "" when the
|
|
// fragment is empty. The first fragment of a new block is prefixed with a blank
|
|
// line so the daemon's concatenation keeps blocks visually separated.
|
|
func (t *cursorThinkingStream) delta(text string) string {
|
|
if text == "" {
|
|
return ""
|
|
}
|
|
if !t.blockOpen && t.anySent {
|
|
text = "\n\n" + text
|
|
}
|
|
t.blockOpen = true
|
|
t.anySent = true
|
|
return text
|
|
}
|
|
|
|
// complete closes the current reasoning block. The terminal event carries no
|
|
// content of its own in the observed protocol, so nothing is forwarded; the
|
|
// next block's first delta gets the separating blank line.
|
|
func (t *cursorThinkingStream) complete() {
|
|
t.blockOpen = false
|
|
}
|
|
|
|
// cursorToolCall is a top-level `tool_call` event projected onto the fields the
|
|
// daemon transcript needs.
|
|
type cursorToolCall struct {
|
|
Name string
|
|
CallID string
|
|
Input map[string]any
|
|
Result string
|
|
}
|
|
|
|
// cursorToolCallKeySuffix is how Cursor names the per-tool payload: the tool is
|
|
// the key of a nested object rather than a `name` field, so `readToolCall`
|
|
// identifies a read and holds that call's `args` and (once completed) `result`.
|
|
const cursorToolCallKeySuffix = "ToolCall"
|
|
|
|
// parseCursorToolCall extracts tool identity, arguments and result from a
|
|
// `tool_call` event:
|
|
//
|
|
// {"type":"tool_call","subtype":"started","call_id":"call-…",
|
|
// "tool_call":{"readToolCall":{"args":{"path":"/x/a.txt"}},"toolCallId":"call-…"}}
|
|
//
|
|
// A payload we cannot decode still yields the call ID, so started/completed
|
|
// events stay paired and the transcript keeps a row for the call.
|
|
func parseCursorToolCall(evt *cursorStreamEvent) cursorToolCall {
|
|
call := cursorToolCall{CallID: cursorCallID(evt.CallID)}
|
|
|
|
var envelope map[string]json.RawMessage
|
|
if len(evt.ToolCall) == 0 || json.Unmarshal(evt.ToolCall, &envelope) != nil {
|
|
return call
|
|
}
|
|
if call.CallID == "" {
|
|
var nestedID string
|
|
if err := json.Unmarshal(envelope["toolCallId"], &nestedID); err == nil {
|
|
call.CallID = cursorCallID(nestedID)
|
|
}
|
|
}
|
|
|
|
key := cursorToolPayloadKey(envelope)
|
|
if key == "" {
|
|
return call
|
|
}
|
|
call.Name = strings.TrimSuffix(key, cursorToolCallKeySuffix)
|
|
|
|
var payload struct {
|
|
Args map[string]any `json:"args"`
|
|
Result json.RawMessage `json:"result"`
|
|
}
|
|
if err := json.Unmarshal(envelope[key], &payload); err != nil {
|
|
return call
|
|
}
|
|
call.Input = payload.Args
|
|
if len(payload.Result) > 0 {
|
|
call.Result = string(payload.Result)
|
|
}
|
|
return call
|
|
}
|
|
|
|
// cursorToolPayloadKey picks the `<name>ToolCall` key of a tool_call envelope.
|
|
// Observed payloads carry exactly one; the sort keeps the choice deterministic
|
|
// if a future CLI ever emits more than one.
|
|
func cursorToolPayloadKey(envelope map[string]json.RawMessage) string {
|
|
var keys []string
|
|
for key := range envelope {
|
|
if len(key) > len(cursorToolCallKeySuffix) && strings.HasSuffix(key, cursorToolCallKeySuffix) {
|
|
keys = append(keys, key)
|
|
}
|
|
}
|
|
if len(keys) == 0 {
|
|
return ""
|
|
}
|
|
sort.Strings(keys)
|
|
return keys[0]
|
|
}
|
|
|
|
// cursorCallID normalizes the CLI's call identifier. Cursor packs two ids into
|
|
// one newline-separated string ("call-…\nfc_…"); the first line alone is
|
|
// unique, and keeping IDs single-line stops the embedded newline from breaking
|
|
// up daemon log lines.
|
|
func cursorCallID(raw string) string {
|
|
id := strings.TrimSpace(raw)
|
|
if idx := strings.IndexByte(id, '\n'); idx >= 0 {
|
|
id = id[:idx]
|
|
}
|
|
return strings.TrimSpace(id)
|
|
}
|
|
|
|
func cursorUsageModel(evtModel, configuredModel string) string {
|
|
if model := strings.TrimSpace(evtModel); model != "" {
|
|
return model
|
|
}
|
|
if model := strings.TrimSpace(configuredModel); model != "" {
|
|
return model
|
|
}
|
|
return "cursor"
|
|
}
|
|
|
|
func (b *cursorBackend) accumulateResultUsage(usage map[string]TokenUsage, evt *cursorStreamEvent, configuredModel string) {
|
|
model := cursorUsageModel(evt.Model, configuredModel)
|
|
u := usage[model]
|
|
|
|
// Cursor agent has emitted token usage in multiple shapes: top-level
|
|
// camelCase fields, nested camelCase usage, and nested legacy snake_case
|
|
// usage. Prefer top-level result totals when present, otherwise use the
|
|
// nested usage object.
|
|
if evt.InputTokens != 0 || evt.OutputTokens != 0 || evt.CacheReadTokens != 0 || evt.CacheWriteTokens != 0 {
|
|
u.InputTokens += evt.InputTokens
|
|
u.OutputTokens += evt.OutputTokens
|
|
u.CacheReadTokens += evt.CacheReadTokens
|
|
u.CacheWriteTokens += evt.CacheWriteTokens
|
|
} else if evt.Usage != nil {
|
|
u.InputTokens += evt.Usage.InputTokens
|
|
u.OutputTokens += evt.Usage.OutputTokens
|
|
u.CacheReadTokens += evt.Usage.CacheReadInputTokens
|
|
u.CacheWriteTokens += evt.Usage.CacheWriteInputTokens
|
|
} else {
|
|
return
|
|
}
|
|
|
|
usage[model] = u
|
|
}
|
|
|
|
// ── Cursor stream-json types ──
|
|
|
|
type cursorStreamEvent struct {
|
|
Type string `json:"type"`
|
|
Subtype string `json:"subtype,omitempty"`
|
|
SessionID string `json:"session_id,omitempty"`
|
|
Model string `json:"model,omitempty"`
|
|
|
|
// assistant fields
|
|
Message json.RawMessage `json:"message,omitempty"`
|
|
|
|
// thinking fields
|
|
Text string `json:"text,omitempty"`
|
|
|
|
// tool_call fields (current CLI): the tool is a nested key under tool_call
|
|
ToolCall json.RawMessage `json:"tool_call,omitempty"`
|
|
CallID string `json:"call_id,omitempty"`
|
|
|
|
// tool_use fields
|
|
ToolName string `json:"tool_name,omitempty"`
|
|
ToolID string `json:"tool_id,omitempty"`
|
|
Parameters json.RawMessage `json:"parameters,omitempty"`
|
|
|
|
// tool_result fields
|
|
Output string `json:"output,omitempty"`
|
|
|
|
// result fields.
|
|
//
|
|
// The result event reports token counts only; cursor-agent's stream-json
|
|
// carries no per-turn cost (no `total_cost_usd`, no per-step `cost`). This
|
|
// was verified against the real CLI (2026.07.20, both stream-json and json
|
|
// output) and Cursor's CLI docs. So — unlike Grok, whose turn cost xAI
|
|
// reports authoritatively — Cursor spend is estimated downstream from the
|
|
// static rate table, never carried through here. Add a cost field only
|
|
// when a real payload proves the CLI emits one.
|
|
ResultText string `json:"result,omitempty"`
|
|
IsError bool `json:"is_error,omitempty"`
|
|
InputTokens int64 `json:"inputTokens,omitempty"`
|
|
OutputTokens int64 `json:"outputTokens,omitempty"`
|
|
CacheReadTokens int64 `json:"cacheReadTokens,omitempty"`
|
|
CacheWriteTokens int64 `json:"cacheWriteTokens,omitempty"`
|
|
Usage *cursorUsage `json:"usage,omitempty"`
|
|
|
|
// error fields
|
|
ErrorMsg string `json:"error,omitempty"`
|
|
Detail string `json:"detail,omitempty"`
|
|
|
|
// legacy compat
|
|
Part json.RawMessage `json:"part,omitempty"`
|
|
}
|
|
|
|
func (evt *cursorStreamEvent) readSessionID() string {
|
|
if s := strings.TrimSpace(evt.SessionID); s != "" {
|
|
return s
|
|
}
|
|
return ""
|
|
}
|
|
|
|
func (evt *cursorStreamEvent) hasResultUsage() bool {
|
|
return evt.Usage != nil || evt.InputTokens != 0 || evt.OutputTokens != 0 || evt.CacheReadTokens != 0 || evt.CacheWriteTokens != 0
|
|
}
|
|
|
|
type cursorUsage struct {
|
|
InputTokens int64 `json:"input_tokens"`
|
|
OutputTokens int64 `json:"output_tokens"`
|
|
CacheReadInputTokens int64 `json:"cached_input_tokens"`
|
|
CacheWriteInputTokens int64
|
|
}
|
|
|
|
func (u *cursorUsage) UnmarshalJSON(data []byte) error {
|
|
var raw struct {
|
|
InputTokensSnake int64 `json:"input_tokens"`
|
|
InputTokensCamel int64 `json:"inputTokens"`
|
|
OutputTokensSnake int64 `json:"output_tokens"`
|
|
OutputTokensCamel int64 `json:"outputTokens"`
|
|
CachedInputTokensSnake int64 `json:"cached_input_tokens"`
|
|
CachedInputTokensCamel int64 `json:"cachedInputTokens"`
|
|
CacheReadTokensCamel int64 `json:"cacheReadTokens"`
|
|
CacheReadInputTokensSnake int64 `json:"cache_read_input_tokens"`
|
|
CacheReadInputTokensCamel int64 `json:"cacheReadInputTokens"`
|
|
CacheWriteTokensCamel int64 `json:"cacheWriteTokens"`
|
|
CacheCreationInputTokensSnake int64 `json:"cache_creation_input_tokens"`
|
|
CacheCreationInputTokensCamel int64 `json:"cacheCreationInputTokens"`
|
|
}
|
|
if err := json.Unmarshal(data, &raw); err != nil {
|
|
return err
|
|
}
|
|
u.InputTokens = firstNonZeroInt64(raw.InputTokensSnake, raw.InputTokensCamel)
|
|
u.OutputTokens = firstNonZeroInt64(raw.OutputTokensSnake, raw.OutputTokensCamel)
|
|
u.CacheReadInputTokens = firstNonZeroInt64(
|
|
raw.CachedInputTokensSnake,
|
|
raw.CachedInputTokensCamel,
|
|
raw.CacheReadTokensCamel,
|
|
raw.CacheReadInputTokensSnake,
|
|
raw.CacheReadInputTokensCamel,
|
|
)
|
|
u.CacheWriteInputTokens = firstNonZeroInt64(
|
|
raw.CacheWriteTokensCamel,
|
|
raw.CacheCreationInputTokensSnake,
|
|
raw.CacheCreationInputTokensCamel,
|
|
)
|
|
return nil
|
|
}
|
|
|
|
func firstNonZeroInt64(values ...int64) int64 {
|
|
for _, v := range values {
|
|
if v != 0 {
|
|
return v
|
|
}
|
|
}
|
|
return 0
|
|
}
|
|
|
|
type cursorAssistantMessage struct {
|
|
Model string `json:"model"`
|
|
Content []cursorContentBlock `json:"content"`
|
|
Usage *cursorUsage `json:"usage,omitempty"`
|
|
}
|
|
|
|
type cursorContentBlock struct {
|
|
Type string `json:"type"`
|
|
Text string `json:"text,omitempty"`
|
|
ID string `json:"id,omitempty"`
|
|
Name string `json:"name,omitempty"`
|
|
Input json.RawMessage `json:"input,omitempty"`
|
|
}
|
|
|
|
type cursorTextPart struct {
|
|
Text string `json:"text"`
|
|
}
|
|
|
|
// cursorStepFinishPart carries per-step token counts. cursor-agent does not
|
|
// report a per-step cost (see cursorStreamEvent's result-fields note), so only
|
|
// tokens are parsed here.
|
|
type cursorStepFinishPart struct {
|
|
Tokens struct {
|
|
Input int `json:"input"`
|
|
Output int `json:"output"`
|
|
Cache struct {
|
|
Read int `json:"read"`
|
|
} `json:"cache"`
|
|
} `json:"tokens"`
|
|
}
|
|
|
|
// ── Helpers ──
|
|
|
|
// normalizeCursorStreamLine handles the stdout:/stderr: prefix that Cursor
|
|
// CLI may emit in stream-json mode. Returns the trimmed JSON line.
|
|
func normalizeCursorStreamLine(raw string) string {
|
|
trimmed := strings.TrimSpace(raw)
|
|
if trimmed == "" {
|
|
return ""
|
|
}
|
|
// Cursor CLI may prefix lines with "stdout:" or "stderr:" — strip it.
|
|
if idx := cursorStreamPrefixRe.FindStringIndex(trimmed); idx != nil {
|
|
return strings.TrimSpace(trimmed[idx[1]:])
|
|
}
|
|
return trimmed
|
|
}
|
|
|
|
var cursorStreamPrefixRe = regexp.MustCompile(`^(?i)(stdout|stderr)\s*[:=]?\s*`)
|
|
|
|
func cursorErrorText(evt *cursorStreamEvent) string {
|
|
if evt.ErrorMsg != "" {
|
|
return evt.ErrorMsg
|
|
}
|
|
if evt.Detail != "" {
|
|
return evt.Detail
|
|
}
|
|
if evt.ResultText != "" {
|
|
return evt.ResultText
|
|
}
|
|
return ""
|
|
}
|
|
|
|
// cursorBlockedArgs are flags hardcoded by the daemon that must not be
|
|
// overridden by user-configured custom_args. Overriding these would break
|
|
// the daemon↔cursor-agent communication protocol.
|
|
var cursorBlockedArgs = map[string]blockedArgMode{
|
|
"-p": blockedStandalone, // non-interactive print mode
|
|
"--output-format": blockedWithValue, // stream-json protocol
|
|
"--yolo": blockedStandalone, // auto-approval for autonomous operation
|
|
}
|
|
|
|
// buildCursorArgs assembles the argv for a one-shot cursor-agent invocation.
|
|
//
|
|
// Usage: cursor-agent -p --output-format stream-json
|
|
//
|
|
// --workspace <cwd> --yolo [--model <m>] [--resume <id>]
|
|
//
|
|
// The prompt is deliberately NOT part of argv. cursor-agent's -p is a boolean
|
|
// print-mode switch and the prompt is a positional argument; when no positional
|
|
// prompt is present and stdin is not a TTY, the CLI reads stdin to EOF and uses
|
|
// that as the prompt. We rely on that path because putting user-controlled text
|
|
// on the command line is not safe on Windows: the official cursor-agent.cmd/.ps1
|
|
// launcher ends in `& node.exe index.js $args`, and PowerShell re-serialises
|
|
// $args onto node's command line. Under Windows PowerShell 5.1 (and pwsh
|
|
// <= 7.2, which default to Legacy native argument passing) an argument holding
|
|
// embedded double quotes is not re-escaped, so a prompt containing e.g.
|
|
// `go build -ldflags "-X main.version=foo"` gets re-tokenised and `-X` reaches
|
|
// commander.js as a standalone flag: "error: unknown option '-X'" (#5649).
|
|
// Routing the prompt through stdin keeps it off every command line, so no shell
|
|
// or launcher on any platform can re-tokenise it. Only fixed, content-free
|
|
// flags remain in argv.
|
|
func buildCursorArgs(opts ExecOptions, logger *slog.Logger) []string {
|
|
args := []string{
|
|
"-p",
|
|
"--output-format", "stream-json",
|
|
"--yolo",
|
|
}
|
|
if opts.Cwd != "" {
|
|
args = append(args, "--workspace", opts.Cwd)
|
|
}
|
|
if opts.Model != "" {
|
|
args = append(args, "--model", opts.Model)
|
|
}
|
|
// NOTE: cursor-agent CLI does not support --system-prompt or --max-turns.
|
|
// Instructions are injected via AGENTS.md and .cursor/skills/ files instead.
|
|
if opts.ResumeSessionID != "" {
|
|
args = append(args, "--resume", opts.ResumeSessionID)
|
|
}
|
|
args = append(args, filterCustomArgs(opts.CustomArgs, cursorBlockedArgs, logger)...)
|
|
return args
|
|
}
|