Files
multica/server/internal/daemon/execenv/runtime_config_kind_test.go
Bohan Jiang 4eb80c8b77 fix(daemon): keep the runtime brief byte-stable across triggers (MUL-5377) (#6021)
* fix(daemon): keep the runtime brief byte-stable across triggers (MUL-5377)

Claude Code loads the runtime brief (CLAUDE.md / AGENTS.md) into messages[0],
ahead of the entire conversation. A cache breakpoint is all-or-nothing, so a
single differing byte there invalidates the prompt cache for the whole history
on every `--resume`. InjectRuntimeConfig rewrites that file on every run and
interpolated nine per-run values into it, so in practice the cache was thrown
away on the first comment that landed on any issue.

Measured on one issue over three runs: run 1 (cold) spent 89.9k cache-write
tokens building 105k of context; runs 2 and 3 each spent ~425k re-creating a
prefix they should have read. 842k of 946.5k cache-write tokens (89%) went into
re-creation, with only tools[]+system[] surviving each resume (a constant
18,085 tokens both times).

Fix: the brief now carries only what is stable for the lifetime of a resumed
session, and per-run state travels in the per-turn user message, which is
appended after the cached prefix.

- Merge kindCommentTriggered + kindAssignmentTriggered into kindIssue, and stop
  reading TriggerCommentID in classifyTask. The brief can no longer diverge by
  trigger type structurally, rather than by convention.
- Replace writeWorkflowComment/writeWorkflowAssignment with one writeWorkflowIssue
  that routes on the per-turn message. The mode-specific status rules live inside
  their own mode block, so "own the status arc" and "do not touch the status"
  can never be read as unarbitrated peers.
- Move Task Initiator, Session Continuity Notice and Connected Apps out of the
  brief into BuildPrompt via BuildTaskInitiatorBlock / SessionContinuityNotice /
  BuildConnectedAppsBlock.
- Drop TriggerCommentID, TriggerThreadID, NewCommentsSince, NewCommentCount,
  PriorSessionResumed and CommentReplyTargets from the brief; BuildPrompt already
  emitted all six from the same helpers, so this is de-duplication.
- Set PriorSessionResumeUnavailable on `task` as well as `taskCtx` in both local
  resume gates, or the notice would silently vanish on exactly the failure path
  it exists to disclose.

Tests: TestInjectRuntimeConfigByteIdenticalAcrossTriggers renders the brief
across nine per-run variants (trigger type, differing comment/thread ids, resume
delta, resume-unavailable, cross-thread fan-out, member/agent initiator,
connected apps) for two providers and requires bytes.Equal, with a non-vacuity
guard so it cannot pass on a function that ignores its input. Daemon-side tests
assert the moved sections still reach the agent through the per-turn prompt.

Co-authored-by: multica-agent <github@multica.ai>

* fix(daemon): route the issue workflow on an explicit turn-mode marker

Review follow-up on the mode router. The brief said Reply mode applies when
the per-turn message "opens with a [NEW COMMENT] block", but buildCommentPrompt
writes two paragraphs before that block and only emits it when
TriggerCommentContent is non-empty. Two ways to get it wrong:

- The message never literally opens with the block, so the router's own
  wording did not match the prompt it describes.
- A comment-triggered run with an empty comment body — or an older server that
  does not send one — emitted no block at all. An agent following the brief
  would fall through to Ownership mode and change the issue status on a turn
  whose rule is "do NOT change the issue status".

BuildPrompt now emits an unconditional `**Turn mode: Reply.**` /
`**Turn mode: Ownership.**` line from the same branches it uses to pick a code
path, and the brief routes on that marker. Brief and prompt can no longer
disagree about the mode, because the value that selects the path also states
it. The router also names a safe fallback (treat an unlabelled turn as Reply
mode and leave the status alone).

Tests: TestTurnModeMarkerAlwaysPresent covers comment-triggered with and
without comment content, plus both assignment shapes;
TestTurnModeMarkerAbsentOnIssuelessKinds keeps the marker off chat /
quick-create / autopilot; TestBriefModeRouterMatchesPromptMarkers fails if the
brief ever describes a marker the prompt does not emit.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-28 16:00:16 +08:00

252 lines
9.3 KiB
Go

package execenv
import (
"strings"
"testing"
)
// TestClassifyTask pins the precedence rule on classifyTask. All four
// kinds plus tiebreak cases for safety.
func TestClassifyTask(t *testing.T) {
t.Parallel()
cases := []struct {
name string
ctx TaskContextForEnv
want taskKind
}{
{"chat", TaskContextForEnv{ChatSessionID: "c"}, kindChat},
{"quick-create", TaskContextForEnv{QuickCreatePrompt: "p"}, kindQuickCreate},
{"autopilot", TaskContextForEnv{AutopilotRunID: "r"}, kindAutopilotRunOnly},
{"issue-comment-triggered", TaskContextForEnv{IssueID: "i", TriggerCommentID: "c"}, kindIssue},
{"issue-assignment-triggered", TaskContextForEnv{IssueID: "i"}, kindIssue},
{"issue-bare", TaskContextForEnv{}, kindIssue},
{"tiebreak-chat-vs-quick", TaskContextForEnv{ChatSessionID: "c", QuickCreatePrompt: "p"}, kindChat},
{"tiebreak-quick-vs-autopilot", TaskContextForEnv{QuickCreatePrompt: "p", AutopilotRunID: "r"}, kindQuickCreate},
{"tiebreak-autopilot-vs-comment", TaskContextForEnv{AutopilotRunID: "r", IssueID: "i", TriggerCommentID: "c"}, kindAutopilotRunOnly},
}
for _, tc := range cases {
tc := tc
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
if got := classifyTask(tc.ctx); got != tc.want {
t.Errorf("classifyTask: got %d, want %d", got, tc.want)
}
})
}
}
// TestTaskKindHasIssueContext pins the predicate that gates Project
// Context / Issue Metadata / Sub-issue Creation in the slim dispatcher.
func TestTaskKindHasIssueContext(t *testing.T) {
t.Parallel()
cases := []struct {
kind taskKind
want bool
}{
{kindIssue, true},
{kindAutopilotRunOnly, false},
{kindQuickCreate, false},
{kindChat, false},
}
for _, tc := range cases {
if got := tc.kind.hasIssueContext(); got != tc.want {
t.Errorf("kind=%d hasIssueContext: got %v, want %v", tc.kind, got, tc.want)
}
}
}
// TestBuildMetaSkillContentBriefContent pins that buildMetaSkillContent
// renders the (now sole) brief: the `issue get` one-liner is present and
// the retired legacy verbose description is not.
func TestBuildMetaSkillContentBriefContent(t *testing.T) {
t.Parallel()
out := buildMetaSkillContent("claude", TaskContextForEnv{
IssueID: "issue-1",
TriggerCommentID: "comment-1",
AgentName: "Eve",
AgentID: "eve-1",
})
if !strings.Contains(out, "- `multica issue get <id> --output json` — full issue.\n") {
t.Errorf("brief is missing the `issue get` one-liner\n---\n%s", out)
}
if strings.Contains(out, "Get full issue details.") {
t.Errorf("brief still carries the retired legacy `issue get` description\n---\n%s", out)
}
}
// TestBuildMetaSkillContentSlimKindMatrix locks in which sections the
// slim brief emits per task kind, machine-checking the matrix documented
// on `buildMetaSkillContentSlim`. Heading is matched as a discrete line
// (preceded by newline + followed by newline) so inline references like
// "see ## Comment Formatting" do not trip the absence assertions.
func TestBuildMetaSkillContentSlimKindMatrix(t *testing.T) {
baseRepo := []RepoContextForEnv{{URL: "https://example.com/x.git", Description: "x"}}
baseSkill := []SkillContextForEnv{{Name: "skill-x", Description: "x"}}
type sectionCheck struct {
heading string
mustHave map[taskKind]bool
}
allKinds := map[taskKind]bool{
kindIssue: true, kindAutopilotRunOnly: true,
kindQuickCreate: true, kindChat: true,
}
issueKinds := map[taskKind]bool{kindIssue: true}
checks := []sectionCheck{
{"# Multica Agent Runtime", allKinds},
{"## Background Task Safety", allKinds},
{"## Agent Identity", allKinds},
{"## Available Commands", allKinds},
{"### Workflow", allKinds},
{"## Important: Always Use the `multica` CLI", allKinds},
{"## Output", allKinds},
{"## Comment Formatting", issueKinds},
{"## Repositories", map[taskKind]bool{
kindIssue: true, kindAutopilotRunOnly: true, kindChat: true,
}},
{"## Issue Metadata", issueKinds},
{"## Instruction Precedence", issueKinds},
{"## Sub-issue Creation", issueKinds},
{"## Skills", map[taskKind]bool{
kindIssue: true, kindAutopilotRunOnly: true, kindChat: true,
}},
{"## Mentions", issueKinds},
{"## Attachments", issueKinds},
}
fixtures := map[taskKind]TaskContextForEnv{
kindChat: {ChatSessionID: "c-1", AgentName: "Eve", AgentID: "eve-1",
Repos: baseRepo, AgentSkills: baseSkill},
kindQuickCreate: {QuickCreatePrompt: "p", AgentName: "Eve", AgentID: "eve-1",
Repos: baseRepo, AgentSkills: baseSkill},
kindAutopilotRunOnly: {AutopilotRunID: "r-1", AgentName: "Eve", AgentID: "eve-1",
Repos: baseRepo, AgentSkills: baseSkill},
kindIssue: {IssueID: "i-1", AgentName: "Eve", AgentID: "eve-1",
Repos: baseRepo, AgentSkills: baseSkill},
}
for kind, ctx := range fixtures {
out := buildMetaSkillContent("claude", ctx)
for _, c := range checks {
needle := "\n" + c.heading + "\n"
firstLine := c.heading + "\n"
present := strings.HasPrefix(out, firstLine) || strings.Contains(out, needle)
want := c.mustHave[kind]
if want && !present {
t.Errorf("kind=%d: expected heading %q in slim brief", kind, c.heading)
}
if !want && present {
t.Errorf("kind=%d: heading %q should NOT be in slim brief (matrix gating regression)", kind, c.heading)
}
}
}
}
// TestSlimQuickCreateAvailableCommands locks the minimal-variant content
// for quick-create's Available Commands: `issue create` present, every
// other Core command absent (the hard guardrails forbid the call).
func TestSlimQuickCreateAvailableCommands(t *testing.T) {
out := buildMetaSkillContent("codex", TaskContextForEnv{
QuickCreatePrompt: "create an issue about flaky tests",
AgentName: "Eve", AgentID: "eve-1",
})
for _, want := range []string{
"## Available Commands",
"multica issue create --title",
"`multica --help`",
} {
if !strings.Contains(out, want) {
t.Errorf("quick_create slim Available Commands missing %q", want)
}
}
for _, banned := range []string{
"multica issue get <id>",
"multica issue comment list <issue-id>",
"multica issue update <id>",
"multica issue status <id> <status>",
"multica issue comment add <issue-id>",
"multica issue metadata list <issue-id>",
"multica issue metadata set <issue-id>",
"multica issue metadata delete <issue-id>",
"multica issue children <id>",
"multica repo checkout <url>",
"### Squad maintenance",
"multica squad member set-role",
} {
if strings.Contains(out, banned) {
t.Errorf("quick_create slim Available Commands should NOT advertise %q (hard guardrails forbid the call)", banned)
}
}
}
// TestBackgroundTaskSafetySlimHardPins asserts the slim brief carries the
// same hardened Background Task Safety pins as the legacy brief (MUL-4140).
// The verbose path is covered by
// TestInjectRuntimeConfigBackgroundTaskSafetyProviderAgnostic; this locks
// the compressed slim path so a future slim-brief trim can't quietly drop
// the no-background-and-yield / no-"standing by" guardrails that address
// the MUL-4091 mechanism.
func TestBackgroundTaskSafetySlimHardPins(t *testing.T) {
out := buildMetaSkillContent("claude", TaskContextForEnv{
IssueID: "i-1", TriggerCommentID: "tc-1",
AgentName: "Eve", AgentID: "eve-1",
})
for _, want := range []string{
"## Background Task Safety",
"Do NOT end your turn while background tasks",
"wait for a future notification/reminder",
"run the work synchronously instead",
"Never background-and-yield",
"foreground tool call that blocks",
// MUL-5274: an explicitly requested persistent local service is a
// completed handoff, not unfinished run-owned work. Pin the narrow
// exception and its readiness / cleanup / honesty requirements.
"persistent service handoff",
"running service itself is the requested deliverable",
"stdio redirected to durable logs",
"PID/profile",
"verify readiness before replying",
"survival as best-effort, not guaranteed",
"does not cover tests, builds, CI polling",
"are not agent-owned background tasks",
"GitHub Actions after a successful push",
"Do not wait for them by default",
// MUL-5223 pins: named tool-shape bans, merge requirements
// denied as acceptance criteria, replacement hand-off phrasing,
// and the scoped escape hatch that keeps an explicitly requested
// CI result both permitted and executable.
"do NOT run `gh pr checks --watch`",
"any sleep / retry loop that polls check status",
"NOT your delivery acceptance criteria",
"CI running: <PR link>",
"unless the explicit exception below applies",
"The one exception",
"ONE foreground blocking call (`gh pr checks <pr> --watch`)",
"running in the background so you can keep working",
"standing by",
} {
if !strings.Contains(out, want) {
t.Errorf("slim Background Task Safety missing hardened pin %q\n---\n%s", want, out)
}
}
// `gh run watch` may only appear as a banned command, never as the
// section's example of how to wait properly.
if strings.Contains(out, "e.g. `gh run watch`") {
t.Errorf("slim Background Task Safety should not suggest waiting for external GitHub CI\n---\n%s", out)
}
// MUL-5274 review: with the persistent-service exception in the list, a
// "The rules above ..." scoping sentence would sweep in work that is
// precisely no longer run-owned after handoff.
if strings.Contains(out, "The rules above") {
t.Errorf("slim Background Task Safety must not reintroduce the ambiguous \"The rules above\" scoping sentence\n---\n%s", out)
}
}