mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-06 01:50:14 +02:00
* fix(daemon): keep the runtime brief byte-stable across triggers (MUL-5377) Claude Code loads the runtime brief (CLAUDE.md / AGENTS.md) into messages[0], ahead of the entire conversation. A cache breakpoint is all-or-nothing, so a single differing byte there invalidates the prompt cache for the whole history on every `--resume`. InjectRuntimeConfig rewrites that file on every run and interpolated nine per-run values into it, so in practice the cache was thrown away on the first comment that landed on any issue. Measured on one issue over three runs: run 1 (cold) spent 89.9k cache-write tokens building 105k of context; runs 2 and 3 each spent ~425k re-creating a prefix they should have read. 842k of 946.5k cache-write tokens (89%) went into re-creation, with only tools[]+system[] surviving each resume (a constant 18,085 tokens both times). Fix: the brief now carries only what is stable for the lifetime of a resumed session, and per-run state travels in the per-turn user message, which is appended after the cached prefix. - Merge kindCommentTriggered + kindAssignmentTriggered into kindIssue, and stop reading TriggerCommentID in classifyTask. The brief can no longer diverge by trigger type structurally, rather than by convention. - Replace writeWorkflowComment/writeWorkflowAssignment with one writeWorkflowIssue that routes on the per-turn message. The mode-specific status rules live inside their own mode block, so "own the status arc" and "do not touch the status" can never be read as unarbitrated peers. - Move Task Initiator, Session Continuity Notice and Connected Apps out of the brief into BuildPrompt via BuildTaskInitiatorBlock / SessionContinuityNotice / BuildConnectedAppsBlock. - Drop TriggerCommentID, TriggerThreadID, NewCommentsSince, NewCommentCount, PriorSessionResumed and CommentReplyTargets from the brief; BuildPrompt already emitted all six from the same helpers, so this is de-duplication. - Set PriorSessionResumeUnavailable on `task` as well as `taskCtx` in both local resume gates, or the notice would silently vanish on exactly the failure path it exists to disclose. Tests: TestInjectRuntimeConfigByteIdenticalAcrossTriggers renders the brief across nine per-run variants (trigger type, differing comment/thread ids, resume delta, resume-unavailable, cross-thread fan-out, member/agent initiator, connected apps) for two providers and requires bytes.Equal, with a non-vacuity guard so it cannot pass on a function that ignores its input. Daemon-side tests assert the moved sections still reach the agent through the per-turn prompt. Co-authored-by: multica-agent <github@multica.ai> * fix(daemon): route the issue workflow on an explicit turn-mode marker Review follow-up on the mode router. The brief said Reply mode applies when the per-turn message "opens with a [NEW COMMENT] block", but buildCommentPrompt writes two paragraphs before that block and only emits it when TriggerCommentContent is non-empty. Two ways to get it wrong: - The message never literally opens with the block, so the router's own wording did not match the prompt it describes. - A comment-triggered run with an empty comment body — or an older server that does not send one — emitted no block at all. An agent following the brief would fall through to Ownership mode and change the issue status on a turn whose rule is "do NOT change the issue status". BuildPrompt now emits an unconditional `**Turn mode: Reply.**` / `**Turn mode: Ownership.**` line from the same branches it uses to pick a code path, and the brief routes on that marker. Brief and prompt can no longer disagree about the mode, because the value that selects the path also states it. The router also names a safe fallback (treat an unlabelled turn as Reply mode and leave the status alone). Tests: TestTurnModeMarkerAlwaysPresent covers comment-triggered with and without comment content, plus both assignment shapes; TestTurnModeMarkerAbsentOnIssuelessKinds keeps the marker off chat / quick-create / autopilot; TestBriefModeRouterMatchesPromptMarkers fails if the brief ever describes a marker the prompt does not emit. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai>
252 lines
9.3 KiB
Go
252 lines
9.3 KiB
Go
package execenv
|
|
|
|
import (
|
|
"strings"
|
|
"testing"
|
|
)
|
|
|
|
// TestClassifyTask pins the precedence rule on classifyTask. All four
|
|
// kinds plus tiebreak cases for safety.
|
|
func TestClassifyTask(t *testing.T) {
|
|
t.Parallel()
|
|
cases := []struct {
|
|
name string
|
|
ctx TaskContextForEnv
|
|
want taskKind
|
|
}{
|
|
{"chat", TaskContextForEnv{ChatSessionID: "c"}, kindChat},
|
|
{"quick-create", TaskContextForEnv{QuickCreatePrompt: "p"}, kindQuickCreate},
|
|
{"autopilot", TaskContextForEnv{AutopilotRunID: "r"}, kindAutopilotRunOnly},
|
|
{"issue-comment-triggered", TaskContextForEnv{IssueID: "i", TriggerCommentID: "c"}, kindIssue},
|
|
{"issue-assignment-triggered", TaskContextForEnv{IssueID: "i"}, kindIssue},
|
|
{"issue-bare", TaskContextForEnv{}, kindIssue},
|
|
{"tiebreak-chat-vs-quick", TaskContextForEnv{ChatSessionID: "c", QuickCreatePrompt: "p"}, kindChat},
|
|
{"tiebreak-quick-vs-autopilot", TaskContextForEnv{QuickCreatePrompt: "p", AutopilotRunID: "r"}, kindQuickCreate},
|
|
{"tiebreak-autopilot-vs-comment", TaskContextForEnv{AutopilotRunID: "r", IssueID: "i", TriggerCommentID: "c"}, kindAutopilotRunOnly},
|
|
}
|
|
for _, tc := range cases {
|
|
tc := tc
|
|
t.Run(tc.name, func(t *testing.T) {
|
|
t.Parallel()
|
|
if got := classifyTask(tc.ctx); got != tc.want {
|
|
t.Errorf("classifyTask: got %d, want %d", got, tc.want)
|
|
}
|
|
})
|
|
}
|
|
}
|
|
|
|
// TestTaskKindHasIssueContext pins the predicate that gates Project
|
|
// Context / Issue Metadata / Sub-issue Creation in the slim dispatcher.
|
|
func TestTaskKindHasIssueContext(t *testing.T) {
|
|
t.Parallel()
|
|
cases := []struct {
|
|
kind taskKind
|
|
want bool
|
|
}{
|
|
{kindIssue, true},
|
|
{kindAutopilotRunOnly, false},
|
|
{kindQuickCreate, false},
|
|
{kindChat, false},
|
|
}
|
|
for _, tc := range cases {
|
|
if got := tc.kind.hasIssueContext(); got != tc.want {
|
|
t.Errorf("kind=%d hasIssueContext: got %v, want %v", tc.kind, got, tc.want)
|
|
}
|
|
}
|
|
}
|
|
|
|
// TestBuildMetaSkillContentBriefContent pins that buildMetaSkillContent
|
|
// renders the (now sole) brief: the `issue get` one-liner is present and
|
|
// the retired legacy verbose description is not.
|
|
func TestBuildMetaSkillContentBriefContent(t *testing.T) {
|
|
t.Parallel()
|
|
|
|
out := buildMetaSkillContent("claude", TaskContextForEnv{
|
|
IssueID: "issue-1",
|
|
TriggerCommentID: "comment-1",
|
|
AgentName: "Eve",
|
|
AgentID: "eve-1",
|
|
})
|
|
|
|
if !strings.Contains(out, "- `multica issue get <id> --output json` — full issue.\n") {
|
|
t.Errorf("brief is missing the `issue get` one-liner\n---\n%s", out)
|
|
}
|
|
if strings.Contains(out, "Get full issue details.") {
|
|
t.Errorf("brief still carries the retired legacy `issue get` description\n---\n%s", out)
|
|
}
|
|
}
|
|
|
|
// TestBuildMetaSkillContentSlimKindMatrix locks in which sections the
|
|
// slim brief emits per task kind, machine-checking the matrix documented
|
|
// on `buildMetaSkillContentSlim`. Heading is matched as a discrete line
|
|
// (preceded by newline + followed by newline) so inline references like
|
|
// "see ## Comment Formatting" do not trip the absence assertions.
|
|
func TestBuildMetaSkillContentSlimKindMatrix(t *testing.T) {
|
|
|
|
baseRepo := []RepoContextForEnv{{URL: "https://example.com/x.git", Description: "x"}}
|
|
baseSkill := []SkillContextForEnv{{Name: "skill-x", Description: "x"}}
|
|
|
|
type sectionCheck struct {
|
|
heading string
|
|
mustHave map[taskKind]bool
|
|
}
|
|
allKinds := map[taskKind]bool{
|
|
kindIssue: true, kindAutopilotRunOnly: true,
|
|
kindQuickCreate: true, kindChat: true,
|
|
}
|
|
issueKinds := map[taskKind]bool{kindIssue: true}
|
|
checks := []sectionCheck{
|
|
{"# Multica Agent Runtime", allKinds},
|
|
{"## Background Task Safety", allKinds},
|
|
{"## Agent Identity", allKinds},
|
|
{"## Available Commands", allKinds},
|
|
{"### Workflow", allKinds},
|
|
{"## Important: Always Use the `multica` CLI", allKinds},
|
|
{"## Output", allKinds},
|
|
{"## Comment Formatting", issueKinds},
|
|
{"## Repositories", map[taskKind]bool{
|
|
kindIssue: true, kindAutopilotRunOnly: true, kindChat: true,
|
|
}},
|
|
{"## Issue Metadata", issueKinds},
|
|
{"## Instruction Precedence", issueKinds},
|
|
{"## Sub-issue Creation", issueKinds},
|
|
{"## Skills", map[taskKind]bool{
|
|
kindIssue: true, kindAutopilotRunOnly: true, kindChat: true,
|
|
}},
|
|
{"## Mentions", issueKinds},
|
|
{"## Attachments", issueKinds},
|
|
}
|
|
|
|
fixtures := map[taskKind]TaskContextForEnv{
|
|
kindChat: {ChatSessionID: "c-1", AgentName: "Eve", AgentID: "eve-1",
|
|
Repos: baseRepo, AgentSkills: baseSkill},
|
|
kindQuickCreate: {QuickCreatePrompt: "p", AgentName: "Eve", AgentID: "eve-1",
|
|
Repos: baseRepo, AgentSkills: baseSkill},
|
|
kindAutopilotRunOnly: {AutopilotRunID: "r-1", AgentName: "Eve", AgentID: "eve-1",
|
|
Repos: baseRepo, AgentSkills: baseSkill},
|
|
kindIssue: {IssueID: "i-1", AgentName: "Eve", AgentID: "eve-1",
|
|
Repos: baseRepo, AgentSkills: baseSkill},
|
|
}
|
|
|
|
for kind, ctx := range fixtures {
|
|
out := buildMetaSkillContent("claude", ctx)
|
|
for _, c := range checks {
|
|
needle := "\n" + c.heading + "\n"
|
|
firstLine := c.heading + "\n"
|
|
present := strings.HasPrefix(out, firstLine) || strings.Contains(out, needle)
|
|
want := c.mustHave[kind]
|
|
if want && !present {
|
|
t.Errorf("kind=%d: expected heading %q in slim brief", kind, c.heading)
|
|
}
|
|
if !want && present {
|
|
t.Errorf("kind=%d: heading %q should NOT be in slim brief (matrix gating regression)", kind, c.heading)
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
// TestSlimQuickCreateAvailableCommands locks the minimal-variant content
|
|
// for quick-create's Available Commands: `issue create` present, every
|
|
// other Core command absent (the hard guardrails forbid the call).
|
|
func TestSlimQuickCreateAvailableCommands(t *testing.T) {
|
|
|
|
out := buildMetaSkillContent("codex", TaskContextForEnv{
|
|
QuickCreatePrompt: "create an issue about flaky tests",
|
|
AgentName: "Eve", AgentID: "eve-1",
|
|
})
|
|
|
|
for _, want := range []string{
|
|
"## Available Commands",
|
|
"multica issue create --title",
|
|
"`multica --help`",
|
|
} {
|
|
if !strings.Contains(out, want) {
|
|
t.Errorf("quick_create slim Available Commands missing %q", want)
|
|
}
|
|
}
|
|
|
|
for _, banned := range []string{
|
|
"multica issue get <id>",
|
|
"multica issue comment list <issue-id>",
|
|
"multica issue update <id>",
|
|
"multica issue status <id> <status>",
|
|
"multica issue comment add <issue-id>",
|
|
"multica issue metadata list <issue-id>",
|
|
"multica issue metadata set <issue-id>",
|
|
"multica issue metadata delete <issue-id>",
|
|
"multica issue children <id>",
|
|
"multica repo checkout <url>",
|
|
"### Squad maintenance",
|
|
"multica squad member set-role",
|
|
} {
|
|
if strings.Contains(out, banned) {
|
|
t.Errorf("quick_create slim Available Commands should NOT advertise %q (hard guardrails forbid the call)", banned)
|
|
}
|
|
}
|
|
}
|
|
|
|
// TestBackgroundTaskSafetySlimHardPins asserts the slim brief carries the
|
|
// same hardened Background Task Safety pins as the legacy brief (MUL-4140).
|
|
// The verbose path is covered by
|
|
// TestInjectRuntimeConfigBackgroundTaskSafetyProviderAgnostic; this locks
|
|
// the compressed slim path so a future slim-brief trim can't quietly drop
|
|
// the no-background-and-yield / no-"standing by" guardrails that address
|
|
// the MUL-4091 mechanism.
|
|
func TestBackgroundTaskSafetySlimHardPins(t *testing.T) {
|
|
|
|
out := buildMetaSkillContent("claude", TaskContextForEnv{
|
|
IssueID: "i-1", TriggerCommentID: "tc-1",
|
|
AgentName: "Eve", AgentID: "eve-1",
|
|
})
|
|
|
|
for _, want := range []string{
|
|
"## Background Task Safety",
|
|
"Do NOT end your turn while background tasks",
|
|
"wait for a future notification/reminder",
|
|
"run the work synchronously instead",
|
|
"Never background-and-yield",
|
|
"foreground tool call that blocks",
|
|
// MUL-5274: an explicitly requested persistent local service is a
|
|
// completed handoff, not unfinished run-owned work. Pin the narrow
|
|
// exception and its readiness / cleanup / honesty requirements.
|
|
"persistent service handoff",
|
|
"running service itself is the requested deliverable",
|
|
"stdio redirected to durable logs",
|
|
"PID/profile",
|
|
"verify readiness before replying",
|
|
"survival as best-effort, not guaranteed",
|
|
"does not cover tests, builds, CI polling",
|
|
"are not agent-owned background tasks",
|
|
"GitHub Actions after a successful push",
|
|
"Do not wait for them by default",
|
|
// MUL-5223 pins: named tool-shape bans, merge requirements
|
|
// denied as acceptance criteria, replacement hand-off phrasing,
|
|
// and the scoped escape hatch that keeps an explicitly requested
|
|
// CI result both permitted and executable.
|
|
"do NOT run `gh pr checks --watch`",
|
|
"any sleep / retry loop that polls check status",
|
|
"NOT your delivery acceptance criteria",
|
|
"CI running: <PR link>",
|
|
"unless the explicit exception below applies",
|
|
"The one exception",
|
|
"ONE foreground blocking call (`gh pr checks <pr> --watch`)",
|
|
"running in the background so you can keep working",
|
|
"standing by",
|
|
} {
|
|
if !strings.Contains(out, want) {
|
|
t.Errorf("slim Background Task Safety missing hardened pin %q\n---\n%s", want, out)
|
|
}
|
|
}
|
|
// `gh run watch` may only appear as a banned command, never as the
|
|
// section's example of how to wait properly.
|
|
if strings.Contains(out, "e.g. `gh run watch`") {
|
|
t.Errorf("slim Background Task Safety should not suggest waiting for external GitHub CI\n---\n%s", out)
|
|
}
|
|
// MUL-5274 review: with the persistent-service exception in the list, a
|
|
// "The rules above ..." scoping sentence would sweep in work that is
|
|
// precisely no longer run-owned after handoff.
|
|
if strings.Contains(out, "The rules above") {
|
|
t.Errorf("slim Background Task Safety must not reintroduce the ambiguous \"The rules above\" scoping sentence\n---\n%s", out)
|
|
}
|
|
}
|