mirror of
https://github.com/multica-ai/multica.git
synced 2026-07-27 04:56:20 +02:00
* fix(task): auto-retry transient provider stream cut like chat (MUL-4910) Claude Code's "API Error: Connection closed mid-response" is a transient network cut. In unattended issue runs it fell through to agent_error.unknown / process_failure — neither in retryableReasons — so the task terminated. Interactive chat only appeared resilient because the CLI's own in-process retry usually recovers first; there is no chat-specific retry in Multica. Both paths share the same finalizeStreamResult -> Classify -> retryableReasons pipeline. - classify: route "connection closed" / "mid-response" to agent_error.provider_network, before the exit-status rule so the "exited with error: exit status N" variant also lands here. - retry: add provider_network to retryableReasons. It is resume-safe, so the retry child inherits the session and continues the truncated conversation instead of restarting. Tests cover the new classification (incl. exit-status variant) and the retry/resume flags. Note: mirror the substrings into the MUL-1949 offline backfill SQL to keep in-flight and historical taxonomies aligned. Co-authored-by: multica-agent <github@multica.ai> * feat(task): defer provider_network's final retry ~5s — three-tier (MUL-4910) Follow-up to the immediate-retry fix: make the connection-closed retry a three-tier schedule — first run + immediate retry + one retry deferred ~5s — so a blip that survives the immediate retry gets a short cooldown before the final attempt instead of firing back-to-back. Reuses the existing deferred/fire_at primitive (the comment-routing escalation mechanism) rather than adding new infrastructure: CreateRetryTask gains an optional fire_at; when set, the child is inserted 'deferred' and the existing PromoteDueDeferredTasksForRuntime sweeper — already run promote-first on every claim poll — flips it to 'queued' at fire_at. No migration, no claim-query change, no daemon change. - retryAttemptCeiling: raise provider_network's ceiling to 3 (other reasons keep max_attempts=2); applied in both retryEligible and MaybeRetryFailedTask's budget pre-check so the primary and sweeper paths agree. - retryDelayForAttempt: only provider_network's final attempt is deferred (5s); every other retry — including provider_network's first — stays immediate. - FailTask + MaybeRetryFailedTask pass fire_at and skip the queued broadcast/notify for a deferred child (promotion emits them at fire time). Timing: the deferred child fires on the first claim poll at/after fire_at, so >= 5s; on an otherwise-idle runtime it can stretch to the poll interval — the same behaviour deferred escalations already have. Tests: pure schedule/eligibility coverage (TestProviderNetworkRetrySchedule) plus a DB test asserting fire_at controls deferred-vs-queued, attempt=3, and resume-safety. Co-authored-by: multica-agent <github@multica.ai> * fix(task): persist reason-aware retry budget; respect max_attempts=1 disable (MUL-4910) Addresses the pre-merge review must-fix: retryAttemptCeiling raised provider_network to 3 unconditionally, which (1) overrode the max_attempts=1 "auto-retry disabled" contract and (2) persisted a self-contradictory child (attempt=3, max_attempts=2) that leaks to the task API, so a naive attempt < max_attempts consumer would misjudge the budget. - retryAttemptCeiling now returns taskMaxAttempts unchanged when it is <= 1 (disabled stays disabled) and only ever WIDENS otherwise (max(col, 3) for provider_network) — a higher configured budget is kept. - CreateRetryTask takes an optional max_attempts; FailTask and MaybeRetryFailedTask write the reason-aware effective ceiling into the child so the whole retry chain self-describes (attempt=3, max_attempts=3). NULL inherits the parent column, so non-provider_network reasons are unchanged. Tests: - pure: ceiling widens to 3, keeps a higher budget, and never revives a disabled (max_attempts=1) task; eligibility rejects the disabled case. - DB (end-to-end FailTask): default budget → deferred final child at attempt=3/max_attempts=3; first failure → immediate child; max_attempts=1 → no child. Plus CreateRetryTask persists the passed budget / inherits on NULL. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai>
257 lines
9.6 KiB
Go
257 lines
9.6 KiB
Go
package taskfailure
|
|
|
|
import (
|
|
"regexp"
|
|
"strings"
|
|
)
|
|
|
|
// providerHTTP5xxRe matches a 3-digit number starting with 5 (5xx HTTP
|
|
// status code) that isn't surrounded by other digits. Mirrors the SQL
|
|
// regex `(^|[^0-9])5[0-9][0-9]([^0-9]|$)` from MUL-1949 — keeps phrases
|
|
// like "1500ms" or "1.5.0" from accidentally landing in
|
|
// provider_server_error.
|
|
//
|
|
// Compiled at package init: the classifier is on the in-flight write
|
|
// path for every failed task, so paying the regex compile cost at
|
|
// startup rather than per-call matters.
|
|
var providerHTTP5xxRe = regexp.MustCompile(`(^|[^0-9])5[0-9][0-9]([^0-9]|$)`)
|
|
|
|
// httpAuthCodeRe / httpQuotaCodeRe / httpCapacityCodeRe match specific 3-digit
|
|
// HTTP status codes only when they are NOT embedded in a longer number, using
|
|
// the same digit-boundary guard as providerHTTP5xxRe. Without this guard the
|
|
// bare substrings "401"/"402"/"403"/"429"/"529" fire on unrelated numbers —
|
|
// e.g. "402913 tokens", "15290ms", "exit status 4030" — misclassifying process
|
|
// or unknown failures as provider billing / rate-limit errors. That pollutes
|
|
// failure observability: a genuine process crash gets filed under a provider
|
|
// bucket, masking the real cause on failure dashboards. (A misfire here still
|
|
// can't cause a spurious retry: the auth / quota / capacity buckets these
|
|
// regexes guard are all non-retryable. The only agent_error.* reason on
|
|
// internal/service/task.go's retryableReasons allowlist is provider_network
|
|
// — MUL-4910 — and these regexes never route into it.) The 5xx bucket was
|
|
// already anchored for exactly this reason (MUL-1949); these codes were not.
|
|
var (
|
|
httpAuthCodeRe = regexp.MustCompile(`(^|[^0-9])(401|403)([^0-9]|$)`)
|
|
httpQuotaCodeRe = regexp.MustCompile(`(^|[^0-9])402([^0-9]|$)`)
|
|
httpCapacityCodeRe = regexp.MustCompile(`(^|[^0-9])(429|529)([^0-9]|$)`)
|
|
)
|
|
|
|
// Classify maps a free-form error string from the agent runtime / CLI
|
|
// to one of the 14 agent_error.* sub-reasons. Always returns a valid
|
|
// Reason; falls back to ReasonAgentUnknown when no rule matches and for
|
|
// empty input.
|
|
//
|
|
// The rule order mirrors the SQL CASE expression in MUL-1949
|
|
// (db-boy's offline backfill query). The SQL is the source of truth:
|
|
// when the two diverge, this Go classifier is wrong and should be
|
|
// updated to match. Keeping them in lock-step is required so that
|
|
// in-flight rows and historically backfilled rows share the same
|
|
// taxonomy.
|
|
//
|
|
// Matching is case-insensitive substring against the lowercased input.
|
|
// More-specific rules come before more-generic ones (e.g.
|
|
// context_overflow before provider_quota_limit, because "token limit"
|
|
// would otherwise be claimed by the quota bucket via "limit").
|
|
//
|
|
// Why a substring classifier rather than structured error codes: the
|
|
// 11 backend wrappers (server/pkg/agent/*) all surface upstream API
|
|
// failures verbatim in Result.Error, often as `API Error: 400 {...}`
|
|
// or as raw stderr tails. Insisting on structured codes would require
|
|
// touching every backend; a string classifier gives us refined
|
|
// failure_reason today and lets per-backend structured upgrades land
|
|
// independently.
|
|
func Classify(rawError string) Reason {
|
|
trimmed := strings.TrimSpace(rawError)
|
|
if trimmed == "" {
|
|
// SQL maps NULL/empty to a separate bucket ("empty_error"),
|
|
// but that bucket is not part of the canonical 21. In-flight
|
|
// callers should never hand us empty input — if they do, the
|
|
// safest landing is the catchall.
|
|
return ReasonAgentUnknown
|
|
}
|
|
lower := strings.ToLower(trimmed)
|
|
|
|
switch {
|
|
// 1. Context / token window overflow. Checked early so "token
|
|
// limit" doesn't get swallowed by the broader "limit" / "quota"
|
|
// rule below.
|
|
case containsAny(lower,
|
|
"context length",
|
|
"context_length_exceeded",
|
|
"maximum context",
|
|
"prompt is too long",
|
|
"context size has been exceeded",
|
|
),
|
|
// SQL had `%token%limit%` — ILIKE wildcard between tokens. We
|
|
// approximate with both substrings present, which catches
|
|
// "token limit", "tokens per minute limit", etc., without the
|
|
// false positives a naive `Contains("token") || Contains("limit")`
|
|
// would generate.
|
|
strings.Contains(lower, "token") && strings.Contains(lower, "limit"):
|
|
return ReasonAgentContextOverflow
|
|
|
|
// 2. Missing config / API key. Checked before auth because
|
|
// "missing API key" partly overlaps with "invalid api key"
|
|
// wording but is structurally a config error, not an auth
|
|
// rejection.
|
|
case strings.Contains(lower, "missing environment variable"),
|
|
strings.Contains(lower, "missing") && strings.Contains(lower, "api_key"),
|
|
strings.Contains(lower, "api key") && strings.Contains(lower, "required"),
|
|
strings.Contains(lower, "no llm provider configured"),
|
|
strings.Contains(lower, "no provider configured"):
|
|
return ReasonAgentMissingConfig
|
|
|
|
// 3. Auth / access. 401 / 403 / "Not logged in" / invalid token
|
|
// / lacks access to the model. Status codes use a digit boundary
|
|
// so "4030" / "1401ms" don't spuriously land here.
|
|
case httpAuthCodeRe.MatchString(lower),
|
|
containsAny(lower,
|
|
"unauthorized",
|
|
"login required",
|
|
"not logged in",
|
|
"please login again",
|
|
"refresh token",
|
|
"invalid api key",
|
|
"access token",
|
|
"subscription access",
|
|
"does not have access",
|
|
"you may not have access",
|
|
):
|
|
return ReasonAgentProviderAuthOrAccess
|
|
|
|
// 4. Quota / billing. 402 / insufficient balance / monthly usage
|
|
// limit / credits exhausted.
|
|
case httpQuotaCodeRe.MatchString(lower),
|
|
containsAny(lower,
|
|
"insufficient_balance",
|
|
"balance is too low",
|
|
"monthly usage limit",
|
|
"usage limit",
|
|
"you've hit your limit",
|
|
// Curly apostrophe variant: providers and copy-pasted error
|
|
// strings sometimes use U+2019 instead of ASCII '. SQL ILIKE
|
|
// would not match the curly form either, so this is a small
|
|
// in-flight improvement on top of the SQL classifier.
|
|
"you\u2019ve hit your limit",
|
|
"credits",
|
|
"quota",
|
|
):
|
|
return ReasonAgentProviderQuotaLimit
|
|
|
|
// 5. Capacity / rate limit. 429 / 529 / overloaded / rate limit.
|
|
case httpCapacityCodeRe.MatchString(lower),
|
|
containsAny(lower,
|
|
"rate limit",
|
|
"overloaded",
|
|
"no capacity available",
|
|
):
|
|
return ReasonAgentProviderCapacityOrRateLimit
|
|
|
|
// 6. Provider 5xx / server error. The 5xx regex is checked here
|
|
// rather than as plain string matches because the SQL uses an
|
|
// anchored regex — see providerHTTP5xxRe's docstring.
|
|
case containsAny(lower,
|
|
"server had an error",
|
|
"provider returned error",
|
|
"internal error",
|
|
"service unavailable",
|
|
"bad gateway",
|
|
),
|
|
providerHTTP5xxRe.MatchString(lower):
|
|
return ReasonAgentProviderServerError
|
|
|
|
// 7. Provider network. Stream cut, dial failures, DNS / I/O
|
|
// timeout below the HTTP layer. "connection closed" / "mid-response"
|
|
// catch the Claude Code CLI's mid-stream disconnect
|
|
// ("API Error: Connection closed mid-response. ...") so a transient cut
|
|
// lands in the retryable provider_network bucket (with session resume)
|
|
// instead of falling through to agent_error.unknown / process_failure
|
|
// and terminating the task (MUL-4910). Checked before rule 13 so the
|
|
// "... exited with error: exit status N ..." variant still routes here.
|
|
// Mirror these substrings into the MUL-1949 offline backfill SQL.
|
|
case containsAny(lower,
|
|
"stream disconnected",
|
|
"connection closed",
|
|
"mid-response",
|
|
"error sending request",
|
|
"unable to connect",
|
|
"dial tcp",
|
|
"connection refused",
|
|
"connectionrefused",
|
|
"dns",
|
|
"i/o timeout",
|
|
):
|
|
return ReasonAgentProviderNetwork
|
|
|
|
// 8. Model not found / unavailable. The SQL uses `%model%not%found%`,
|
|
// which matches "model … not found" with anything in between;
|
|
// we approximate with both substrings present, which captures
|
|
// typical phrasings like "model X not found" and "the requested
|
|
// model was not found".
|
|
case strings.Contains(lower, "model") && strings.Contains(lower, "not found"),
|
|
containsAny(lower,
|
|
"unknown model",
|
|
"selected model",
|
|
"http 404",
|
|
"404 page not found",
|
|
):
|
|
return ReasonAgentModelNotFoundOrUnavailable
|
|
|
|
// 9. Empty / unparseable output from the agent CLI itself. These
|
|
// strings come from server/pkg/agent/*.go wrappers and are
|
|
// stable.
|
|
case containsAny(lower,
|
|
"returned empty output",
|
|
"returned no parseable output",
|
|
):
|
|
return ReasonAgentEmptyOrUnparseableOutput
|
|
|
|
// 10. Agent subprocess hard timeout (per-task wall clock).
|
|
case strings.Contains(lower, "timed out after"):
|
|
return ReasonAgentTimeout
|
|
|
|
// 11. Runner CLI binary missing.
|
|
case strings.Contains(lower, "executable not found"):
|
|
return ReasonAgentRuntimeMissingExecutable
|
|
|
|
// 12. Runner CLI version too old / incompatible protocol.
|
|
case containsAny(lower,
|
|
"below the minimum supported version",
|
|
"requires a newer version",
|
|
):
|
|
return ReasonAgentRuntimeVersionUnsupported
|
|
|
|
// 13. Agent / runner process-level failure. Checked last among
|
|
// specific rules because "exit status" / "signal" can co-occur
|
|
// with more specific upstream errors that SHOULD win (e.g. an
|
|
// agent that crashed *because* the provider rate-limited it
|
|
// should be classified as rate-limited, not as a process
|
|
// failure).
|
|
case containsAny(lower,
|
|
"exit status",
|
|
"signal",
|
|
"panic",
|
|
"sigsegv",
|
|
"process exited",
|
|
"pipe has been ended",
|
|
"file already closed",
|
|
"initialize failed",
|
|
):
|
|
return ReasonAgentProcessFailure
|
|
}
|
|
|
|
return ReasonAgentUnknown
|
|
}
|
|
|
|
// containsAny reports whether s contains any of the supplied substrings.
|
|
// Caller is responsible for lowercasing s ahead of time so the helper
|
|
// stays cheap on the hot path — pre-lowercasing once is faster than
|
|
// case-folding inside each substring scan.
|
|
func containsAny(s string, subs ...string) bool {
|
|
for _, sub := range subs {
|
|
if strings.Contains(s, sub) {
|
|
return true
|
|
}
|
|
}
|
|
return false
|
|
}
|