Files
Multica Eve aa349fed02 MUL-5637 fix(mcp): agent mcp_config is an authoritative allowlist (#6292)
* fix(mcp): treat agent mcp_config as an authoritative allowlist

An agent's saved mcp_config was silently widened with the runtime host's
own user-level MCP servers, so an explicitly empty `{"mcpServers":{}}`
resolved to the COMPLETE host set instead of no servers at all — the
opposite of what the operator configured (GitHub #6283).

`--strict-mcp-config` was being passed correctly; the merge happened
before it, in the daemon, so strict mode constrained an already-widened
set. Introduced by #5277 and present in v0.4.16 through main.

Restore the three-state contract in resolveEffectiveMcpConfig:

  null / unset        -> inherit the provider's native MCP configuration
  {"mcpServers":{}}   -> strict empty, no host servers
  non-empty object    -> strict allowlist, exactly those servers

Two explicit inherit paths keep the additive behaviour reachable without
weakening the default:

- runtime_config.mcp.inherit_runtime = true opts an agent back in.
- The claim response now carries mcp_config_overlay_only so the daemon
  can tell an agent-authored config from a per-task Composio overlay.
  Without it, enabling an integration on an agent that never configured
  MCP would have stripped the host servers it was already inheriting.

Both decode paths fail closed: malformed runtime_config never enables
inheritance, and a failed runtime merge falls back to the agent's own
config.

The web MCP tab and the `agent create/update --mcp-config` help text
described the old additive behaviour, which is how a tightened config
could look correct while exposing every host server; both now state
which mode is in effect.

Note for rollout: the fix lives in the daemon, so self-hosted users must
upgrade the daemon — a server/UI upgrade alone does not apply it.

Co-authored-by: multica-agent <github@multica.ai>

* fix(mcp): close review gaps in the authoritative mcp_config change

Addresses the four must-fix findings from review of #6292.

1. Deleting the last managed server no longer widens access.
   removeManagedMcpServer cleared the config to null, which now means
   "inherit the host's MCP servers" — so a delete took the agent from one
   allowed server to every server on the host. It now leaves an explicit
   `{"mcpServers":{}}`. Restoring inheritance moved to a separate
   clearManagedMcpConfig action behind its own confirmation that states the
   widening. The delete dialog no longer claims "Runtime servers are not
   affected", which was the opposite of the truth.

2. The UI no longer promises a boundary an old daemon does not enforce.
   The strict semantics live in the daemon, so a config saved against an
   older daemon is not yet in effect. Adds the authoritative-mcp-v1 daemon
   capability:

   - The daemon advertises it and reports authoritative_mcp on the
     runtime-capabilities response.
   - The claim path fails closed: a managed, non-inheriting mcp_config
     claimed by a daemon without the capability cancels the task and
     returns 412 with an actionable message, instead of letting that daemon
     merge the host's servers in. runtime_config.mcp.inherit_runtime is the
     documented escape hatch, and it is honest — it declares that the
     operator accepts the host's servers.
   - The MCP tab shows "needs upgrade" rather than "Not exposed" while the
     bound runtime lacks the capability.

3. Saving OpenClaw settings no longer drops the inherit opt-in.
   parseOpenclawRuntimeConfig discarded unknown keys and the tab persisted
   the result as the whole runtime_config, so one unrelated routing save
   silently deleted mcp.inherit_runtime. Unknown keys now round-trip
   through OpenclawRuntimeConfig.passthrough, excluded from the dirty check
   so they cannot make the form look edited.

4. Documents the new semantics in the built-in creating-agents skill and
   its source map: the three states, the persisted
   runtime_config.mcp.inherit_runtime field, and the claim-time capability
   gate.

Also corrects the PR's rollout claim: there is no database migration, but
this does add a persisted JSON field and change the meaning of an existing
one.

Co-authored-by: multica-agent <github@multica.ai>

* fix(mcp): stop the authoritative-daemon gate from blocking valid claims

CI's backend job failed three handler claim tests with the new 412. Two
distinct problems, both real:

1. The gate fired for a non-object mcp_config. 66 handler fixtures seed
   `[]`, which is not a valid MCP config and cannot carry `mcpServers`, so
   it expresses no boundary to protect. An old daemon does not widen it
   either: mergeRuntimeAndAgentMcpConfig fails to unmarshal a non-object
   and falls back to the agent config alone (verified directly). Gating
   these blocked tasks with no security benefit, so the gate now requires a
   JSON object.

2. The shared daemon test-request helper advertised no capabilities, so
   every claim test was accidentally simulating a pre-#6283 daemon. It now
   defaults authoritative-mcp-v1 on, matching what every current daemon
   sends. Only that capability — skill-bundles / coalesced-comments / rpc
   are feature negotiations whose absence tests real legacy behaviour, so
   they stay opt-in per test.

Adds claim-level coverage for the gate itself, which is what the unit tests
alone could not catch: an outdated daemon gets 412 with an actionable
message and the task is cancelled; a capability-advertising daemon gets
200; the inherit_runtime opt-in lets an outdated daemon through; and an
unmanaged or non-object config is never gated.

Verified against a real migrated schema this time (throwaway Postgres),
which is how the three failures were reproduced locally and confirmed
fixed: `go test ./internal/handler ./internal/daemon` both ok.

Co-authored-by: multica-agent <github@multica.ai>

* fix(mcp): surface the daemon-upgrade refusal and stop gating safe providers

Addresses the second review round on #6292.

1. The refusal is now visible wherever the operator looks. The default
   claim path is the machine-level BATCH endpoint, which skips build
   failures and still answers 200 {"tasks":[]}, so the previous bare
   CancelTask showed a task that vanished with no stated reason — turning an
   explicit upgrade requirement into an unexplained failure. The claim path
   now fails the task with a new classified reason,
   mcp_config_daemon_outdated, plus the actionable message. That reaches the
   user on all three claim paths and on any daemon version, which a new
   response field could not: the audience is by definition a daemon too old
   to read one. The per-runtime path keeps its 412.

   The reason is deliberately not auto-retryable — the same outdated daemon
   would claim the retry and fail it again.

2. The gate no longer cancels safe tasks. It applied to every provider, but
   only claude / codebuddy / codex / cursor / opencode / openclaw were ever
   merged with host MCP by an old daemon (loadRuntimeMcpServerConfigs).
   Qwen was never merged and already had strict semantics, so its tasks were
   being failed for a risk that does not exist. Scoped via
   providersOldDaemonsMergedRuntimeMcp; an unknown provider does not gate,
   because the gate should only fire where the old behaviour is concrete.

3. The new authoritative_mcp flag now goes through the API schema layer.
   Both local-skills responses were returning raw network JSON, so the flag
   that decides whether the UI may assert an MCP boundary rested on an
   unchecked type assertion. Adds RuntimeLocalSkillListRequestSchema with
   authoritative_mcp and mcp_supported defaulting to FALSE — the fail-closed
   direction — and a MALFORMED_ fallback that cannot express a guarantee.

Claim-level tests now cover all three paths, which is what the previous
helper-only tests missed: per-runtime 412, batch recording the refusal on
the task while still delivering the healthy tasks in the same batch, WS RPC
refusing and accepting, the qwen negative case, the inherit_runtime escape
hatch, and unmanaged / non-object configs.

Verified against a real migrated schema (throwaway Postgres):
go test ./internal/handler ./internal/daemon ./pkg/agent ./pkg/taskfailure
./internal/service all ok.

Co-authored-by: multica-agent <github@multica.ai>

* fix(mcp): register the new failure reason and wire its copy into the UI

Addresses the third review round on #6292.

1. mcp_config_daemon_outdated was declared but never registered in
   taskfailure.allReasons, so metrics.NormalizeFailureReason missed the
   known-value map and fell through to free-text Classify() — relabelling a
   platform-side refusal as `agent_error.unknown` (verified directly) and
   leaving the Prometheus series un-pre-warmed. Registered it, canonical
   count 22 → 23 (platform 8 → 9), with the wire value and IsAgentError split
   pinned. New test pins the WHOLE canonical set through
   NormalizeFailureReason so forgetting the next reason fails a test instead
   of quietly mislabelling a metric; NormalizeFailureReason had no coverage
   at all before.

2. The upgrade copy was dead. The locale strings landed last round but
   neither consumer mapped the reason: chatFailureCopy fell back to generic
   failure text with the actionable detail buried in the collapsed raw
   error, and task-failure.ts rendered the bare wire value
   `mcp_config_daemon_outdated` in the agent activity list and issue
   execution log. Both are mapped now, with regression tests, plus the
   runtime class pinned in failure-class.test.ts. This directly contradicted
   the claim in the claim-path comment that every path reaches the user, so
   that is now actually true.

3. providersOldDaemonsMergedRuntimeMcp is documented as what it is: a FROZEN
   record of what pre-capability daemons merged, not a mirror of the daemon's
   current provider switch. The old "keep the two lists in lockstep" note was
   actively harmful advice — runtime MCP discovery for a new provider can only
   ship in a daemon that already advertises the capability (never gated), so
   adding it here would fail tasks on old daemons that never merged for it,
   re-creating the qwen false-positive. Pinned with a test.

Also corrects a stale count in task-failure.ts (7 → 9 platform reasons).

Verified against a real migrated schema (throwaway Postgres): full backend
suite green apart from the pre-existing environmental cmd/multica guard; all
9 TestMcpGate_* integration tests pass.

Co-authored-by: multica-agent <github@multica.ai>

* docs(taskfailure): correct taxonomy counts and finish the reason registration

Non-blocking nits from the fourth review round on #6292.

- Taxonomy counts now say 23 reasons / 9 platform-side. Registering
  mcp_config_daemon_outdated last round updated the assertions but not the
  prose. Swept the whole repo rather than only the flagged lines, which
  turned up four more that were already stale at 21 and drifted further:
  handler/dashboard.go, daemon/poisoned.go, core/types/agent.ts, and the
  db/queries/task_usage.sql comment sqlc copies into the generated file.

  The generated file's comment was updated by hand to match its source.
  Running `sqlc generate` churned 58 lines across 47 unrelated files — the
  local sqlc version differs from the one that produced the checked-in
  output — so that churn was reverted and only the one intended line kept.

- failure_test.go's `required` list now includes
  ReasonMcpConfigDaemonOutdated. Length and label assertions already covered
  the reason, but the list is documented as the complete canonical set, so
  the omission contradicted its own comment.

- Restored the line break in chat-message-list.test.tsx that a previous edit
  of mine collapsed.

Comment, test-fixture and formatting only; no behaviour change.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Eve <eve@multica-ai.local>
Co-authored-by: multica-agent <github@multica.ai>
2026-08-03 16:06:07 +08:00
..