mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-06 10:50:54 +02:00
* feat(agents): remove custom_env from agent resources, add audited env endpoint (MUL-2600)
The agent resource shape (list / get / create / update / archive /
restore responses + WebSocket events) no longer carries `custom_env`
values. Reads/writes of env now flow exclusively through a dedicated
`/api/agents/{id}/env` endpoint that is owner/admin-only, rejects
agent-actor sessions, applies a "****" sentinel preserve guard on
PUT, and writes a persistent audit row per reveal/update.
Why
- `multica agent list --output json` historically returned plaintext
`custom_env` for owner/admin callers (the redaction gate gave only
members the masked map). Any agent token running on the workspace
inherits its owner's role and could read every other agent's
secrets just by listing.
- Patching list/get redaction alone (PR #3175 direction) left
symmetric leaks via mutation responses, WS events, the "reveal"
path itself (no actor-aware auth), and a `****` overwrite footgun
on UpdateAgent.
What changed
- Backend: drop `custom_env` from AgentResponse; add coarse
`has_custom_env` + `custom_env_key_count`. Strip env handling from
UpdateAgent (silently ignored if sent). Keep CreateAgent's
custom_env acceptance.
- Backend: new GET/PUT `/api/agents/{id}/env` handlers in
`internal/handler/agent_env.go`:
- resolveActor → 403 for agent actors (closes the lateral-movement
path).
- Owner/admin role gate via existing helper.
- PUT honours value == "****" as "preserve existing value".
- Both write to `activity_log` with `agent_env_revealed` /
`agent_env_updated` actions. Audit details record key names only,
never values.
- Daemon claim path (`ClaimAgentTask`) unchanged — `TaskAgentData`
still carries plaintext env for runtime injection.
- SQL: new `UpdateAgentCustomEnv` query; sqlc regenerated (v1.31.1).
- CLI: new `multica agent env get|set` subcommands. `--custom-env*`
flags removed from `multica agent update`; the no-fields error
now points to the new path.
- Frontend: drop env fields from `Agent` + `UpdateAgentRequest`; add
`getAgentEnv` / `updateAgentEnv` client methods; rewrite env-tab
to show "N variables configured" + explicit "Reveal & edit"
button, fetching values only on intentional reveal.
- Locales: parity-safe additions to en + zh-Hans.
- Docs: agents-create.{mdx,zh.mdx} reflect the new threat model and
endpoint.
- Mobile: schema drops `custom_env` / `custom_env_redacted`, adds
metadata fields.
Tests
- Handler tests pinned the new invariants: no env in list/get
responses, owner reveal happy-path + audit row, agent-actor 403,
`****` sentinel preserves real values, UpdateAgent silently
ignores `custom_env`, pure `mergeAgentEnv` cases.
- CLI tests pivot to the new flag surface: `agent update` MUST NOT
expose the env flags; `agent env set` MUST expose
--custom-env-stdin/--custom-env-file.
- Frontend test fixtures updated; pnpm typecheck / test / lint
pass cleanly.
This is a breaking API change. Scripts that read `custom_env` from
`/api/agents` must migrate to `GET /api/agents/{id}/env`.
Co-authored-by: multica-agent <github@multica.ai>
* fix(agents): close actor-spoofing + audit fail-closed in env endpoints (MUL-2600)
Addresses Elon's review of #3209:
* Mint a task-scoped `mat_` token per claim, bound to (agent, task,
workspace, owner). Daemon injects it into the agent process in place
of its own credential. Auth middleware authoritatively rebuilds
X-User-ID / X-Agent-ID / X-Task-ID from the token row and sets
X-Actor-Source=task_token; that header is server-set only — incoming
values are stripped before any auth branch runs. resolveActor honors
the header so an agent that strips X-Agent-ID / X-Task-ID still
resolves as actor=agent.
* GetAgentEnv / UpdateAgentEnv are now fail-closed on audit-log
failures: GET refuses to return plaintext, PUT persists inside the
same tx as the audit row so they commit/roll back together.
* PUT /api/agents/{id} returns 400 when the body carries custom_env
instead of silently dropping it — directs callers to the audited env
endpoint.
* Agent actors never see mcp_config, even when the underlying member
is owner/admin; mutation broadcasts go through a redaction shim so
WS subscribers don't pick it up either.
* Fix backend test that asserted dense JSON (jsonb::text renders
whitespace) and frontend test that assumed a unique "Test User"
match.
Co-authored-by: multica-agent <github@multica.ai>
* fix(agents): close residual MUL-2600 gaps from review (MUL-2600)
Migration 108 FK now correctly references agent_task_queue(id) instead
of the non-existent agent_task table; the previous name blocked CI
backend migrations.
Task-token-authenticated requests can no longer be re-routed at a
different workspace by passing workspace_slug / workspace_id /
?workspace_id / a URL workspace param. ResolveWorkspaceIDFromRequest
and resolveWorkspaceUUID both short-circuit on X-Actor-Source=task_token
and return only the token-bound X-Workspace-ID; buildMiddleware adds a
defence-in-depth 403 if any URL-resolved workspace disagrees with the
token binding.
mcp_config no longer leaks back to agent actors through UpdateAgent /
CreateAgent / ArchiveAgent / RestoreAgent HTTP responses — the same
redactAgentResponseForActor helper that GetAgent/ListAgents use is now
applied to mutation responses too. WS broadcasts were already redacted
via broadcastAgentResponse.
FailTask and every TaskService cancel path (CancelTask /
CancelTasksForIssue / CancelTasksForAgent / CancelTasksByTriggerComment
/ BroadcastCancelledTasks) now eagerly DeleteTaskTokensByTask so the
mat_ token's 24h window doesn't outlive a terminated task. Failure is
non-fatal — the FK cascade and expiry remain durable guards.
Doc-only: clarify that PUT /api/agents/{id} now hard-rejects bodies
that carry custom_env (was previously "silently ignores").
Tests:
- middleware: TestResolveWorkspaceIDFromRequest gains a task_token
case asserting client-supplied slug/id/query cannot override the
bound workspace.
- handler: TestUpdateAgent_RedactsMcpConfigForAgentActor and
TestUpdateAgent_KeepsMcpConfigForMemberActor pin the mutation-
response redaction contract per actor type.
Co-authored-by: multica-agent <github@multica.ai>
* fix(agents): match redacted mcp_config as JSON null, not Go nil (MUL-2600)
`AgentResponse.McpConfig` is `json.RawMessage` without `omitempty`, so
the redacted response serialises as `"mcp_config": null`. On decode,
`json.RawMessage` keeps the literal bytes `null` rather than collapsing
to Go nil, which made the assertion fire on a non-leak.
The product contract (field always present, distinguished from "no
config" via `mcp_config_redacted`) is intentional, so adjust the test
to check for "no secret-bearing content" instead of weakening the
contract via `omitempty`.
Co-authored-by: multica-agent <github@multica.ai>
---------
Co-authored-by: multica-agent <github@multica.ai>
190 lines
6.9 KiB
Go
190 lines
6.9 KiB
Go
package middleware
|
|
|
|
import (
|
|
"context"
|
|
"log/slog"
|
|
"net/http"
|
|
"strings"
|
|
"time"
|
|
|
|
"github.com/golang-jwt/jwt/v5"
|
|
"github.com/jackc/pgx/v5/pgtype"
|
|
"github.com/multica-ai/multica/server/internal/auth"
|
|
"github.com/multica-ai/multica/server/internal/util"
|
|
db "github.com/multica-ai/multica/server/pkg/db/generated"
|
|
)
|
|
|
|
func uuidToString(u pgtype.UUID) string { return util.UUIDToString(u) }
|
|
|
|
// Auth middleware validates JWT tokens or Personal Access Tokens.
|
|
// Token sources (in priority order):
|
|
// 1. Authorization: Bearer <token> header (PAT or JWT)
|
|
// 2. multica_auth HttpOnly cookie (JWT) — requires valid CSRF token for state-changing requests
|
|
//
|
|
// Sets X-User-ID and X-User-Email headers on the request for downstream handlers.
|
|
//
|
|
// patCache is optional; when non-nil, PAT lookups are cached with a short
|
|
// TTL (auth.AuthCacheTTL). On cache hit the middleware skips both the DB
|
|
// SELECT and the last_used_at UPDATE — last_used_at is therefore refreshed
|
|
// at most once per TTL window per token, not per request.
|
|
func Auth(queries *db.Queries, patCache *auth.PATCache) func(http.Handler) http.Handler {
|
|
return func(next http.Handler) http.Handler {
|
|
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
|
// X-Actor-Source is server-set only — any value supplied by
|
|
// the client is untrusted and discarded before the auth
|
|
// branches run. Only the mat_ branch below re-sets it. This
|
|
// is what prevents a client from sending a normal mul_ PAT
|
|
// plus a forged `X-Actor-Source: member` (or anything else)
|
|
// to convince a downstream handler that its request came
|
|
// from a non-task-token path.
|
|
r.Header.Del("X-Actor-Source")
|
|
|
|
tokenString, fromCookie := extractToken(r)
|
|
if tokenString == "" {
|
|
slog.Debug("auth: no token found", "path", r.URL.Path)
|
|
http.Error(w, `{"error":"missing authorization"}`, http.StatusUnauthorized)
|
|
return
|
|
}
|
|
|
|
// Cookie-based auth requires CSRF validation for state-changing methods.
|
|
if fromCookie && !auth.ValidateCSRF(r) {
|
|
slog.Debug("auth: CSRF validation failed", "path", r.URL.Path)
|
|
http.Error(w, `{"error":"CSRF validation failed"}`, http.StatusForbidden)
|
|
return
|
|
}
|
|
|
|
// Agent task token: "mat_" prefix. Minted by the server at
|
|
// task-claim time and injected by the daemon into the agent
|
|
// process. Authoritative for actor identity — the bound
|
|
// (user_id, agent_id, task_id, workspace_id) triple is
|
|
// written into request headers here, OVERRIDING whatever the
|
|
// client sent, so a downstream actor-resolver cannot be
|
|
// tricked by a client that strips or forges X-Agent-ID /
|
|
// X-Task-ID. Owner-only endpoints (e.g. agent env
|
|
// management) reject requests authenticated this way; see
|
|
// `actorSourceFromRequest`. MUL-2600.
|
|
if strings.HasPrefix(tokenString, "mat_") {
|
|
if queries == nil {
|
|
http.Error(w, `{"error":"invalid token"}`, http.StatusUnauthorized)
|
|
return
|
|
}
|
|
hash := auth.HashToken(tokenString)
|
|
tt, err := queries.GetTaskTokenByHash(r.Context(), hash)
|
|
if err != nil {
|
|
slog.Warn("auth: invalid task token", "path", r.URL.Path, "error", err)
|
|
http.Error(w, `{"error":"invalid token"}`, http.StatusUnauthorized)
|
|
return
|
|
}
|
|
r.Header.Set("X-User-ID", uuidToString(tt.UserID))
|
|
r.Header.Set("X-Agent-ID", uuidToString(tt.AgentID))
|
|
r.Header.Set("X-Task-ID", uuidToString(tt.TaskID))
|
|
r.Header.Set("X-Workspace-ID", uuidToString(tt.WorkspaceID))
|
|
// X-Actor-Source flags the auth path so resolveActor and
|
|
// any owner-only handler can deny without re-querying the
|
|
// token table. The value "task_token" is the only signal
|
|
// this header is allowed to carry — strip anything else a
|
|
// client tried to send.
|
|
r.Header.Set("X-Actor-Source", "task_token")
|
|
next.ServeHTTP(w, r)
|
|
return
|
|
}
|
|
|
|
// PAT: tokens starting with "mul_"
|
|
if strings.HasPrefix(tokenString, "mul_") {
|
|
hash := auth.HashToken(tokenString)
|
|
|
|
// Cache hit: TTL has not expired, the token was valid the
|
|
// last time we looked, and nothing has invalidated the
|
|
// entry since. Skip the DB SELECT and the last_used_at
|
|
// UPDATE — last_used_at is bumped once per TTL window.
|
|
if userID, ok := patCache.Get(r.Context(), hash); ok {
|
|
r.Header.Set("X-User-ID", userID)
|
|
next.ServeHTTP(w, r)
|
|
return
|
|
}
|
|
|
|
if queries == nil {
|
|
http.Error(w, `{"error":"invalid token"}`, http.StatusUnauthorized)
|
|
return
|
|
}
|
|
pat, err := queries.GetPersonalAccessTokenByHash(r.Context(), hash)
|
|
if err != nil {
|
|
slog.Warn("auth: invalid PAT", "path", r.URL.Path, "error", err)
|
|
http.Error(w, `{"error":"invalid token"}`, http.StatusUnauthorized)
|
|
return
|
|
}
|
|
|
|
userID := uuidToString(pat.UserID)
|
|
r.Header.Set("X-User-ID", userID)
|
|
|
|
// Clamp cache TTL to the token's remaining lifetime so a
|
|
// PAT expiring in <AuthCacheTTL can't continue passing
|
|
// auth on a cache hit after expires_at.
|
|
var expiresAt time.Time
|
|
if pat.ExpiresAt.Valid {
|
|
expiresAt = pat.ExpiresAt.Time
|
|
}
|
|
patCache.Set(r.Context(), hash, userID, auth.TTLForExpiry(time.Now(), expiresAt))
|
|
|
|
// Cache miss = TTL expired (or first use after revoke /
|
|
// process restart). Refresh last_used_at; subsequent hits
|
|
// within the TTL window skip this write entirely.
|
|
go queries.UpdatePersonalAccessTokenLastUsed(context.Background(), pat.ID)
|
|
|
|
next.ServeHTTP(w, r)
|
|
return
|
|
}
|
|
|
|
// JWT
|
|
token, err := jwt.Parse(tokenString, func(token *jwt.Token) (any, error) {
|
|
if _, ok := token.Method.(*jwt.SigningMethodHMAC); !ok {
|
|
return nil, jwt.ErrSignatureInvalid
|
|
}
|
|
return auth.JWTSecret(), nil
|
|
})
|
|
if err != nil || !token.Valid {
|
|
slog.Warn("auth: invalid token", "path", r.URL.Path, "error", err)
|
|
http.Error(w, `{"error":"invalid token"}`, http.StatusUnauthorized)
|
|
return
|
|
}
|
|
|
|
claims, ok := token.Claims.(jwt.MapClaims)
|
|
if !ok {
|
|
slog.Warn("auth: invalid claims", "path", r.URL.Path)
|
|
http.Error(w, `{"error":"invalid claims"}`, http.StatusUnauthorized)
|
|
return
|
|
}
|
|
|
|
sub, ok := claims["sub"].(string)
|
|
if !ok || strings.TrimSpace(sub) == "" {
|
|
slog.Warn("auth: invalid claims", "path", r.URL.Path)
|
|
http.Error(w, `{"error":"invalid claims"}`, http.StatusUnauthorized)
|
|
return
|
|
}
|
|
r.Header.Set("X-User-ID", sub)
|
|
if email, ok := claims["email"].(string); ok {
|
|
r.Header.Set("X-User-Email", email)
|
|
}
|
|
|
|
next.ServeHTTP(w, r)
|
|
})
|
|
}
|
|
}
|
|
|
|
// extractToken returns the bearer token and whether it came from a cookie.
|
|
// Priority: Authorization header > multica_auth cookie.
|
|
func extractToken(r *http.Request) (token string, fromCookie bool) {
|
|
if authHeader := r.Header.Get("Authorization"); authHeader != "" {
|
|
tokenString := strings.TrimPrefix(authHeader, "Bearer ")
|
|
if tokenString != authHeader {
|
|
return tokenString, false
|
|
}
|
|
}
|
|
|
|
if cookie, err := r.Cookie(auth.AuthCookieName); err == nil && cookie.Value != "" {
|
|
return cookie.Value, true
|
|
}
|
|
|
|
return "", false
|
|
}
|