mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-13 03:15:34 +02:00
* feat(project): add local_directory project_resource type (MUL-2662)
Adds a second project_resource type alongside github_repo so a project
can be pinned to an existing directory on a specific daemon (the v1 of
the local-working-directory flow tracked in MUL-2618). The ref schema is
{ local_path, daemon_id, label? }; local_path must be absolute and
daemon_id is required. The same (daemon_id, local_path) pair is allowed
on multiple projects by design — no UNIQUE constraint is added.
Implementation reuses the existing project_resource API surface: the new
type is wired through the validator switch with no migration, no new
events, and no daemon-handler changes (daemon already passes through
arbitrary resource types via ProjectResources). The CLI gains
--local-path / --daemon-id / --ref-label shortcuts so
`multica project resource add --type local_directory` mirrors the
existing `--type github_repo --url ...` ergonomics; the generic --ref
flag still works for both types.
Tests cover the full CRUD lifecycle, the same-path-across-projects
allowance, the same-path-same-project conflict, the validator rejections
(missing/blank/relative path, missing daemon_id, wrong payload type),
and the cross-platform isAbsoluteLocalPath helper.
Co-authored-by: multica-agent <github@multica.ai>
* feat(project): add update endpoint + label-shadow guard for project_resource (MUL-2662)
Addresses the Elon review on PR #3263:
- Add PUT /api/projects/{id}/resources/{resourceId} with sqlc query,
matching handler, CLI `project resource update`, and a new
EventProjectResourceUpdated WS event. resource_type stays immutable;
ref/label/position are all individually optional.
- Catch same-project (daemon_id, local_path) collisions where only the
embedded label differs — the row-level UNIQUE only matches the full
ref JSON, so a label typo would otherwise let the same working
directory bind twice.
- Tests cover the update lifecycle (label-only / ref / clear / 404 /
invalid path) and the label-shadow conflict on both create and
update; the in-place rename still succeeds because the conflict
scan ignores the row being edited.
Incidental: regenerating sqlc picked up a missing skills_local scan in
UpdateAgentCustomEnv that drifted in from #3200.
Co-authored-by: multica-agent <github@multica.ai>
* fix(project): close bundled-create label-shadow gap + merge resource_ref on CLI update (MUL-2662)
Two follow-ups from MUL-2662 review round 2:
- CreateProject inline resources path now dedupes local_directory entries on
(daemon_id, local_path) before opening the transaction. The DB-level
UNIQUE(project_id, resource_type, resource_ref) constraint only fires on a
full JSON match, so two rows with the same target but different `label`
would otherwise slip past. Standalone POST/PUT already cover this via
findLocalDirectoryConflict; bundled create was the missing surface.
- `multica project resource update` now seeds resource_ref from the existing
row before applying per-type shortcut flags, so `--default-branch-hint x`
on its own no longer constructs a payload missing `url` (which the server
400s on). Local_directory partial edits get the same merge behavior.
Co-authored-by: multica-agent <github@multica.ai>
* feat(desktop): local_directory project_resource UI (MUL-2665) (#3273)
* feat(desktop): local_directory project_resource UI (MUL-2665)
First UI surface for the local-working-directory flow tracked in MUL-2618.
Lets users on the desktop pin a project to an existing folder on this
machine; web stays read-only since the per-daemon check can't be done in
the browser.
What's new for the renderer:
- ProjectResourcesSection grows a desktop-only "Add local directory"
button next to the existing GitHub-repo popover. Clicking it opens
Electron's native folder picker, validates the path through a new
IPC pair (existence + r/w), and submits a project_resource of
resource_type=local_directory with daemon_id pulled live from
daemonAPI.getStatus.
- LocalDirectoryRow renders the rename pencil + path tooltip, and
greys out when ref.daemon_id != this machine's daemon_id (with a
"only available on the machine that registered this directory"
tooltip). Delete stays enabled so users can drop stale registrations
from any device.
- LocalDirectoryHint sits above the issue-detail comment composer and
shows "Agent will work in-place at {label} ({path})" when the issue's
project has a local_directory matching this daemon. Hidden on web.
- TaskStatusPill picks up a new "waiting_for_directory_release" stage
that the daemon will publish when it dequeues a task but can't
acquire the path lock. The render is in place now so the daemon
sibling subtask can wire the status string without an additional UI
PR.
Plumbing:
- @multica/core/types gains LocalDirectoryResourceRef +
UpdateProjectResourceRequest, and the api client gets the matching
PUT method backed by the server endpoint that landed in
2ac3faebb (MUL-2662). A useUpdateProjectResource hook drives the
in-place label edit.
- New Electron handlers under apps/desktop/src/main/local-directory.ts:
local-directory:pick -> dialog.showOpenDialog (openDirectory)
local-directory:validate -> stat + access(R_OK + W_OK)
exposed through the preload as desktopAPI.pickDirectory /
validateLocalDirectory. View code talks to them via a thin
packages/views/platform helper that returns reason=unsupported on
web instead of crashing.
- useLocalDaemonStatus exposes the local daemon's id, device name, and
running flag from daemonAPI.onStatusChange so the renderer can do the
cross-device match without coupling to the desktop preload typings.
Tests:
- pickStageKeys gets a unit test covering the new stage and proving
the directory-release status outranks availability hints.
- LocalDirectoryHint tests cover the four render branches (no project,
no daemon, foreign daemon, matching daemon).
- i18n parity stays green; new keys added under projects.resources.*
and chat.status_pill.stages.waiting_for_directory_release in both
locales.
Out of scope (will land separately):
- The daemon-side waiting/lock signal that flips the pill into the
new state.
- Adding local_directory to the create-project modal's bulk
attach flow.
- Docs page refresh for project-resources.mdx — left for the
MUL-2618 umbrella sweep.
Co-authored-by: multica-agent <github@multica.ai>
* fix(desktop): hide rename for foreign daemon local_directory rows (MUL-2618)
Address review nit on #3273: the rename pencil was gated only by
`canEdit`, so a foreign / unknown-daemon row still showed it even
though the spec says cross-device rows are disabled. Gate rename on
`!mismatch` so it disappears on those rows; delete stays available
so a stale registration can still be dropped from any device.
Co-authored-by: multica-agent <github@multica.ai>
---------
Co-authored-by: multica-agent <github@multica.ai>
* feat(daemon): local_directory execution + path mutex + GC exception (MUL-2663) (#3274)
* feat(daemon): local_directory execution + path mutex + GC exception (MUL-2663)
Wires up the daemon side of the local_directory project_resource introduced
in MUL-2662. When a task is dispatched against a project whose resources
include a local_directory pinned to this daemon's UUID, the daemon now:
- Validates the path (absolute, exists, daemon process can read+write,
not in the system-root / $HOME blacklist) and fails the task fast on
any precondition violation, with a user-readable reason.
- Serialises concurrent tasks on the same on-disk path via a
daemon-local LocalPathLocker keyed by symlink-resolved realpath. The
lock is held for the entire task lifetime (claim → context write →
agent → result report).
- When the lock is contended, the daemon flips the row to a new
waiting_local_directory status on the server (carrying a wait_reason
like "<path> (held by task <short id>)") so the UI can render
"等待本地目录释放" instead of leaving the row silently in dispatched
past the sweeper timeout. The status accepts being woken into running
once the lock is acquired.
- Sets execenv.WorkDir to the user's path (no copy, no mount). envRoot
still lives under workspacesRoot/<wsID>/ and hosts output/, logs/, and
.gc_meta.json — the daemon's logbook for the run.
- Stamps GCMeta.LocalDirectory=true so the GC loop never RemoveAlls
envRoot for these tasks (gcActionClean → gcActionCleanArtifacts,
gcActionOrphan → gcActionSkip). The user's directory was never under
envRoot to begin with, so this is defense in depth.
- Skips execenv.Reuse for local_directory tasks because the prior
WorkDir is the user's path and reusing it through that code path
loses the envRoot association the GC loop needs. Prepare is cheap
here (no clone, no copy), so always running it is fine.
Server-side protocol changes:
- New CHECK value 'waiting_local_directory' on agent_task_queue.status
plus a wait_reason TEXT column (migration 109).
- All cancel / active / counted-as-running / orphan-recovery queries
expanded to include the new status; FailStaleTasks intentionally
excludes it (the daemon owns the wait).
- New SQL MarkAgentTaskWaitingLocalDirectory(id, reason) and a relaxed
StartAgentTask that accepts both dispatched and
waiting_local_directory as preconditions (and clears wait_reason on
the way through).
- New POST /api/daemon/tasks/{taskId}/wait-local-directory endpoint,
TaskService.MarkTaskWaitingLocalDirectory broadcaster, and matching
daemon Client.MarkTaskWaitingLocalDirectory.
Tests cover: path blacklist + R/W enforcement, mutex serialisation +
ctx-cancelled wait, lock handover between two tasks, GC never returns
gcActionClean / gcActionOrphan for local_directory rows (with negative
control for the standard path), and Prepare/Cleanup correctly substitute
+ protect the user's WorkDir.
The desktop UI side (UI for adding a local_directory resource, surfacing
the "等待本地目录" badge) is MUL-2665; the agent-task lifecycle changes
(no branch switch, dirty-tree tolerant, auto-commit) are MUL-2664.
This PR targets the shared MUL-2618 v1 feature branch agent/j/912b8cb1,
not main; the whole v1 will be merged to main together when complete.
Co-authored-by: multica-agent <github@multica.ai>
* fix(daemon): tighten local_directory status, symlink, cancel handling (MUL-2618)
Address the 3 must-fix items from Elon's review of PR #3274.
1. Status string unified. The server / daemon publish
`waiting_local_directory`; align views, locales, and the
pickStageKeys test (PR #3273 had used `waiting_for_directory_release`
on a placeholder string). Without this, the daemon's wait state
never reached the pill once the two siblings merged.
2. validateLocalPath now also runs the blacklist against the
symlink-resolved realpath, with macOS's `/etc` -> `/private/etc`
redirect handled via `isBlacklistedRealPath` which compares
canonical forms. Without this, a symlink such as
`/Users/me/proj/home -> /Users/me` slipped the literal $HOME check
while every daemon write still landed in the user's home. Tests
cover symlink-to-home, symlink-to-system-root, and the negative
case (symlink to a regular subdirectory).
3. acquireLocalDirectoryLockIfNeeded now spins up a cancellation
watcher inside `onWait` (lazy — the fast path stays free) so the
gap between dispatch and StartTask responds to server-side cancel
or row deletion. If the watcher fires while the daemon is parked
on the path mutex, the lock-wait context is cancelled, Acquire
returns promptly, and the helper exits silently the same way the
run-phase poller does. New TestAcquireLocalDirectoryLock_CancelDuringWait
exercises the path end-to-end with a fake server.
Co-authored-by: multica-agent <github@multica.ai>
* fix(daemon): unconditional canonical blacklist + Windows drive-root generalisation (MUL-2618)
- validateLocalPath now always runs isBlacklistedRealPath on the
symlink-resolved path, not only when it differs from absPath. The old
guard let users type the canonical form of an OS-symlinked banned root
(e.g. /private/tmp, /private/etc, /private/var on macOS) straight
through, since EvalSymlinks is a no-op on already-canonical input.
- Windows drive-root rejection moved off the static C/D/E/F enumeration
onto filepath.VolumeName via a new isDriveRoot helper, so removable /
network drives mounted at G:..Z: and UNC \\server\share roots are also
blocked. systemRootBlacklist keeps the well-known C:\ trees only.
- Tests: macOS-only case exercises direct /private/{tmp,etc,var}; a
new TestIsDriveRoot covers the Windows generalisation (skipped on
POSIX runners by runtime guard).
Co-authored-by: multica-agent <github@multica.ai>
---------
Co-authored-by: multica-agent <github@multica.ai>
* feat(views): wire waiting_local_directory end-to-end in issue UI + presence (MUL-2618)
Connect the daemon-emitted `task:waiting_local_directory` and `task:running`
events through to issue execution log, sticky agent banner, activity indicator,
and agent presence so a parked task is no longer invisible on the issue page.
- Add `waiting_local_directory` to `AgentTask.status` and the typed
`task:running` / `task:waiting_local_directory` WS event payloads.
- Chat realtime sync writes both new statuses into the pending-task cache so
the chat StatusPill flips out of a stale `dispatched` frame.
- ExecutionLogSection: count `waiting_local_directory` as active, add tone +
status label, treat parked tasks the same as dispatched for time anchor /
transcript visibility / terminate-confirm note.
- AgentLiveCard: subscribe to both new events, rank the parked state between
dispatched and queued, and surface a "is waiting for the local directory"
banner with the muted "Clock" treatment used for queued.
- IssueAgentActivityIndicator: route parked tasks into the queued bucket so
the hover stack and chip stay visible.
- derive-presence: parked tasks count toward `queuedCount` so the agent
workload chip stays out of `idle` while the daemon waits on the path lock.
- Locales: add `agent_live.is_waiting_local_directory` and
`execution_log.status_waiting_local_directory` (en + zh-Hans).
Co-authored-by: multica-agent <github@multica.ai>
* feat(project): enforce one local_directory per (project, daemon) (MUL-2618)
The daemon-side resolver picks the first matching local_directory by
daemon_id, so allowing two rows on the same daemon — even at different
paths — let the agent silently write into whichever sorted first. Tighten
the invariant top to bottom:
- server: `findLocalDirectoryConflict` rejects any second row sharing a
daemon_id, regardless of `local_path` or label. Bundled-create surface in
`CreateProject` runs the same daemon-scoped dedupe up front.
- daemon: `findLocalDirectoryAssignment` fails fast when it finds more than
one row pinned to the current daemon (older API client / direct DB
writes can still produce that state — refuse to guess).
- desktop UI: hide the "Add local directory" action once the current
daemon owns a row on this project, with a hint and a defensive toast on
the call path; foreign-daemon rows stay visible read-only as before.
- Tests:
* daemon: new `two local_directory rows on this daemon fail fast` /
`local_directory rows on different daemons coexist` cases.
* handler: rewrite the legacy `LabelShadow` cases as
`DaemonScopedConflict` / `BundledLocalDirectoryDaemonConflict` —
asserts 409 on same-daemon different-path, 201 on per-daemon bundles.
- Locales: en + zh-Hans copy for the new hint + toast.
Co-authored-by: multica-agent <github@multica.ai>
* chore(sqlc): drop stale skills_local in UpdateAgentCustomEnv (MUL-2618)
Follow-up to the main-merge in 0f8e8ca7: the auto-merge preserved most
of main's skills_local revert but kept the column reference inside the
UpdateAgentCustomEnv scanner because that block hadn't been touched by
either side. Re-running `sqlc generate` regenerates the file without
skills_local in this query, matching the rest of the file and the
post-revert schema.
Co-authored-by: multica-agent <github@multica.ai>
* feat(create-project): binary source picker — repos OR local directory
Turn the create-project dialog's "Repos" pill into a binary Source
picker. A project's source is mutually exclusive: either a set of
GitHub repos (worktree mode, default) or a single local working
directory (local mode, desktop-only). Mirrors the constraint the
backend will enforce next.
Behavior:
- Pill shows the active mode's selection (GitHub icon + repo count, or
folder icon + local label/path).
- Popover has a 2-tab segmented control at the top; the Local tab is
hidden entirely on web (local_directory needs a daemon_id).
- Local tab requires the daemon online — amber notice + disabled picker
when offline, re-renders automatically via useLocalDaemonStatus.
- Switching tabs preserves the other side's stash, but handleSubmit
only emits the resource matching the active sourceMode, so abandoned
picks never leak into the created project.
Backend mutual-exclusion validation + the resources-section
conditional-add-button still to come — this PR just unblocks the
dialog so it can be demoed.
* fix(mobile): cover waiting_local_directory in run row status maps (MUL-2618)
---------
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Multica J <j@multica.ai>
607 lines
19 KiB
Go
607 lines
19 KiB
Go
package daemon
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
"net/http"
|
|
"os"
|
|
"os/exec"
|
|
"path/filepath"
|
|
"strings"
|
|
"time"
|
|
|
|
"github.com/multica-ai/multica/server/internal/daemon/execenv"
|
|
)
|
|
|
|
// gcLoop periodically scans local workspace directories and removes those
|
|
// whose issue is done/cancelled and hasn't been updated within the configured TTL.
|
|
func (d *Daemon) gcLoop(ctx context.Context) {
|
|
if !d.cfg.GCEnabled {
|
|
d.logger.Info("gc: disabled")
|
|
return
|
|
}
|
|
d.logger.Info("gc: started",
|
|
"interval", d.cfg.GCInterval,
|
|
"ttl", d.cfg.GCTTL,
|
|
"orphan_ttl", d.cfg.GCOrphanTTL,
|
|
"artifact_ttl", d.cfg.GCArtifactTTL,
|
|
"artifact_patterns", d.cfg.GCArtifactPatterns,
|
|
)
|
|
|
|
// Run once at startup after a short delay (let the daemon finish initializing).
|
|
if err := sleepWithContext(ctx, 30*time.Second); err != nil {
|
|
return
|
|
}
|
|
d.runGC(ctx)
|
|
|
|
ticker := time.NewTicker(d.cfg.GCInterval)
|
|
defer ticker.Stop()
|
|
|
|
for {
|
|
select {
|
|
case <-ctx.Done():
|
|
return
|
|
case <-ticker.C:
|
|
d.runGC(ctx)
|
|
}
|
|
}
|
|
}
|
|
|
|
// gcStats accumulates byte counts and per-pattern hit counts for one GC cycle.
|
|
type gcStats struct {
|
|
cleaned int // whole task dirs removed (issue done/cancelled)
|
|
orphaned int // whole task dirs removed (no meta / unreachable issue)
|
|
skipped int // task dirs left untouched
|
|
artifactDirs int // task dirs that had at least one artifact reclaimed
|
|
artifactRemoved int // count of removed artifact subdirs
|
|
bytesReclaimed int64 // total bytes freed in this cycle
|
|
byPattern map[string]int // basename -> reclaim count, for visibility
|
|
}
|
|
|
|
// runGC performs a single GC scan across all workspace directories.
|
|
func (d *Daemon) runGC(ctx context.Context) {
|
|
root := d.cfg.WorkspacesRoot
|
|
entries, err := os.ReadDir(root)
|
|
if err != nil {
|
|
if os.IsNotExist(err) {
|
|
return
|
|
}
|
|
d.logger.Warn("gc: read workspaces root failed", "error", err)
|
|
return
|
|
}
|
|
|
|
stats := &gcStats{byPattern: map[string]int{}}
|
|
for _, wsEntry := range entries {
|
|
if !wsEntry.IsDir() || wsEntry.Name() == ".repos" {
|
|
continue
|
|
}
|
|
wsDir := filepath.Join(root, wsEntry.Name())
|
|
d.gcWorkspace(ctx, wsDir, stats)
|
|
}
|
|
|
|
// Prune stale worktree references from all bare repo caches.
|
|
d.pruneRepoWorktrees(root)
|
|
|
|
if stats.cleaned > 0 || stats.orphaned > 0 || stats.artifactDirs > 0 {
|
|
d.logger.Info("gc: cycle complete",
|
|
"cleaned", stats.cleaned,
|
|
"orphaned", stats.orphaned,
|
|
"skipped", stats.skipped,
|
|
"artifact_dirs", stats.artifactDirs,
|
|
"artifact_removed", stats.artifactRemoved,
|
|
"bytes_reclaimed", stats.bytesReclaimed,
|
|
"by_pattern", stats.byPattern,
|
|
)
|
|
}
|
|
}
|
|
|
|
// gcWorkspace scans task directories inside a single workspace directory.
|
|
func (d *Daemon) gcWorkspace(ctx context.Context, wsDir string, stats *gcStats) {
|
|
taskEntries, err := os.ReadDir(wsDir)
|
|
if err != nil {
|
|
d.logger.Warn("gc: read workspace dir failed", "dir", wsDir, "error", err)
|
|
return
|
|
}
|
|
|
|
cleanedHere := 0
|
|
for _, entry := range taskEntries {
|
|
if ctx.Err() != nil {
|
|
return
|
|
}
|
|
if !entry.IsDir() {
|
|
continue
|
|
}
|
|
taskDir := filepath.Join(wsDir, entry.Name())
|
|
action := d.shouldCleanTaskDir(ctx, taskDir)
|
|
switch action {
|
|
case gcActionClean:
|
|
bytes := dirSize(taskDir)
|
|
d.cleanTaskDir(taskDir)
|
|
stats.cleaned++
|
|
stats.bytesReclaimed += bytes
|
|
cleanedHere++
|
|
case gcActionOrphan:
|
|
bytes := dirSize(taskDir)
|
|
d.cleanTaskDir(taskDir)
|
|
stats.orphaned++
|
|
stats.bytesReclaimed += bytes
|
|
cleanedHere++
|
|
case gcActionCleanArtifacts:
|
|
removed, bytes, perPattern := d.cleanTaskArtifacts(taskDir, d.cfg.GCArtifactPatterns)
|
|
if removed > 0 {
|
|
stats.artifactDirs++
|
|
stats.artifactRemoved += removed
|
|
stats.bytesReclaimed += bytes
|
|
for k, v := range perPattern {
|
|
stats.byPattern[k] += v
|
|
}
|
|
}
|
|
stats.skipped++ // task dir itself preserved
|
|
default:
|
|
stats.skipped++
|
|
}
|
|
}
|
|
|
|
// Remove the workspace directory itself if it's now empty.
|
|
if cleanedHere > 0 {
|
|
remaining, _ := os.ReadDir(wsDir)
|
|
if len(remaining) == 0 {
|
|
os.Remove(wsDir)
|
|
}
|
|
}
|
|
}
|
|
|
|
type gcAction int
|
|
|
|
const (
|
|
gcActionSkip gcAction = iota
|
|
gcActionClean // issue is done/cancelled and stale
|
|
gcActionOrphan // no meta or unknown issue and dir is old
|
|
gcActionCleanArtifacts // task completed long enough ago; drop regenerable artifacts only
|
|
)
|
|
|
|
// shouldCleanTaskDir decides whether a task directory should be removed.
|
|
// Dispatches on meta.Kind so chat / autopilot / quick-create tasks each
|
|
// follow the parent record that actually governs their lifecycle.
|
|
func (d *Daemon) shouldCleanTaskDir(ctx context.Context, taskDir string) gcAction {
|
|
// A task currently running on this env root must never be reclaimed —
|
|
// not even on the done/cancelled or orphan-404 paths. A new comment on
|
|
// an already-done issue can dispatch a follow-up task that reuses the
|
|
// prior workdir without bumping the issue's updated_at, so the regular
|
|
// TTL check alone wouldn't notice the resumed activity.
|
|
if d.isActiveEnvRoot(taskDir) {
|
|
return gcActionSkip
|
|
}
|
|
|
|
meta, err := execenv.ReadGCMeta(taskDir)
|
|
if err != nil {
|
|
return d.orphanByMTime(taskDir, "no meta")
|
|
}
|
|
|
|
action := d.shouldCleanTaskDirForKind(ctx, taskDir, meta)
|
|
if !meta.LocalDirectory {
|
|
return action
|
|
}
|
|
// local_directory tasks keep their envRoot indefinitely so the user
|
|
// can inspect output/ and logs/ for forensic context. The WorkDir is
|
|
// the user's own path and lives outside taskDir, so the envRoot
|
|
// itself is just the daemon's logbook for the run — never large, and
|
|
// safe to keep.
|
|
//
|
|
// gcActionClean → demote to artifact-pattern cleanup so envRoot
|
|
// (and especially the logbook) survives.
|
|
// gcActionOrphan → skip outright; we don't ever wipe a
|
|
// local_directory envRoot via the mtime path,
|
|
// since the parent issue / chat record going
|
|
// away should not collateral-delete the user's
|
|
// own audit trail.
|
|
//
|
|
// gcActionCleanArtifacts and gcActionSkip already obey the
|
|
// "no full envRoot RemoveAll" rule.
|
|
switch action {
|
|
case gcActionClean:
|
|
return gcActionCleanArtifacts
|
|
case gcActionOrphan:
|
|
return gcActionSkip
|
|
default:
|
|
return action
|
|
}
|
|
}
|
|
|
|
// shouldCleanTaskDirForKind runs the per-Kind dispatch without applying the
|
|
// local_directory override. Split out so shouldCleanTaskDir can intercept
|
|
// the result.
|
|
func (d *Daemon) shouldCleanTaskDirForKind(ctx context.Context, taskDir string, meta *execenv.GCMeta) gcAction {
|
|
switch meta.Kind {
|
|
case execenv.GCKindIssue:
|
|
return d.gcDecisionIssue(ctx, taskDir, meta)
|
|
case execenv.GCKindChat:
|
|
return d.gcDecisionChat(ctx, taskDir, meta)
|
|
case execenv.GCKindAutopilotRun:
|
|
return d.gcDecisionAutopilotRun(ctx, taskDir, meta)
|
|
case execenv.GCKindQuickCreate:
|
|
return d.gcDecisionQuickCreate(ctx, taskDir, meta)
|
|
default:
|
|
// Unknown kind: fall back to mtime-based orphan cleanup so a future
|
|
// daemon writing a kind we don't recognize doesn't get insta-wiped.
|
|
return d.orphanByMTime(taskDir, "unknown kind")
|
|
}
|
|
}
|
|
|
|
// orphanByMTime returns gcActionOrphan if the directory is older than
|
|
// GCOrphanTTL, gcActionSkip otherwise. Centralizes the "we have no parent
|
|
// record signal so just look at the disk" fallback used by every kind.
|
|
func (d *Daemon) orphanByMTime(taskDir, reason string) gcAction {
|
|
info, err := os.Stat(taskDir)
|
|
if err != nil {
|
|
return gcActionSkip
|
|
}
|
|
if time.Since(info.ModTime()) > d.cfg.GCOrphanTTL {
|
|
d.logger.Info("gc: orphan directory", "dir", taskDir, "reason", reason, "age", time.Since(info.ModTime()).Round(time.Hour))
|
|
return gcActionOrphan
|
|
}
|
|
return gcActionSkip
|
|
}
|
|
|
|
// isAccessNotFound detects the 404 returned by gc-check endpoints. The same
|
|
// status covers "row deleted" and "daemon token can't see this workspace"
|
|
// (the requireDaemonWorkspaceAccess anti-enumeration shape), so callers
|
|
// can't tell the two apart from the response alone.
|
|
func isAccessNotFound(err error) bool {
|
|
var reqErr *requestError
|
|
return errors.As(err, &reqErr) && reqErr.StatusCode == http.StatusNotFound
|
|
}
|
|
|
|
func (d *Daemon) gcDecisionIssue(ctx context.Context, taskDir string, meta *execenv.GCMeta) gcAction {
|
|
status, err := d.client.GetIssueGCCheck(ctx, meta.IssueID)
|
|
if err != nil {
|
|
if isAccessNotFound(err) {
|
|
// 404 is ambiguous: server returns it for both "issue deleted"
|
|
// and "daemon token has no access to the workspace". Fall back
|
|
// to the mtime-gated orphan cleanup so a scoped-down token
|
|
// can't instantly wipe dirs whose issues are still live.
|
|
return d.orphanByMTime(taskDir, "issue not accessible")
|
|
}
|
|
return gcActionSkip
|
|
}
|
|
|
|
if (status.Status == "done" || status.Status == "cancelled") &&
|
|
time.Since(status.UpdatedAt) > d.cfg.GCTTL {
|
|
d.logger.Info("gc: eligible for cleanup",
|
|
"dir", filepath.Base(taskDir),
|
|
"kind", "issue",
|
|
"issue", meta.IssueID,
|
|
"status", status.Status,
|
|
"updated_at", status.UpdatedAt.Format(time.RFC3339),
|
|
)
|
|
return gcActionClean
|
|
}
|
|
|
|
if d.cfg.GCArtifactTTL > 0 && len(d.cfg.GCArtifactPatterns) > 0 &&
|
|
!meta.CompletedAt.IsZero() && time.Since(meta.CompletedAt) > d.cfg.GCArtifactTTL {
|
|
d.logger.Info("gc: eligible for artifact cleanup",
|
|
"dir", filepath.Base(taskDir),
|
|
"kind", "issue",
|
|
"issue", meta.IssueID,
|
|
"status", status.Status,
|
|
"completed_at", meta.CompletedAt.Format(time.RFC3339),
|
|
)
|
|
return gcActionCleanArtifacts
|
|
}
|
|
|
|
return gcActionSkip
|
|
}
|
|
|
|
func (d *Daemon) gcDecisionChat(ctx context.Context, taskDir string, meta *execenv.GCMeta) gcAction {
|
|
status, err := d.client.GetChatSessionGCCheck(ctx, meta.ChatSessionID)
|
|
if err != nil {
|
|
if isAccessNotFound(err) {
|
|
// 404 means the chat_session row is gone — DeleteChatSession is
|
|
// a real DELETE, so a hard delete propagates here as soon as
|
|
// the user clicks the button. This is the strongest reclaim
|
|
// signal we get and it's exactly acceptance criterion #3:
|
|
// reclaim within one GC cycle (≤ GCInterval), not 72h.
|
|
//
|
|
// We don't gate on mtime: every chat_session_id in a meta file
|
|
// was written by this daemon under its current token, so there
|
|
// is no cross-workspace probe to defend against.
|
|
d.logger.Info("gc: eligible for cleanup",
|
|
"dir", filepath.Base(taskDir),
|
|
"kind", "chat",
|
|
"chat_session", meta.ChatSessionID,
|
|
"reason", "session not accessible (hard-deleted)",
|
|
)
|
|
return gcActionClean
|
|
}
|
|
return gcActionSkip
|
|
}
|
|
|
|
switch status.Status {
|
|
case "active":
|
|
// An active chat session must never be reclaimed by mtime — that
|
|
// would silently kill a user's idle session and break "PriorWorkDir"
|
|
// resume on their next message. This is the explicit short-circuit
|
|
// the issue body called out as verifyable behavior #2.
|
|
return gcActionSkip
|
|
case "archived":
|
|
if time.Since(status.UpdatedAt) > d.cfg.GCTTL {
|
|
d.logger.Info("gc: eligible for cleanup",
|
|
"dir", filepath.Base(taskDir),
|
|
"kind", "chat",
|
|
"chat_session", meta.ChatSessionID,
|
|
"status", status.Status,
|
|
"updated_at", status.UpdatedAt.Format(time.RFC3339),
|
|
)
|
|
return gcActionClean
|
|
}
|
|
}
|
|
return gcActionSkip
|
|
}
|
|
|
|
func (d *Daemon) gcDecisionAutopilotRun(ctx context.Context, taskDir string, meta *execenv.GCMeta) gcAction {
|
|
status, err := d.client.GetAutopilotRunGCCheck(ctx, meta.AutopilotRunID)
|
|
if err != nil {
|
|
if isAccessNotFound(err) {
|
|
return d.orphanByMTime(taskDir, "autopilot run not accessible")
|
|
}
|
|
return gcActionSkip
|
|
}
|
|
|
|
// Terminal states per the autopilot_run CHECK constraint:
|
|
// completed, failed, skipped — the run finished its own work.
|
|
// issue_created — the run produced an issue task that owns
|
|
// its own workdir; this run's workdir is
|
|
// dead weight from here on.
|
|
// Non-terminal: pending, running. Skip until they reach a terminal state
|
|
// rather than trying to bound them by mtime — long autopilots are real.
|
|
if isAutopilotRunTerminal(status.Status) {
|
|
anchor := status.CompletedAt
|
|
if anchor.IsZero() {
|
|
// Defensive: terminal status without completed_at means the
|
|
// run finished but the column wasn't stamped (older code path).
|
|
// Fall back to the meta's CompletedAt so we still GC eventually.
|
|
anchor = meta.CompletedAt
|
|
}
|
|
if !anchor.IsZero() && time.Since(anchor) > d.cfg.GCTTL {
|
|
d.logger.Info("gc: eligible for cleanup",
|
|
"dir", filepath.Base(taskDir),
|
|
"kind", "autopilot_run",
|
|
"autopilot_run", meta.AutopilotRunID,
|
|
"status", status.Status,
|
|
"completed_at", anchor.Format(time.RFC3339),
|
|
)
|
|
return gcActionClean
|
|
}
|
|
}
|
|
return gcActionSkip
|
|
}
|
|
|
|
// isAutopilotRunTerminal mirrors the run.status CHECK in
|
|
// migrations/042_autopilot.up.sql. Non-terminal states are pending/running;
|
|
// every other value the schema allows is a final resting state from the
|
|
// daemon's POV (the run is no longer producing work in this workdir).
|
|
func isAutopilotRunTerminal(status string) bool {
|
|
switch status {
|
|
case "completed", "failed", "skipped", "issue_created":
|
|
return true
|
|
default:
|
|
return false
|
|
}
|
|
}
|
|
|
|
func (d *Daemon) gcDecisionQuickCreate(ctx context.Context, taskDir string, meta *execenv.GCMeta) gcAction {
|
|
status, err := d.client.GetTaskGCCheck(ctx, meta.TaskID)
|
|
if err != nil {
|
|
if isAccessNotFound(err) {
|
|
// Task row was hard-deleted, or token can't see it. Either way,
|
|
// fall back to mtime-gated orphan to stay safe across scoped
|
|
// tokens — same reasoning as the issue path.
|
|
return d.orphanByMTime(taskDir, "task not accessible")
|
|
}
|
|
return gcActionSkip
|
|
}
|
|
|
|
// Quick-create workdirs are not reused by the issue task that
|
|
// LinkTaskToIssue eventually attaches — that issue gets its own
|
|
// envRoot. So as soon as the quick-create task itself reaches a
|
|
// terminal state we can reclaim the directory immediately, without
|
|
// waiting for GCTTL. If the user wants to revisit, the linked issue
|
|
// has the agent's output already.
|
|
if isAgentTaskTerminal(status.Status) {
|
|
d.logger.Info("gc: eligible for cleanup",
|
|
"dir", filepath.Base(taskDir),
|
|
"kind", "quick_create",
|
|
"task", meta.TaskID,
|
|
"status", status.Status,
|
|
)
|
|
return gcActionClean
|
|
}
|
|
return gcActionSkip
|
|
}
|
|
|
|
// isAgentTaskTerminal reports whether a value of agent_task_queue.status
|
|
// represents a final state. Mirrors the status enum used across the
|
|
// task service — see service/task.go for the canonical list.
|
|
func isAgentTaskTerminal(status string) bool {
|
|
switch status {
|
|
case "completed", "failed", "cancelled":
|
|
return true
|
|
default:
|
|
return false
|
|
}
|
|
}
|
|
|
|
// cleanTaskDir removes a task directory and logs the result.
|
|
func (d *Daemon) cleanTaskDir(taskDir string) {
|
|
if err := os.RemoveAll(taskDir); err != nil {
|
|
d.logger.Warn("gc: remove task dir failed", "dir", taskDir, "error", err)
|
|
} else {
|
|
d.logger.Info("gc: removed", "dir", taskDir)
|
|
}
|
|
}
|
|
|
|
// cleanTaskArtifacts walks taskDir and deletes every directory whose basename
|
|
// matches one of patterns. Returns (removedCount, bytesReclaimed, perPattern).
|
|
//
|
|
// Safety contract:
|
|
// - patterns are basename-only; entries with a path separator are dropped.
|
|
// - .git subtrees are never descended into, so the agent's git history stays
|
|
// intact even if a pattern would otherwise match.
|
|
// - symlinks are skipped entirely — neither the link nor its target is
|
|
// touched, so a malicious or stale link can't redirect the GC outside the
|
|
// workdir.
|
|
// - every removal target is verified to live inside taskDir, so a tampered
|
|
// .gc_meta.json can't trick the daemon into deleting outside its sandbox.
|
|
func (d *Daemon) cleanTaskArtifacts(taskDir string, patterns []string) (removed int, bytes int64, perPattern map[string]int) {
|
|
perPattern = map[string]int{}
|
|
if taskDir == "" || len(patterns) == 0 {
|
|
return
|
|
}
|
|
patternSet := make(map[string]struct{}, len(patterns))
|
|
for _, p := range patterns {
|
|
p = strings.TrimSpace(p)
|
|
if p == "" || strings.ContainsAny(p, "/\\") {
|
|
continue
|
|
}
|
|
patternSet[p] = struct{}{}
|
|
}
|
|
if len(patternSet) == 0 {
|
|
return
|
|
}
|
|
|
|
absRoot, err := filepath.Abs(taskDir)
|
|
if err != nil {
|
|
return
|
|
}
|
|
|
|
walkErr := filepath.WalkDir(absRoot, func(path string, entry os.DirEntry, err error) error {
|
|
if err != nil {
|
|
return nil // best-effort — keep walking
|
|
}
|
|
if path == absRoot {
|
|
return nil
|
|
}
|
|
if !entry.IsDir() {
|
|
return nil
|
|
}
|
|
// Never descend into .git — preserves agent commits even if a pattern
|
|
// like "objects" would otherwise match.
|
|
if entry.Name() == ".git" {
|
|
return filepath.SkipDir
|
|
}
|
|
// Refuse to follow symlinked directories. WalkDir reports them as type
|
|
// Dir on some platforms; lstat to be sure.
|
|
info, statErr := os.Lstat(path)
|
|
if statErr != nil {
|
|
return nil
|
|
}
|
|
if info.Mode()&os.ModeSymlink != 0 {
|
|
return filepath.SkipDir
|
|
}
|
|
if _, ok := patternSet[entry.Name()]; !ok {
|
|
return nil
|
|
}
|
|
// Containment check: target must remain inside taskDir.
|
|
rel, relErr := filepath.Rel(absRoot, path)
|
|
if relErr != nil || rel == "" || rel == "." || strings.HasPrefix(rel, "..") {
|
|
return filepath.SkipDir
|
|
}
|
|
size := dirSize(path)
|
|
if rmErr := os.RemoveAll(path); rmErr != nil {
|
|
d.logger.Warn("gc: artifact remove failed", "path", path, "error", rmErr)
|
|
return filepath.SkipDir
|
|
}
|
|
removed++
|
|
bytes += size
|
|
perPattern[entry.Name()]++
|
|
d.logger.Info("gc: artifact removed", "path", path, "bytes", size)
|
|
// Don't descend into the now-deleted subtree.
|
|
return filepath.SkipDir
|
|
})
|
|
if walkErr != nil {
|
|
d.logger.Warn("gc: artifact walk failed", "dir", taskDir, "error", walkErr)
|
|
}
|
|
return
|
|
}
|
|
|
|
// dirSize returns the total size of all regular files under root, in bytes.
|
|
// Non-fatal: errors during the walk are ignored so callers can report a
|
|
// best-effort byte count without aborting the whole GC cycle.
|
|
func dirSize(root string) int64 {
|
|
var total int64
|
|
_ = filepath.WalkDir(root, func(_ string, entry os.DirEntry, err error) error {
|
|
if err != nil {
|
|
return nil
|
|
}
|
|
if entry.IsDir() {
|
|
return nil
|
|
}
|
|
info, infoErr := entry.Info()
|
|
if infoErr != nil {
|
|
return nil
|
|
}
|
|
if info.Mode().IsRegular() {
|
|
total += info.Size()
|
|
}
|
|
return nil
|
|
})
|
|
return total
|
|
}
|
|
|
|
const gitCmdTimeout = 30 * time.Second
|
|
|
|
// pruneRepoWorktrees runs `git worktree prune` on all bare repos in the cache.
|
|
func (d *Daemon) pruneRepoWorktrees(workspacesRoot string) {
|
|
reposRoot := filepath.Join(workspacesRoot, ".repos")
|
|
wsEntries, err := os.ReadDir(reposRoot)
|
|
if err != nil {
|
|
return
|
|
}
|
|
|
|
for _, wsEntry := range wsEntries {
|
|
if !wsEntry.IsDir() {
|
|
continue
|
|
}
|
|
wsRepoDir := filepath.Join(reposRoot, wsEntry.Name())
|
|
repoEntries, err := os.ReadDir(wsRepoDir)
|
|
if err != nil {
|
|
continue
|
|
}
|
|
for _, repoEntry := range repoEntries {
|
|
if !repoEntry.IsDir() {
|
|
continue
|
|
}
|
|
barePath := filepath.Join(wsRepoDir, repoEntry.Name())
|
|
if !isBareRepo(barePath) {
|
|
continue
|
|
}
|
|
d.pruneWorktree(barePath)
|
|
}
|
|
}
|
|
}
|
|
|
|
func (d *Daemon) pruneWorktree(barePath string) {
|
|
ctx, cancel := context.WithTimeout(context.Background(), gitCmdTimeout)
|
|
defer cancel()
|
|
cmd := exec.CommandContext(ctx, "git", "-C", barePath, "worktree", "prune")
|
|
|
|
if out, err := cmd.CombinedOutput(); err != nil {
|
|
d.logger.Warn("gc: worktree prune failed",
|
|
"repo", barePath,
|
|
"output", strings.TrimSpace(string(out)),
|
|
"error", err,
|
|
)
|
|
}
|
|
}
|
|
|
|
// isBareRepo checks if a path looks like a bare git repository.
|
|
func isBareRepo(path string) bool {
|
|
if _, err := os.Stat(filepath.Join(path, "HEAD")); err != nil {
|
|
return false
|
|
}
|
|
if _, err := os.Stat(filepath.Join(path, "objects")); err != nil {
|
|
return false
|
|
}
|
|
return true
|
|
}
|