Files
multica/packages/core/agents/use-agent-activity.ts
Naiyuan Qing 21e3cfaa01 Agent runtime status redesign: split presence into availability + last-task (#1794)
* feat(agent-status): add workspace live-tasks endpoint and TaskFailureReason type

Lays the API + type contract for the front-end agent presence cache:

- New `GET /api/active-tasks` returns active (queued/dispatched/running)
  tasks plus failed tasks within the last 2 minutes for the current
  workspace. The 2-minute window powers a UI-side auto-clearing "Failed"
  agent state without back-end pollers.
- `agent_task_queue` has no workspace_id column, so the query JOINs agent;
  `SELECT atq.*` keeps `failure_reason` (migration 055) on the wire.
- Adds `TaskFailureReason` to `AgentTask` so the UI can map the 5 backend
  classifiers (agent_error / timeout / runtime_offline / runtime_recovery
  / manual) to copy without parsing free-text errors.
- New `api.getActiveTasksForWorkspace()` client method; workspace is
  resolved server-side from the X-Workspace-Slug header (no path param,
  matching /api/agents and /api/runtimes conventions).

Includes the joint engineering plan and designer brief that scope the
broader Agent / Runtime status redesign — Phase 0 is this contract plus
the front-end derivation layer landing in the next commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(agent-status): derive presence/health states with WS sync and desktop IPC bridge

Adds the front-end derivation layer that turns raw server data into the
user-facing 5-state agent / 4-state runtime enums. UI files are
deliberately untouched in this commit — derivation lives behind hooks
(useAgentPresence, useRuntimeHealth) that any component can call with
zero additional network traffic.

Architecture:
- Derivation is pure functions in packages/core/{agents,runtimes}; the
  back-end stays free of UI translation. Agents algorithm: runtime
  offline > recent failed (2-min window) > running > queued > available.
  Runtimes algorithm: status + last_seen_at -> online / recently_lost /
  offline / about_to_gc.
- A single workspace-wide active-tasks query backs all per-agent
  presence reads, eliminating N+1 across hover cards, list rows, and
  pickers. 30-second tick re-renders the hooks so the failed window
  expires even when no underlying data changes.
- WS task lifecycle events (dispatch / completed / failed / cancelled)
  invalidate active-tasks via the prefix dispatcher. completed/failed
  were removed from specificEvents so they go through both the prefix
  invalidate and the existing chat ws.on() handlers. Reconnect refetch
  picks up active-tasks too.
- Desktop bridges window.daemonAPI.onStatusChange directly into the
  runtimes cache via setQueryData, giving the local daemon sub-second
  feedback (vs. 75s server sweep). Bridge is wsId-bound so workspace
  switches automatically rebind the subscription; daemon_id matching
  covers the same-daemon-multiple-providers case.

24 derivation unit tests cover all branches plus null/empty/boundary
inputs (FAILED_WINDOW_MS edges, null last_seen_at, missing
completed_at). Full core suite: 112 tests passing. Typecheck green
across all 8 workspace packages.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(agent-status): redesign agent runtime status as two orthogonal dimensions

Splits the conflated 5-state agent presence into two independent axes:

- AgentAvailability (3-state): online / unstable / offline — drives the
  dot indicator everywhere a dot appears. Pure runtime reachability;
  never sticky-red because of a past task outcome.

- LastTaskState (5-state): running / completed / failed / cancelled /
  idle — surfaced as text + icon on focused surfaces (hover card,
  agent detail page, agents list, runtime detail). Never colours the dot.

Major changes:

* Domain layer: AgentPresence union → AgentAvailability + LastTaskState.
  derive-presence split into deriveAgentAvailability + deriveLastTaskState
  + deriveAgentPresenceDetail orchestrator. Tests reorganised into three
  groups (availability invariants, last-task invariants, composition).

* Visual config: presenceConfig (5 entries) → availabilityConfig (3) +
  taskStateConfig (5). availabilityOrder + lastTaskOrder for filter chips.

* Workspace-level presence prefetch: new useWorkspacePresencePrefetch
  hook + WorkspacePresencePrefetch mount component, wired into
  DashboardLayout (web) and WorkspaceRouteLayout (desktop). Hover cards
  render synchronously with no skeleton flash on first hover.

* ActorAvatar hover: flipped default — disableHoverCard removed,
  enableHoverCard added (default false). Opt-in at ~14 decision-moment
  surfaces; pickers / decoration sub-chips stay plain. Status dot
  decoupled (showStatusDot prop) so picker rows can show presence
  without nesting popovers.

* Hover cards: AgentProfileCard simplified — availability dot only,
  Detail link top-right (logs live on the detail page). New
  MemberProfileCard mirrors the structure: name + role + email +
  top-2 owned agents (sorted by 30d run count) with click-through to
  agent detail.

* Agents list: split Status into two columns — availability (3-color
  dot + label) and Last run (task icon + label, optional running
  counts). Two independent filter chip groups (Status + Last run);
  combination acts as intersection ("online + failed" finds broken-
  but-alive agents).

* Other UI surfaces (issue list/board/detail, comments, autopilots,
  projects, runtimes, mention autocomplete, subscribers picker)
  updated to the new dot semantics; status dot now strictly 3-color.

Server changes accompany the client redesign — workspace-wide
agent-task-snapshot endpoint, runtime usage queries, etc. — to feed
the derive layer with the data it needs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(agent-detail): drop last-task chip from detail header + inspector

The Recent work section on the agent detail page already shows the same
data (with task titles, timestamps, error context) — surfacing
"Completed" / "Failed" / etc. up in the header was redundant chrome.
Detail surfaces now show only the 3-state availability dot.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tables): handle narrow viewports across agents / skills / runtimes

Three table layouts were squeezing content into adjacent cells at
intermediate widths. Each fix is small and targeted:

* runtime-list: the Runtime cell's base name had `shrink-0`, so it
  refused to truncate when its grid column was narrowed under width
  pressure — the name visually overflowed into the Health column
  ("ClaudeOnline" etc). Removed shrink-0, added truncate. The Health
  column was also a fixed 9.5rem reservation for the worst-case
  "Recently lost · 2m 14s ago" copy; switched to minmax(0,1fr) so it
  competes fairly with Runtime.

* skills-page: had a single grid template with no responsive
  breakpoints — all 6 columns were rendered at any width and got
  visually jammed below md. Added a <md template that drops Source +
  Updated; the row markup hides those cells via `hidden md:block` /
  `md:contents`.

* agent-list-item: the new Last run column was reserved at minmax(8rem,
  max-content); on narrow md viewports the 8rem floor pushed the row
  past available width. Changed to minmax(0,max-content) so the cell
  shrinks under pressure (its content already truncates).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(agent-card): hover-only Detail + add Runtime row + breathing room

Three small polish tweaks to the agent hover card:

- Detail link gets `mr-1` + fades in only on card hover (group-hover).
  It was visually flush against the popover edge and competing for
  attention; now it stays out of the way during a quick glance and
  surfaces only when the user is dwelling on the card.

- Runtime row is back, in the meta block (cloud/local icon + runtime
  name). The earlier removal was over-aggressive — knowing where an
  agent runs is part of "who is this agent". The wifi badge stays
  dropped because the availability dot in the header already conveys
  reachability.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(runtime): wifi-style health icon (4-state) for runtime list + agent card

Replaces the 6px coloured dot with a wifi-shape icon that carries both
state (Wifi vs WifiOff) and severity (success/warning/muted/destructive).

Mapping:
- online        → Wifi (success)
- recently_lost → WifiHigh (warning) — transient hiccup, fewer bars
- offline       → WifiOff (muted)    — long unreachable
- about_to_gc   → WifiOff (destructive) — sweeper coming soon

Used in two places:

- Runtime list: replaces HealthDot in the dedicated leading-icon column.
  Bumped the column from 0.5rem (dot-sized) to 0.875rem (icon-sized).

- Agent profile card RuntimeRow: derives runtime health from runtime +
  clock (matching the 4-state semantics) and renders HealthIcon next
  to the runtime name. Cloud runtimes always read as online. The
  duplicate signal with the header availability dot is intentional —
  it confirms WHICH runtime is the one currently in the dot's state.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 19:21:13 +08:00

205 lines
6.3 KiB
TypeScript
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"use client";
import { useMemo } from "react";
import { useQuery } from "@tanstack/react-query";
import type { Agent, AgentActivityBucket } from "../types";
import { agentListOptions } from "../workspace/queries";
import { agentActivity30dOptions } from "./queries";
const DAYS = 30;
const DAY_MS = 24 * 60 * 60 * 1000;
/** One day's tally for the sparkline. */
export interface ActivityBucket {
total: number;
failed: number;
}
export interface AgentActivity {
/**
* 30 daily buckets, oldest → newest. Days with no activity are
* zero-filled. Each surface picks how much of the tail to render: the
* Agents list uses 7, the agent detail uses all 30. Reading is the
* caller's job (see `summarizeActivityWindow` for the standard
* tail-slice + roll-up).
*/
buckets: ActivityBucket[];
/**
* Days the agent has existed, capped at DAYS. Pure cosmetic — used by
* tooltip copy ("Created 3 days ago"). The sparkline doesn't change
* shape for young agents on purpose; pre-life days look the same as
* zero days.
*/
daysSinceCreated: number;
}
/**
* Window-sized roll-up of an agent's activity series. Both the Agents
* list (windowDays=7) and the detail "Last 30 days" panel (windowDays=30)
* read through this so the totals can never drift from the bars they
* label.
*/
export interface ActivityWindowSummary {
/** Trailing-N buckets from the activity series (newest end). */
buckets: ActivityBucket[];
/** Sum of `bucket.total` across the window. */
totalRuns: number;
/** Sum of `bucket.failed` across the window. */
totalFailed: number;
/** Echo of the input window — the renderer uses it for copy. */
windowDays: number;
}
const EMPTY: AgentActivity = {
buckets: Array.from({ length: DAYS }, () => ({ total: 0, failed: 0 })),
daysSinceCreated: DAYS,
};
const EMPTY_SUMMARY: ActivityWindowSummary = {
buckets: [],
totalRuns: 0,
totalFailed: 0,
windowDays: 0,
};
/**
* Workspace-wide activity map keyed by `agent.id`. Single-pass batch:
* one fetch + one derivation pass backs every row's sparkline on the
* list AND the detail panel — adding rows costs O(1) HTTP and O(N)
* compute (not O(N) HTTP).
*/
export function useWorkspaceActivityMap(wsId: string | undefined): {
byAgent: Map<string, AgentActivity>;
loading: boolean;
} {
const { data: agents, isPending: agentsPending } = useQuery({
...agentListOptions(wsId ?? ""),
enabled: !!wsId,
});
const { data: buckets, isPending: bucketsPending } = useQuery({
...agentActivity30dOptions(wsId ?? ""),
enabled: !!wsId,
});
const byAgent = useMemo(() => {
if (!agents || !buckets) return new Map<string, AgentActivity>();
return buildActivityMap(agents, buckets, Date.now());
}, [agents, buckets]);
return { byAgent, loading: agentsPending || bucketsPending };
}
export function buildActivityMap(
agents: readonly Agent[],
buckets: readonly AgentActivityBucket[],
now: number,
): Map<string, AgentActivity> {
// Group buckets by agent once so per-agent derivation is O(buckets) not
// O(agents × buckets).
const bucketsByAgent = new Map<string, AgentActivityBucket[]>();
for (const b of buckets) {
const list = bucketsByAgent.get(b.agent_id);
if (list) list.push(b);
else bucketsByAgent.set(b.agent_id, [b]);
}
const out = new Map<string, AgentActivity>();
for (const agent of agents) {
out.set(
agent.id,
deriveAgentActivity(
bucketsByAgent.get(agent.id) ?? [],
agent.created_at,
now,
),
);
}
return out;
}
/**
* Pure derivation: filter the workspace-wide buckets to one agent and
* normalise to a fixed 30-element series ending at `now`. Exported for
* unit-testing and direct reuse on surfaces that already have the
* workspace-wide buckets in hand.
*/
export function deriveAgentActivity(
buckets: readonly AgentActivityBucket[],
agentCreatedAt: string,
now: number,
): AgentActivity {
const series: ActivityBucket[] = Array.from({ length: DAYS }, () => ({
total: 0,
failed: 0,
}));
// Newest slot is the start of "today" in local time; we walk back DAYS
// slots so index 0 = oldest, index DAYS-1 = today.
const today = startOfDay(now);
for (const b of buckets) {
const ts = new Date(b.bucket_at).getTime();
if (Number.isNaN(ts)) continue;
const daysAgo = Math.floor((today - startOfDay(ts)) / DAY_MS);
if (daysAgo < 0 || daysAgo >= DAYS) continue;
const slot = DAYS - 1 - daysAgo;
series[slot]!.total += b.task_count;
series[slot]!.failed += b.failed_count;
}
const createdAt = new Date(agentCreatedAt).getTime();
const ageMs = Number.isFinite(createdAt) ? now - createdAt : Infinity;
const daysSinceCreated = Math.min(
DAYS,
Math.max(0, Math.floor(ageMs / DAY_MS)),
);
return {
buckets: series,
daysSinceCreated,
};
}
/**
* Take the trailing N buckets and roll up totals over them. This is the
* single entry point both surfaces (list + detail) read through, so the
* numbers can never disagree with the bars they label.
*
* `windowDays` is clamped to the available bucket count, so passing a
* value larger than `activity.buckets.length` returns the full series
* rather than an out-of-range slice.
*/
export function summarizeActivityWindow(
activity: AgentActivity | undefined,
windowDays: number,
): ActivityWindowSummary {
if (!activity) return { ...EMPTY_SUMMARY, windowDays };
const safeWindow = Math.min(
Math.max(0, windowDays),
activity.buckets.length,
);
// `slice(-0)` returns the full array (JS quirk: -0 === 0), so guard
// explicitly when no window is requested.
const slice =
safeWindow === 0 ? [] : activity.buckets.slice(-safeWindow);
let totalRuns = 0;
let totalFailed = 0;
for (const b of slice) {
totalRuns += b.total;
totalFailed += b.failed;
}
return { buckets: slice, totalRuns, totalFailed, windowDays };
}
function startOfDay(ts: number): number {
// Local-time day boundary. The back-end truncates to UTC midnight, but
// the user's mental model is "today/yesterday in the timezone they're
// looking at"; using local matches that and keeps "today" stable across
// a working session even when buckets cross UTC midnight.
const d = new Date(ts);
d.setHours(0, 0, 0, 0);
return d.getTime();
}
export const __EMPTY_ACTIVITY = EMPTY;