mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-05 01:19:42 +02:00
* fix(github): surface CI status on PR cards (MUL-5180) The CI mirroring pipeline (MUL-2228, MUL-2392) has never received a single event in production. The GitHub App setup docs only ever asked operators to grant `pull_requests: read` and subscribe to `pull_request`, so GitHub never delivered `check_suite` — `handleCheckSuiteEvent` sat dead behind a subscription nobody was told to enable. Every linked PR reports checks_passed/failed/pending = 0 and the sidebar row falls through to "Checks haven't reported yet" forever. Docs (the root cause), all four locales: - add `Checks: Read-only` permission + `Check suite` event to the App setup table - drop the stale "CI check states are not modeled" claim, which predates MUL-2228 and is what let the setup table stay incomplete - add a "PR rows show no CI status" troubleshooting entry with the public `/apps/<slug>` probe to confirm what an App is actually subscribed to, and a warning that existing installations must accept the new permission before any `check_suite` is delivered UI: - give the actionable status kinds (checks failed/pending/passed, conflicts, ready) their own icon + color. CI outcome previously rendered as plain muted 11px text, visually identical to the diff stats beside it — a failing build read the same as "+437 −6 · 6 files". Terminal and unknown kinds stay muted; the row's state icon already carries that meaning. Co-authored-by: multica-agent <github@multica.ai> * fix(github): unbreak docs build, stop overclaiming CI completeness (MUL-5180) Both must-fixes from review. 1. docs production build failed. `<your App>` in prose was parsed as a JSX tag, so `pnpm --filter @multica/docs build` died with `Expected a closing tag for <your>`. Dropped the angle brackets. Repo CI never caught this because no workflow runs the docs production build — only Vercel does, which is why the PR's GitHub checks were green while the deployment errored. 2. `Checks: Read-only` cannot support the pending status the docs promised. GitHub's webhook contract delivers `check_suite.requested` / `.rerequested` only to Apps holding Checks *write*; read-level access receives `completed` only. Verified against GitHub's published docs. Direction chosen: keep read-only, degrade honestly to final-results-only. Checks *write* is a repo-write capability (create/update check runs), not a wider read — escalating every installation to it just to render an in-flight spinner is not a trade to make on the operator's behalf, and it contradicts the integration's read-only posture. The concrete bug this leaves is premature green: with two reporting apps, the first to complete makes total=1/passed=1 and the row claimed "All checks passed" while the second was still running and might fail. Copy is now "Checks passed" in all four locales — it reports what reported and never asserts completeness. `derivePullRequestStatusKind` documents why. Docs gain a "what CI status can and cannot tell you" section (all four locales) with the read-vs-write delivery table, both consequences stated plainly, and the opt-in path for teams that do want in-flight status: set Checks to Read and write on their own App and the existing pending code lights up with no code change. The pending promise is removed from the read-only setup path. Co-authored-by: multica-agent <github@multica.ai> * fix(github): ignore non-completed check_suite actions (MUL-5180) Review was right: the `Read and write` opt-in the previous commit documented does not produce reliable pending, and following it would break the card. `check_suite.requested` / `.rerequested` are not observations that some CI provider started. GitHub sends them only to Apps holding Checks write, and per the CI-checks App docs they mean "GitHub has created a check suite for YOUR app on this commit; now add your check runs to it". Multica observes other apps' results and never creates check runs. Recording such a suite parks a `queued` row nothing can ever complete, and since `checks_pending` outranks `checks_passed` in derivePullRequestStatusKind, one stuck row freezes every PR on that installation at "checks running" and hides the real pass/fail result. Any self-hoster who already grants Checks write hits this on every push, so the gate is on the action, not the permission. - handleCheckSuiteEvent drops every action except `completed`, with the reasoning and the "don't resurrect requested as a running signal" warning recorded at the gate. - TestWebhook_CheckSuite_QueuedCountsAsPending encoded the wrong delivery semantics (two external apps sending `requested`, which GitHub never does). Replaced by TestWebhook_CheckSuite_NonCompletedActionsIgnored, which pins the drop and checks a later `completed` suite still lands. - The two out-of-order stash tests used `requested` payloads to exercise paths that are really about completed suites; both now use `completed` and assert the same guarantees. - Docs (four locales): the write opt-in is gone. In-flight CI is documented as unsupported at any permission level, with the actual reason and the note that real running status needs polling or a check_run model instead. Co-authored-by: multica-agent <github@multica.ai> * fix(github): make legacy non-completed check suites inert (MUL-5180) Review was right again: the previous commit gated the webhook entry point but left the pre-upgrade state — and the people it was meant to protect (self- hosters who already granted Checks write) are exactly the ones holding it. Two leftovers, both now closed: 1. Rows already in github_pull_request_check_suite. The old handler stored GitHub's `requested` suites as `queued`; nothing will ever complete them. ListPullRequestsByIssue still counted them, so `checks_pending` kept outranking `checks_passed` and the PR stayed pinned to "checks running" for as long as its head SHA stood. The aggregation now selects only `completed` suites. Filtering beats deleting here: recovery is automatic on deploy, needs no migration over a table that can be large, and holds for any writer that misses a gate — not just for today's legacy rows. DISTINCT ON runs after the filter, so an app whose newest suite is a stuck `queued` still reports its most recent completed verdict instead of disappearing. 2. Rows already in github_pending_check_suite. replayPendingCheckSuitesForPR is a second write path into the live table that never passes through handleCheckSuiteEvent, so the next `pull_request` event would re-inject a permanently-queued suite after the fix shipped. It now skips non-completed rows; the drain is DELETE ... RETURNING, so skipping discards them. Both are covered by regression tests that seed the legacy row directly — the fixed handler can no longer produce one — and both were confirmed to fail with their respective fix reverted. The stash test additionally asserts its fixture landed under the repo address the drain keys on; the first draft used the wrong owner and passed vacuously. Also corrects the aggregateChecksConclusion doc comment, which still described "pending" as a not-yet-completed suite. It is now reachable only for a completed suite carrying a null conclusion, and is explicitly not a "CI is running" signal. Co-authored-by: multica-agent <github@multica.ai> * test(github): assert the legacy stash row is consumed by the drain (MUL-5180) Review nits. The stash test proved its fixture existed before the webhook but never that the drain consumed it, so a future change to firePullRequestWebhookWithHead's repo address would make the assertions pass for the wrong reason again — the same way the first draft of this test did. Asserting the stash is empty afterwards closes that gap from the other side. Also fixes two comment typos: `an "CI is running"` -> `a`, and drops the "merged-but-open PR" state, which cannot exist. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai>
114 lines
4.9 KiB
TypeScript
114 lines
4.9 KiB
TypeScript
import type { GitHubPullRequest } from "../types";
|
||
|
||
// Status kinds rendered in the PR sidebar row's detail line. Order in the
|
||
// pass-through table matters — the first matching rule wins. The order is
|
||
// chosen so terminal PR states (closed / merged) short-circuit before any
|
||
// transient CI/conflict signal, since those signals are no longer actionable
|
||
// on a terminal PR.
|
||
//
|
||
// Priority (high → low):
|
||
// 1. closed (not merged) → status_closed
|
||
// 2. merged → status_merged
|
||
// 3. mergeable_state = "dirty" → status_conflicts
|
||
// 4. any failed suite → status_checks_failed
|
||
// 5. any pending suite → status_checks_pending
|
||
// 6. any passed suite → status_checks_passed
|
||
// 7. no suite + mergeable=clean → status_ready
|
||
// 8. otherwise → status_unknown
|
||
//
|
||
// Completeness caveat (MUL-5180): these counts only ever cover check suites
|
||
// Multica has actually observed, which is NOT the same as "every suite that
|
||
// will run". GitHub delivers `check_suite.requested` / `.rerequested` only to
|
||
// Apps holding Checks *write*; the Multica App deliberately stays read-only,
|
||
// so in practice only `completed` suites arrive. A PR with two reporting apps
|
||
// therefore passes through a window where the first has completed and the
|
||
// second has not yet been seen at all.
|
||
//
|
||
// That is why rule 6 renders as "Checks passed" and not "All checks passed" —
|
||
// we can report what reported, never that everything did. Do not reintroduce
|
||
// completeness wording here without also fixing the underlying signal.
|
||
//
|
||
// Note: this table is the single source of truth for the sidebar PR row. The
|
||
// older row-with-badges implementation used a separate "hide status row for
|
||
// terminal PRs" branch — the current row renders
|
||
// with status_closed / status_merged text, never falling through to a
|
||
// conflicts / checks line on a terminal PR. Keep this priority order in sync
|
||
// with the i18n keys `pull_request_card_status_*` and with the progress-strip
|
||
// derivation in `derivePullRequestProgressSegments` (terminal kinds get a
|
||
// solid bar; the rest map onto the per-suite counts).
|
||
export type PullRequestStatusKind =
|
||
| "closed"
|
||
| "merged"
|
||
| "conflicts"
|
||
| "checks_failed"
|
||
| "checks_pending"
|
||
| "checks_passed"
|
||
| "ready"
|
||
| "unknown";
|
||
|
||
export interface PullRequestStatusInput {
|
||
state: GitHubPullRequest["state"];
|
||
mergeable_state?: string | null;
|
||
checks_failed?: number;
|
||
checks_pending?: number;
|
||
checks_passed?: number;
|
||
}
|
||
|
||
export function derivePullRequestStatusKind(input: PullRequestStatusInput): PullRequestStatusKind {
|
||
if (input.state === "closed") return "closed";
|
||
if (input.state === "merged") return "merged";
|
||
if (input.mergeable_state === "dirty") return "conflicts";
|
||
if ((input.checks_failed ?? 0) > 0) return "checks_failed";
|
||
if ((input.checks_pending ?? 0) > 0) return "checks_pending";
|
||
if ((input.checks_passed ?? 0) > 0) return "checks_passed";
|
||
if (input.mergeable_state === "clean") return "ready";
|
||
return "unknown";
|
||
}
|
||
|
||
export interface PullRequestProgressSegment {
|
||
kind: "failed" | "pending" | "passed";
|
||
ratio: number;
|
||
}
|
||
|
||
// Segmented progress bar input. Returns null when:
|
||
// - the PR is terminal (closed/merged) — the card paints a solid bar
|
||
// in a state-specific color, no segmentation needed;
|
||
// - no check_suite has been observed (total === 0) — the card hides
|
||
// the bar entirely.
|
||
// Otherwise emits the segments left-to-right: failed → pending → passed.
|
||
// "Failure first" is intentional: problems should be visible before signal
|
||
// that everything is fine.
|
||
export function derivePullRequestProgressSegments(
|
||
input: PullRequestStatusInput,
|
||
): PullRequestProgressSegment[] | null {
|
||
if (input.state === "closed" || input.state === "merged") return null;
|
||
const failed = input.checks_failed ?? 0;
|
||
const pending = input.checks_pending ?? 0;
|
||
const passed = input.checks_passed ?? 0;
|
||
const total = failed + pending + passed;
|
||
if (total === 0) return null;
|
||
const segments: PullRequestProgressSegment[] = [];
|
||
if (failed > 0) segments.push({ kind: "failed", ratio: failed / total });
|
||
if (pending > 0) segments.push({ kind: "pending", ratio: pending / total });
|
||
if (passed > 0) segments.push({ kind: "passed", ratio: passed / total });
|
||
return segments;
|
||
}
|
||
|
||
export interface PullRequestStatsInput {
|
||
additions?: number;
|
||
deletions?: number;
|
||
changed_files?: number;
|
||
}
|
||
|
||
// shouldShowPullRequestStats encodes the "old backend → new frontend" guard:
|
||
// when the backend that served this PR row doesn't know about the stats
|
||
// columns yet, every numeric field defaults to 0. Rendering "+0 −0 · 0 files"
|
||
// in that case would be a lie (the PR almost certainly has real changes),
|
||
// so we hide the entire stats row until at least one signal is non-zero.
|
||
export function shouldShowPullRequestStats(input: PullRequestStatsInput): boolean {
|
||
const a = input.additions ?? 0;
|
||
const d = input.deletions ?? 0;
|
||
const f = input.changed_files ?? 0;
|
||
return a + d + f > 0;
|
||
}
|