mirror of
https://github.com/multica-ai/multica.git
synced 2026-08-04 17:18:35 +02:00
* perf(ci): split frontend build and test onto separate runners (MUL-5347) The frontend job is the critical path of every CI run that touches web code: p50 576s, versus 267s for backend. 94% of it is a single `turbo build typecheck lint test` invocation. Profiling showed the cost is contention, not inefficient tasks. A standard runner has 4 vCPUs, and the job scheduled two CPU-saturating tasks onto them concurrently: `@multica/views:test` (259 jsdom files) and `@multica/web:build` (a webpack production build). Measured against the same suites running with the machine to themselves: @multica/views:test 104s owning 4 cores -> 500s sharing them @multica/web:build 26s owning 10 cores -> 342s sharing 4 Split the work across two runners so neither starves the other. The split is weighted rather than even, because the halves are not close: the test group costs ~575 CPU-seconds and everything else ~250, so tests get a runner alone. Also drop `^typecheck` from the `test` and `lint` task definitions. No package in this repo emits build artifacts -- core/ui/views export raw .ts that the consumer transpiles -- so that edge ordered tasks without producing anything they consumed, and left the heaviest task in the graph idle for ~38s while core/ui ran `tsc --noEmit`. Removing it also drops three redundant typecheck tasks from the test job, taking it from 7 tasks / 798 CPU-seconds to 4 / 575. Type errors still fail CI through the `typecheck` task, which frontend-build owns. `frontend` is retained as an aggregate gate so the existing status-check name keeps reporting, including the established behaviour that it goes green rather than pending on PRs the path filter excludes. Co-authored-by: multica-agent <github@multica.ai> * perf(ci): cache turbo task results across runs (MUL-5347) Every CI run rebuilt all 16 frontend tasks cold -- turbo reported `Remote caching disabled` / `Cached: 0 cached, 16 total` on every run, because only the pnpm store was cached. Persist turbo's filesystem cache with actions/cache instead. No Vercel remote-cache credentials exist in this repo, so this is the local cache keyed per job. Measured, simulating a fresh checkout against a warm cache: frontend-build 54.7s -> 184ms (12/12 cached) frontend-test 115.7s -> 50ms (7/7 cached) Cache footprint is 20MB for the build job and ~60KB for the test job -- `test` declares no outputs, so turbo stores exit codes and logs, yet a hit still skips the most expensive task in the graph. Turning the cache on promotes two latent hashing bugs into real ones, so both are fixed here: `build` declared narrow `inputs` globs that matched neither `content/**/*.mdx` (compiled by fumadocs-mdx during the build), `public/**` assets, nor `package.json` / `tsconfig.json`. Editing any of them left the task hash byte-identical, which is inert when nothing is ever restored and a stale-build bug the moment something is. Dropped in favour of turbo's default input set. `test` regains its `^typecheck` edge, reverting part of the previous commit. That change was justified on the basis that the edge only imposed ordering, which was true without a cache: a turbo task hash covers a workspace dependency's sources only if a task edge reaches them, so with no edge, editing packages/views or packages/ui left `@multica/web#test` and `@multica/desktop#test` unchanged in hash and a cached pass would be replayed over changed code. `typecheck` is the only task every package defines, so it is the edge that closes the gap; `^test` would miss packages/ui, which has no test script. `lint` keeps no edge -- eslint here is not type-aware, so it cannot observe a dependency's sources. Also adds .github/workflows/ci.yml to globalDependencies: it pins the Node version, and without it a toolchain bump would replay results produced by the old runtime behind a green check. Co-authored-by: multica-agent <github@multica.ai> * fix(ci): scope turbo cache to the resolved runtime, drop typecheck from tests (MUL-5347) Addresses review feedback on the caching commit. Must-fix: the cache was not isolated by interpreter. `node-version: 22` floats across patch releases and turbo's global hash does not include the interpreter at all (`engines` is null in its dry-run cache inputs), so a runner moving to another 22.x would leave the workflow text and every task hash unchanged, restore the old cache through `restore-keys` and report green without executing anything. Both jobs now resolve `node --version` into a step output and carry it, plus `runner.arch`, in the cache key and the restore prefix. Must-fix: `actions/cache@v4` runs on the deprecated Node 20 runtime and was being force-migrated to Node 24 with a warning on every run. Moved to v6 -- the review suggested v5, but v6.1.0 is current and satisfies the same runner floor (>= 2.327.1; hosted runners are on 2.335.1). `test` no longer depends on `^typecheck`. It needs dependency sources in its hash, not a type check: a new hash-only `cache-inputs` transit task carries them. No package implements that script, so every node resolves to <NONEXISTENT> and nothing executes, while the edges still pull each dependency's files into the hash. Declared recursively so the chain survives a package gaining a purely transitive dependency. This returns the ~223 CPU-seconds the previous commit gave up: the cold test job goes from 7 tasks / 115.7s back to 4 tasks / 67.2s, and editing packages/ui still correctly re-runs the views, web and desktop suites while core:test stays cached. Corrects a comment that claimed `^test` would miss packages/ui because it has no test script. It would not -- turbo materialises a <NONEXISTENT> node that participates in hashing, verified against the dry graph. `^test` is still the wrong edge here, but because it serialises the suites behind each other, not for the stated reason. Also drops the stale note that tests no longer depend on `^typecheck`, and swaps the aggregate gate's `always()` for `!cancelled()` so a run superseded by a newer push does not spend a runner on a verdict nobody reads. Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: Bohan-J <bohan@devv.ai> Co-authored-by: multica-agent <github@multica.ai>
365 lines
15 KiB
YAML
365 lines
15 KiB
YAML
name: CI
|
|
|
|
on:
|
|
push:
|
|
branches: [main]
|
|
pull_request:
|
|
branches: [main]
|
|
|
|
concurrency:
|
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
|
cancel-in-progress: true
|
|
|
|
jobs:
|
|
# Decides whether the (heavy, ~6min) frontend job has anything to do.
|
|
# The frontend job validates the web/desktop apps, the shared packages,
|
|
# the install graph, and the selfhost / reserved-slugs scripts it runs;
|
|
# a pure backend-only or docs-only PR touches none of those and gains
|
|
# nothing from a full web build. This job emits a single `frontend`
|
|
# output consumed by the frontend job below.
|
|
changes:
|
|
runs-on: ubuntu-latest
|
|
permissions:
|
|
contents: read
|
|
pull-requests: read
|
|
outputs:
|
|
frontend: ${{ steps.decide.outputs.frontend }}
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v6
|
|
|
|
- name: Filter paths
|
|
id: filter
|
|
uses: dorny/paths-filter@v3
|
|
with:
|
|
# apps/docs is excluded from the frontend turbo run, so a
|
|
# docs-only change does not need this job. apps/mobile has its
|
|
# own mobile-verify workflow. Everything else the frontend job
|
|
# touches is listed here; bias toward over-matching since a
|
|
# missed path silently skips validation.
|
|
filters: |
|
|
frontend:
|
|
- 'apps/web/**'
|
|
- 'apps/desktop/**'
|
|
- 'packages/**'
|
|
- 'package.json'
|
|
- '.npmrc'
|
|
- 'pnpm-lock.yaml'
|
|
- 'pnpm-workspace.yaml'
|
|
- 'turbo.json'
|
|
- '.github/workflows/ci.yml'
|
|
- 'scripts/generate-reserved-slugs.mjs'
|
|
- 'server/internal/handler/reserved_slugs.json'
|
|
- 'scripts/selfhost-config.test.sh'
|
|
- 'scripts/check.sh'
|
|
- 'scripts/dev.sh'
|
|
- 'scripts/local-env.sh'
|
|
- '.env.example'
|
|
- 'docker-compose.selfhost.yml'
|
|
|
|
- name: Decide
|
|
id: decide
|
|
# Always run the frontend job on push to main (full validation);
|
|
# on pull_request, run only when frontend-relevant paths changed.
|
|
# The frontend job itself always runs and reports success — its
|
|
# steps are gated on this output rather than the job being skipped
|
|
# — so the required "frontend" status check is satisfied with a
|
|
# genuine green instead of being left pending on filtered PRs.
|
|
env:
|
|
EVENT_NAME: ${{ github.event_name }}
|
|
FRONTEND_CHANGED: ${{ steps.filter.outputs.frontend }}
|
|
run: |
|
|
if [ "$EVENT_NAME" != "pull_request" ] || [ "$FRONTEND_CHANGED" = "true" ]; then
|
|
echo "frontend=true" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "frontend=false" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# The frontend validation is split across two runners on purpose. Both
|
|
# `@multica/web:build` (a webpack production build) and `@multica/views:test`
|
|
# (259 jsdom files) are CPU-saturating, and a standard runner only has
|
|
# 4 vCPUs. Running them in one job made them starve each other: the views
|
|
# suite needs ~104s wall when it owns 4 cores but took ~500s sharing them,
|
|
# and the identical webpack compile went from ~26s to ~342s. Splitting buys
|
|
# a second 4-vCPU box rather than reducing the work; total runner-minutes go
|
|
# up slightly, wall-clock feedback time goes down.
|
|
#
|
|
# The split is weighted, not even: the test group is by far the heavier half
|
|
# (~575 CPU-seconds vs ~250 for build + typecheck + lint), so it gets a
|
|
# runner to itself and everything else shares the other one.
|
|
frontend-build:
|
|
needs: changes
|
|
runs-on: ubuntu-latest
|
|
env:
|
|
# Pin turbo's filesystem cache somewhere actions/cache can address.
|
|
TURBO_CACHE_DIR: .turbo/cache
|
|
steps:
|
|
- name: Checkout
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
uses: actions/checkout@v6
|
|
|
|
- name: Setup pnpm
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
uses: pnpm/action-setup@v4
|
|
|
|
- name: Setup Node.js
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
uses: actions/setup-node@v6
|
|
with:
|
|
node-version: 22
|
|
cache: pnpm
|
|
|
|
- name: Install dependencies
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
run: pnpm install
|
|
|
|
# `node-version: 22` above floats across patch releases, and turbo's
|
|
# global hash does not include the interpreter at all (`engines` is null
|
|
# in its dry-run cache inputs). Without the resolved version in the key,
|
|
# a runner silently moving to another 22.x would restore a cache built by
|
|
# the old interpreter and report green without executing anything.
|
|
- name: Resolve runtime for cache key
|
|
id: runtime
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
run: echo "node=$(node --version)" >> "$GITHUB_OUTPUT"
|
|
|
|
# Cache entries are immutable, so the key carries the commit SHA to make
|
|
# every run publish a fresh one and `restore-keys` falls back to the most
|
|
# recent prefix match. The two frontend jobs run disjoint task sets, so
|
|
# they get their own prefixes rather than racing to save one key.
|
|
# GitHub scopes caches by branch: a PR reads main's entries (so unchanged
|
|
# tasks hit on the first push) and writes its own (so re-pushes hit too).
|
|
- name: Restore turbo cache
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
uses: actions/cache@v6
|
|
with:
|
|
path: .turbo/cache
|
|
key: turbo-build-${{ runner.os }}-${{ runner.arch }}-${{ steps.runtime.outputs.node }}-${{ github.sha }}
|
|
restore-keys: |
|
|
turbo-build-${{ runner.os }}-${{ runner.arch }}-${{ steps.runtime.outputs.node }}-
|
|
|
|
- name: Test self-host env derivation
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
run: bash scripts/selfhost-config.test.sh
|
|
|
|
- name: Verify reserved-slugs.ts is up to date
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
# Re-runs the generator and fails on any drift from the
|
|
# checked-in TypeScript output. The Go side embeds the JSON
|
|
# source directly, so a passing diff here proves both sides
|
|
# share one source of truth.
|
|
run: |
|
|
pnpm generate:reserved-slugs
|
|
git diff --exit-code -- packages/core/paths/reserved-slugs.ts
|
|
|
|
- name: Build, type check, and lint
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
# Mobile lives in a parallel mobile-verify workflow (path-filtered
|
|
# to apps/mobile/** + packages/core/**) so it doesn't add
|
|
# ~50s of expo-lint + tsc to every web/desktop PR. Keep this
|
|
# filter in sync with the root package.json scripts, which also
|
|
# exclude @multica/mobile.
|
|
run: pnpm exec turbo build typecheck lint --filter='!@multica/docs' --filter='!@multica/mobile'
|
|
|
|
frontend-test:
|
|
needs: changes
|
|
runs-on: ubuntu-latest
|
|
env:
|
|
TURBO_CACHE_DIR: .turbo/cache
|
|
steps:
|
|
- name: Checkout
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
uses: actions/checkout@v6
|
|
|
|
- name: Setup pnpm
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
uses: pnpm/action-setup@v4
|
|
|
|
- name: Setup Node.js
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
uses: actions/setup-node@v6
|
|
with:
|
|
node-version: 22
|
|
cache: pnpm
|
|
|
|
- name: Install dependencies
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
run: pnpm install
|
|
|
|
- name: Resolve runtime for cache key
|
|
id: runtime
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
run: echo "node=$(node --version)" >> "$GITHUB_OUTPUT"
|
|
|
|
# See frontend-build for the key strategy. These entries are tiny (~60KB
|
|
# measured): `test` declares no outputs, so turbo caches exit codes and
|
|
# logs rather than artifacts -- yet a hit still skips the whole suite,
|
|
# which is the single most expensive task in the graph.
|
|
- name: Restore turbo cache
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
uses: actions/cache@v6
|
|
with:
|
|
path: .turbo/cache
|
|
key: turbo-test-${{ runner.os }}-${{ runner.arch }}-${{ steps.runtime.outputs.node }}-${{ github.sha }}
|
|
restore-keys: |
|
|
turbo-test-${{ runner.os }}-${{ runner.arch }}-${{ steps.runtime.outputs.node }}-
|
|
|
|
- name: Test
|
|
if: ${{ needs.changes.outputs.frontend == 'true' }}
|
|
# Same filter rationale as frontend-build. Type errors are not this
|
|
# job's responsibility -- frontend-build owns the `typecheck` task.
|
|
# `test` reaches dependency sources through the hash-only
|
|
# `^cache-inputs` edge (see turbo.json), so no `tsc` runs here.
|
|
run: pnpm exec turbo test --filter='!@multica/docs' --filter='!@multica/mobile'
|
|
|
|
# Aggregate gate. `frontend` is the status-check name the repository's
|
|
# branch rules refer to, so it has to survive the split above: this job
|
|
# keeps reporting under that name and simply fails when either half fails.
|
|
# It also inherits the old contract that the check goes green (rather than
|
|
# staying pending) on PRs the path filter excluded — both halves succeed
|
|
# trivially in that case because every step is gated off.
|
|
frontend:
|
|
needs: [frontend-build, frontend-test]
|
|
# `!cancelled()` rather than `always()`: a run superseded by a newer push
|
|
# is cancelled by the concurrency group above, and there is no reason to
|
|
# spend a runner reporting a verdict nobody will read.
|
|
if: ${{ !cancelled() }}
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- name: Check frontend job results
|
|
env:
|
|
BUILD_RESULT: ${{ needs.frontend-build.result }}
|
|
TEST_RESULT: ${{ needs.frontend-test.result }}
|
|
run: |
|
|
echo "frontend-build: $BUILD_RESULT"
|
|
echo "frontend-test: $TEST_RESULT"
|
|
if [ "$BUILD_RESULT" != "success" ] || [ "$TEST_RESULT" != "success" ]; then
|
|
echo "::error::frontend validation failed"
|
|
exit 1
|
|
fi
|
|
|
|
backend:
|
|
runs-on: ubuntu-latest
|
|
services:
|
|
postgres:
|
|
image: pgvector/pgvector:pg17
|
|
env:
|
|
POSTGRES_DB: multica
|
|
POSTGRES_USER: multica
|
|
POSTGRES_PASSWORD: multica
|
|
ports:
|
|
- 5432:5432
|
|
options: >-
|
|
--health-cmd "pg_isready -U multica -d multica"
|
|
--health-interval 5s
|
|
--health-timeout 5s
|
|
--health-retries 20
|
|
redis:
|
|
image: redis:7-alpine
|
|
ports:
|
|
- 6379:6379
|
|
options: >-
|
|
--health-cmd "redis-cli ping"
|
|
--health-interval 5s
|
|
--health-timeout 5s
|
|
--health-retries 10
|
|
env:
|
|
DATABASE_URL: postgres://multica:multica@localhost:5432/multica?sslmode=disable
|
|
# Wires up the RedisLocalSkill*_test.go suite. Distinct from REDIS_URL
|
|
# (which would flip the server binary itself onto the Redis-backed
|
|
# realtime relay + request stores); the tests talk to this Redis
|
|
# directly so they run alongside the Postgres-backed suite.
|
|
REDIS_TEST_URL: redis://localhost:6379/1
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v6
|
|
|
|
- name: Setup Go
|
|
uses: actions/setup-go@v5
|
|
with:
|
|
go-version: "1.26.1"
|
|
cache-dependency-path: server/go.sum
|
|
|
|
- name: Setup Helm
|
|
uses: azure/setup-helm@v4
|
|
|
|
- name: Test Helm chart
|
|
run: bash scripts/helm-config.test.sh
|
|
|
|
- name: Build
|
|
run: cd server && go build ./...
|
|
|
|
- name: Run migrations
|
|
run: cd server && go run ./cmd/migrate up
|
|
|
|
- name: Verify Go test wrapper
|
|
run: bash scripts/test-go.test.sh
|
|
|
|
- name: Test
|
|
run: bash scripts/test-go.sh --race
|
|
|
|
windows-execenv:
|
|
# The environment-preparation deadline owns a process tree, not just a Go
|
|
# process. This Windows runtime test verifies Job Object cancellation kills
|
|
# a delayed descendant before an immediate retry can reuse the same root.
|
|
runs-on: windows-latest
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v6
|
|
|
|
- name: Setup Go
|
|
uses: actions/setup-go@v5
|
|
with:
|
|
go-version: "1.26.1"
|
|
cache-dependency-path: server/go.sum
|
|
|
|
- name: Test Windows execution-environment isolation
|
|
working-directory: server
|
|
# Keep this job scoped to the runtime regression it exists to prove.
|
|
# The package's legacy OpenClaw HOME tests are not Windows-safe and are
|
|
# outside this PR; the normal backend job still runs the full package.
|
|
run: go test ./internal/daemon/execenv -run '^TestPrepareIsolated_WindowsKillsDescendantBeforeRetry$' -count=1 -timeout=5m
|
|
|
|
- name: Test Windows agent launcher argv/stdin handling
|
|
working-directory: server
|
|
# Agent prompts must never reach a Windows launcher through argv: the
|
|
# official cursor-agent.ps1 ends in `& node.exe index.js $args`, and
|
|
# PowerShell re-serialises $args onto the child command line. Under
|
|
# Legacy native argument passing (powershell.exe 5.1, pwsh <= 7.2) a
|
|
# prompt holding embedded quotes is re-tokenised and fragments like
|
|
# `-X` become flags (#5649). Only a real PowerShell host proves this,
|
|
# so it cannot live in the ubuntu backend job. Scoped to the launcher
|
|
# tests, which are windows-tagged and therefore run nowhere else today;
|
|
# the backend job still runs the full package on Linux.
|
|
# -v so a silent skip (no PowerShell host resolved, or a -run pattern
|
|
# that stops matching) is visible in the log instead of passing as "ok".
|
|
run: go test ./pkg/agent -v -run '^(TestCursorExecutePromptSurvivesPowerShellShim|TestPlatformCursorInvocation|TestPlatformCopilotInvocation|TestPlatformPiInvocation)' -count=1 -timeout=5m
|
|
|
|
- name: Test bounded Codex cleanup with inherited stdout descendant
|
|
working-directory: server
|
|
# Windows cannot prove whole-tree termination without a Job Object,
|
|
# but a descendant holding inherited stdout must never keep Result
|
|
# blocked forever. -v makes RUN/PASS evidence explicit in CI logs.
|
|
run: go test ./pkg/agent -v -run '^TestCodexWindowsInheritedStdoutDescendantCleanupIsBounded$' -count=1 -timeout=5m
|
|
|
|
- name: Build Windows CLI helper entrypoint
|
|
working-directory: server
|
|
run: go build ./cmd/multica
|
|
|
|
installer:
|
|
# Stub-driven shell tests for scripts/install.sh. Kept off the heavy
|
|
# backend job so installer regressions surface independently, and
|
|
# exercised on macOS too because the installer targets macOS/Homebrew
|
|
# and `tar` / `sed` / `mktemp` differ between BSD and GNU userlands.
|
|
strategy:
|
|
fail-fast: false
|
|
matrix:
|
|
os: [ubuntu-latest, macos-latest]
|
|
runs-on: ${{ matrix.os }}
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v6
|
|
|
|
- name: Test shell installers
|
|
run: bash scripts/install.test.sh
|