# CI/CD Optimizations, Part 1: PR Checks

> Two cache bugs, one lint flag, and a vulnerability database downloaded forty times a day. Fixing them cut Go jobs by 65% and PR time barely moved. Finding the job that actually set PR time took it from 5.3 minutes to 2.5.

- Date: 2026-10-10
- Tags: go, ci, github-actions, devops, performance
- Source: https://logicalbytes.dev/posts/ci-critical-path/

---
Statio's monorepo has a Go backend, a React web app, a desktop client, and a pile of GitHub Actions workflows that run on every pull request. I've seen a lot of talk on X recently about optimizing (usually a rust migration). So I decided that I might as well take a crack at what I have going on. Over the last week I spent a few evenings making those workflows faster. Go Build & Test went from 5.2 minutes to 1.8. Go Lint went from 6.0 to 3.0. The time from opening a PR to seeing all green went from about 5.3 to 2.5 which was slightly better than my 3 minute target which was based on looking away for a minute then every check had completed.

All numbers below are medians of successful jobs, pulled from the GitHub API across 493 workflow runs between August 5 and October 7, plus 125 PR runs from October 7 to 9 for the last section. Runners are the standard `ubuntu-latest` hosted ones.


## The PR pipeline

Three workflows trigger on a pull request:

- **CI** (`ci.yml`): Go Build & Test, Go Lint, Frontend Check, OpenAPI codegen drift, a contract test against platform-core, deploy config validation, and a PR title check.
- **Security Scan**: Grype against the Go modules and against the Node lockfile.
- **E2E MCP Pipeline**: builds the gateway, stands up a test MCP server, and runs the pipeline end to end.

Everything runs in parallel. The important thing to remember is that **the PR is done when the slowest job is done.** Shaving time off any other job changes nothing you can feel.


## Mid-September

In mid-September the slowest jobs were the Go ones. Go Lint ran 3.8 minutes, Go Build & Test 3.6, and Frontend Check about a minute. A typical PR finished in about 3.6 to 4 minutes, and Go Lint or Go Build was the last job to finish in 47 of 49 runs. This was good enough so I didn't worry about it. However the number began climbing as I added deeper checks. 

**Frontend tests landed in CI (#118, Sep 17).** Before this, Frontend Check ran `tsc` and nothing else, so the vitest suite could go red and no check would notice. After it, the job also ran vitest and an eslint ratchet (`--max-warnings 441`, a number that's only allowed to go down). At the time the suite had 9 test files and the job went to about 2 minutes. That's a fine trade.

**Lint lost its cache (#124, Oct 1).** `golangci-lint` caches analysis results between runs. Ours flapped: `nolintlint` would report a `//nolint` directive as unused on one run and fine on the next, depending on what the cache held. A lint check that fails at random teaches everyone to hit re-run. I set `skip-cache: true` and Go Lint went from 3.8 to 5.0 minutes. I'd make that call again. A slow honest check beats a fast one nobody trusts.

**The web app grew.** This one wasn't a CI change at all. Between October 1 and October 6 we shipped onboarding, a help menu, container registry UI, new pricing pages, and tests for all of it. Test files went from 22 to 42. Frontend Check's daily median went 2.5, 2.8, 3.6, 3.8, 4.4, 5.0 minutes. A good consequence of ensuring we had testing coverage and ensured that nothing would break in the future, bad because now we were at over 5 minutes, not where I wanted to be.

By October 6 a PR took about 5.3 minutes and every job felt slow. So I started with what I knew best, Go, because Go was the part I was sure was broken.


## Bug one: a cache key that can hit never updates

The Go jobs used `actions/setup-go` with its built-in cache. It keys the cache on a hash of `go.sum`. Which seemed right until I saw cache misses and looked into how `actions/cache` decides whether to save:

**It only saves when the primary key missed.**

So the first run after a `go.sum` change misses, builds, and saves. Every run after that hits the same key, restores that snapshot, and never saves again. The module cache stays fine because modules only change when `go.sum` does. The build cache, `~/.cache/go-build`, goes stale. Every commit after the snapshot recompiles the packages that changed since then, and the list grows with every PR until someone bumps a dependency.

The fix (#166) turns off setup-go's cache and uses `actions/cache` directly, with `github.run_id` at the end of the key:

```yaml
- uses: actions/setup-go@v7
  with:
    go-version-file: go.mod
    cache: false

- name: Go cache
  uses: actions/cache@v4
  with:
    path: |
      ~/.cache/go-build
      ~/go/pkg/mod
    key: go-${{ runner.os }}-${{ hashFiles('**/go.sum') }}-${{ github.run_id }}
    restore-keys: |
      go-${{ runner.os }}-${{ hashFiles('**/go.sum') }}-
      go-${{ runner.os }}-
```

The run ID guarantees the primary key misses every time, so every run saves. `restore-keys` does prefix matching and returns the newest entry, so every run starts from the most recent cache. Go's build cache is content-addressed, so a stale entry is never used for a file that changed, it sits there until GitHub evicts it.

Proud of myself for the fix I merged it and the numbers got worse. Go Build went from 4.6 to 5.2 minutes and Go Lint from 5.0 to 6.0. That was expected because we had some other bugs that were preventing this fix from fully working.


## Bug two: seven jobs, one key, and the fastest one wins

Every job in `ci.yml` that touches Go used that same cache step, so every job in a run computed the same key. GitHub's cache won't let two writers reserve one key. The first job to reach its post step saves. The rest log this and move on:

```
Failed to save: Unable to reserve cache with key go-Linux-<hash>-<run_id>,
another job may be creating this cache.
```

It's a warning, so the job stays green.

Which job gets to its post step first? The fastest one, which is also the job that compiles the least. In our workflow that's a job like OpenAPI codegen drift or deploy config validation, both done in under a minute. The cache that won was 8 KB. That became the newest entry, so `restore-keys` handed 8 KB to Go Build, Go Lint, and everything else on every run. They all logged "Cache restored," then compiled the whole tree from scratch, then lost the race to save.

The fix (#170) is one more variable in the key:

```yaml
key: go-${{ runner.os }}-${{ github.job }}-${{ hashFiles('**/go.sum') }}-${{ github.run_id }}
restore-keys: |
  go-${{ runner.os }}-${{ github.job }}-${{ hashFiles('**/go.sum') }}-
  go-${{ runner.os }}-${{ github.job }}-
```

Each job now has its own cache, and that makes sense for more than the race. Build & Test compiles the whole module and its test binaries, Lint loads packages for analysis, and codegen builds one tool. Sharing one cache between them was never going to work and something I had missed previously. However, that fix that made things slower before would now come into play. 

Results:

```chart
type: bars
title: Go jobs, median minutes
unit: m
series: Before #166, After #166, After #170
Go Build & Test: 4.6, 5.2, 1.8
Go Lint: 5.0, 6.0, 3.0
OpenAPI Codegen: 0.8, 1.0, 0.5
```

Go Lint still runs with `skip-cache: true`, so its 3.0 minutes is a warm Go build cache plus a cold lint analysis. That's the expected trade-off from #124.

**When to use a per-job, run-keyed cache:** any job whose build output changes often, which is most Go projects. **When not to:** if a job's cache is only downloaded dependencies, a key on the lockfile hash is fine because the cache only needs to change when the lockfile does. The run ID also means one cache entry per run. GitHub caps a repo at 10 GB and evicts the least recently used entries, so keep an eye on size. I thought ours sat well under the cap. It didn't, so more lessons were to be learned in the near future. Also Github, maybe like show when we have caches full in some sort of billing page with more identification around them? Regardless we should have checked our cache sizes because we would have caught that 8kb for a Go monorepo is probably wrong. 


## Three smaller fixes (#173)

With the main cache working, I went looking for other jobs paying the same costs.

**Lint only what changed on PRs.** `golangci-lint` has `--new-from-rev`, which reports only issues on lines that differ from a given revision. On pull requests we now pass `--new-from-rev=origin/main`. Push to main still lints the whole tree, so anything the scoped run misses gets caught on merge.

```yaml
- uses: actions/checkout@v7
  with:
    fetch-depth: 0   # --new-from-rev needs main's history

- uses: golangci/golangci-lint-action@v9
  with:
    skip-cache: true
    args: >-
      --timeout=10m --fix=false
      ${{ github.event_name == 'pull_request' && '--new-from-rev=origin/main' || '' }}
```

This flag still type-checks and analyzes every package, because most linters need the whole program. It only filters which issues get reported. The gain is smaller than "lint 5 files instead of 500" sounds. On the first run the golangci step went from about 130 seconds to 102.

**When not to use it:** if main isn't always clean. Scoped linting assumes the baseline passes. If main has existing failures, PRs will skip them and nobody will fix them.

**Cache the Grype vulnerability database.** Both Grype jobs downloaded the full vulnerability DB on every run. It's 462 MB and upstream publishes about once a day. Now it's cached with a date in the key:

```yaml
- name: Date for Grype DB cache key
  id: date
  run: echo "day=$(date -u +%Y-%m-%d)" >> "$GITHUB_OUTPUT"

- uses: actions/cache@v4
  with:
    path: ~/.cache/grype/db
    key: grype-db-${{ runner.os }}-${{ steps.date.outputs.day }}
    restore-keys: grype-db-${{ runner.os }}-
```

The first job each day downloads the DB and saves it. Every job after that restores it. Grype checks for a newer DB when it starts, so a day-old copy from `restore-keys` gets updated rather than scanned against as-is. For Grype a lockfile hash would be the wrong key here because the DB changes when new CVEs are published, not when our code does.

**Give the E2E workflow the same Go cache fix.** The E2E workflow lives in its own file and never got #166 or #170. It was still on setup-go's built-in cache with the stale-snapshot problem. It now has the job-scoped, run-keyed cache plus caches for Bun's install directory and the Grype DB, since the last stage of that pipeline shells out to Grype.

The first run of #173 was cold on purpose: new cache keys, so the second was the test because we had built our caches:

```chart
type: bars
title: PR #173, cold run vs warm run
unit: m
series: Before #173, #173 cold, #173 warm
Grype Go: 2.0, 2.3, 1.2
Grype Node: 2.3, 2.9, 1.8
E2E pipeline: 4.8, 5.8, 3.1
Go Lint: 3.0, 3.1, -
note: Go Lint did not rerun on the warm attempt.
```

The Grype DB cache is 462 MB, and restoring it is not free. It took 13 seconds in one job and 28 in another. That's still well ahead of downloading and unpacking the same file from upstream, but a cache is still a download, from a closer server. E2E restored a 324 MB Go cache and dropped by 1.7 minutes against its pre-#173 number. The lint job didn't rerun on the second attempt; its number from the first run is the one above, and since `skip-cache` stays on, there's no warm state for it to gain.


## The PR didn't get faster

Here's the per-era picture across all the PR-triggered jobs. Every job runs in parallel, so the PR finishes when the longest bar does. Switch between the eras:

```chart
type: race
title: PR-triggered jobs, median minutes
unit: m
series: Mid-Sept, Oct 1 to 6, After #170
Go Build & Test: 3.6, 4.6, 1.8
Go Lint: 3.8, 5.0, 3.0
Frontend Check: 1.0, 3.7, 5.1
E2E pipeline: 3.6, 4.6, 4.8
Grype (slower scan): 2.0, 2.2, 2.3
```

Going by job and it's a big win. However if are focused on the finish line, and it's flat. In mid-September the slowest job was Go Lint at 3.8 minutes. Today it's Frontend Check at 5.1. The Go fixes took two jobs off the critical path, and the frontend job, which had doubled in the same week, was already standing behind them so we saved 12 seconds, not quite good enough but that means we only had one more job to optimize. Some of the jobs we could have left as is, they do not affect our PR times and checks at all but none of that means the Go work was wasted. CI minutes cost money and runner slots, and Go Build failing at 1.8 minutes instead of 5.2 gets you a red X three minutes sooner. Creating earlier signals means faster turn around time on a fix for it. That matters, because a failing test is the most common reason to look at CI at all. The goals was still "my PRs are faster", so despite all the optimizations and minutes saved, we again were at a measurable 12 seconds faster PR times.  

A better prioritization would have been to pull the jobs from your last 20 PR runs, find which one finished last in each, and count. If one job is last 80% of the time, that's where we could have started and the GitHub API returns `started_at` and `completed_at` for every job, and a 20-line script against `/actions/runs/{id}/jobs` could have given us a better priority list. However, I think it was better to knock out one languages performance then switch onto the next one. If you are thinking of pure speed, knock down the slowest jobs first, if you are more like me and you want to start with where you are the most familiar, knocking things down is only beneficial in the long run. 


## The actual bottleneck

Frontend Check runs for 5.1 minutes and the vitest step is 3.9 minutes of it. Vitest's summary from the latest run:

```
Duration  233.89s (transform 1.94s, setup 2.31s, collect 72.96s, tests 125.55s, environment 21.60s, prepare 4.07s)
```

**Collect is 73 seconds.** That's vitest loading and transforming every test file and its imports before anything runs. Our components pull in Tamagui, and the first render in each file pays a one-time config compile. With 43 test files that cost is paid 43 times.

**Some tests sit around 2.6 seconds each.** EgressPolicyPanel, ServerSuspension, HostedServerSlots, LimitExceeded, and HelpMenu all cluster at that number. My first guess was real timers, since tests that all take the same odd length of time are usually waiting on a debounce or a `waitFor` running to its timeout. I was just wrong, sad I couldn't get a quick win, there were only two test files touch timers at all, and both already use `vi.useFakeTimers()`. So the 2.6 seconds was the first render in each file, and almost all of it was jsdom applying Tamagui's styles. More on that below.

**We already tried the obvious config change.** #171 switched vitest from its default `forks` pool to `threads`. On a laptop that cut the suite from 59.5 to 31.5 seconds because the V8 JIT stays warm across files. On CI the job didn't move: 5.1 minutes before, 5.1 after, demoralizing but hey just have to keep digging. A hosted runner on a private repo has 2 vCPUs and the laptop has a lot more, so there was almost no parallelism to gain. We also tried `isolate: false` and it broke 24 files on state leaking between them, the kind of leak a shared in-memory `localStorage` mock produces, so isolation stays. A lot of testing and not a lot of results, but again, see what sticks. 

From there the plan was:

1. **Make the test environment cheaper.** Every file pays for a DOM, and the DOM is where the first-render time goes. Low risk if the suite stays green.
2. **Shard the suite across jobs.** `vitest --shard=1/3` splits files across three parallel jobs. It's the one change that scales with suite size instead of fighting it. The cost is three runners instead of one, plus repeating the `pnpm install` and collect phase per shard.
3. **Move lint and type-check out of the test job.** They don't depend on each other, so running them in parallel takes them off the critical path entirely.

The target is a 3-minute PR, with 2 as the stretch.


## Fixing the tests: happy-dom

Vitest gives every test file a fake browser, and the default is jsdom. jsdom aims to be a faithful browser, and that includes a real CSS engine. Tamagui injects a lot of styles on first render, and jsdom parses and applies all of them. happy-dom is a lighter DOM built for tests that does far less of that work.

The switch was one line in `vitest.config.ts`, plus a stub for `navigator.clipboard` in two test files, because happy-dom's clipboard behaves differently. I also told Vite to pre-bundle Tamagui (`deps.optimizer.web`) so each file doesn't transform the same package again, and aliased `react-native` to `react-native-web` so that pre-bundle resolves.

The clearest single number: LimitExceeded's first test took 2,078 ms on jsdom and 48 ms on happy-dom.

Whole suite, 521 tests, median of 3 runs on the same M3 Pro. Every run passed all 521:

```chart
type: bars
title: vitest, 521 tests, seconds (M3 Pro)
unit: s
series: jsdom, happy-dom, happy-dom + Tamagui pre-bundle
Node 22: 30.3, 19.5, 16.5
Node 24: 27.4, 15.0, 13.7
Node 26: 27.4, 14.7, 12.4
```

From where CI was (Node 22, jsdom) to where the branch is (Node 24, happy-dom, pre-bundle), the suite went from 30.3 to 13.7 seconds, 55% less. About two thirds of that is happy-dom. Node 24 and the pre-bundle split the rest. Node 26 is a little faster again, but it isn't LTS yet, so CI moves to 24.

Still that is my laptop's number. On CI, with 2 vCPUs instead of 12, the same suite took 3.9 minutes. On the PR (#174) the vitest step took 106 seconds, down from the 233 to 244 range on the last three runs on `main`. That's 55% less, the same share as on the laptop, which makes sense: the work that went away is per-file setup, and every file pays it no matter how many cores there are. Collect is still the biggest phase at 60 seconds, so it's what to work on next.

While I was in there I removed some older dependencies that nothing imported anymore. I ensured the type check, lint, the full suite, and the production build all pass without them.


## Swapping the frontend toolchain

With the tests on a better footing, I wanted to know what the rest of the frontend job would cost with faster tools. Three candidates: TypeScript 7 (the Go port of the compiler), `bun check` (a new type checker in Bun canary, shipping in 1.4.3), and the Oxc tools, oxlint and oxfmt, for lint and format. Biome is in the chart as the other Rust option.

Every number below is the median of 5 runs on an M3 Pro with 12 cores, against the same 246 files in the web app. "Cold" means every `.tsbuildinfo` deleted first. The 2-thread bars are my best stand-in for a 2 vCPU hosted runner, not a measurement from one.

```chart
type: bars
title: Frontend toolchain, seconds (M3 Pro, 246 files)
unit: s
series: 12 cores, 2 threads
group: Type check
tsc 6.0.2 -b (cold): 12.04, -
tsc 6.0.2 -b (warm): 6.73, -
tsc 7.0.2 -b (cold): 1.86, 2.13
tsc 7.0.2 -b (warm): 1.16, -
bun check -b (canary): 1.10, 1.87
group: Lint
ESLint (current config): 4.44, -
ESLint (project off): 2.42, -
oxlint: 0.08, 0.09
oxlint --type-aware: 0.40, -
Biome (recommended rules): 0.14, -
group: Format check
Prettier 3.2.5: 2.26, -
oxfmt: 0.08, 0.10
Biome: 0.14, -
note: Median of 5 runs. tsc 6 is single-threaded, so it has no 2-thread number. Blank 2-thread cells were not measured.
```

Before trusting a faster type checker I wanted to know it reports the same errors. On our config all three report zero. That proves less than it sounds, so I turned on `--strict` for the web app, which produces 12 errors. tsc 6, tsc 7, and `bun check` reported the same 12, at the same file, line, column, and error code. Bun's issue tracker has a run of open false-positive reports against `bun check` from this week, and none of them hit this codebase. That's one codebase, so I'm treating it as "works for us," not "works."

**ESLint was paying for type information it never used.** The config sets `parserOptions.project`, which makes typescript-eslint build a full TypeScript program, but it only enables `tseslint.configs.recommended`, which has no type-aware rules. Turning `project` off cut lint from 4.4 to 2.4 seconds and changed nothing it reports. Finally got my quick win with a one-line fix that needs no new tool or testing!

**oxlint is about 50 times faster than ESLint here.** `@oxlint/migrate` converted 83 rules from the existing config. The one that matters and didn't come across is `import/order`: Oxc moved import sorting into the formatter. On a probe file with known problems, oxlint caught everything ESLint did except an unused `React` import. So my one line fix was now out the door again, we can see that quick wins and CI/CD are just not meshing well. 

**The format check wasn't enforcing anything.** Prettier says 199 of 246 files aren't formatted. Prettier only ran from a pre-commit hook, and that hook listed its file types under `types`, which pre-commit ANDs together. No file is JSON and CSS at the same time, so the hook never matched a single file. Nobody noticed because a hook that matches nothing passes.

Here's how much of the CI job these steps account for. The `main` column is the last three successful runs before the change. The PR column is the first run of #174, with happy-dom, Node 24, tsc 7, oxlint, and oxfmt all in:

```chart
type: bars
title: Frontend Check on CI, seconds
unit: s
series: main, PR #174
Install: 14-16, 23
Type check: 21-23, 6
Tests: 233-244, 106
Lint: 13-14, 1
Format check: -, 0.1
**Frontend Check job**: 323, 162
note: main is tsc 6, jsdom and ESLint, with no format check; ranges span the last three runs. PR #174 is tsc 7, happy-dom, oxlint and oxfmt.
```

The job went from 5.4 minutes to 2.7, half the time. Install is slower because the lockfile changed and the pnpm cache missed. It should be back to about 15 seconds once `main` saves a new cache.

On a 2 vCPU runner tsc 7 ran in 6 seconds against tsc 6's 22. oxlint reported "Finished in 106ms on 246 files using 2 threads." Type check and lint went from 35 seconds to 7, so the toolchain saved about half a minute. The biggest win was the tests saving more than two minutes. The toolchain swap is worth it for the local loop, where `tsc -b` dropping from 12 seconds to under 2 changes how often I run it. The PR got faster because of the tests.

So did the PR, this time. All of #174's checks started at 17:15:00, and the last one finished at 17:19:13: 4.2 minutes, down from about 5.1. Frontend Check finished fourth, behind Go Lint at 3.6 minutes and e2e at 4.2. That's the lesson from the first half of this post working the other way around. Cutting the frontend job in half bought 0.9 minutes, because that's how far ahead of e2e it was. e2e was now our critical path now, and its cache fix is sitting in #173.

Where the faster tools do pay off:

- **The local loop.** Type check, lint, and format check together drop from about 19 seconds cold (12.0 + 4.4 + 2.3) to about 2 (1.9 + 0.1 + 0.1). That's the difference between running all three on every save and running them before a push.
- **Pre-commit.** The hook runs ESLint over the whole web app on every commit. At 0.1 seconds instead of 4.4 it stops being the hook people skip, and the format check can actually be enforced.
- **The CI job shape.** Plan item 3 was to move lint and type check into their own job so they run in parallel with tests. At 2 to 3 seconds combined there's no point paying for a second runner and a second `pnpm install`. They can stay in the test job and add almost nothing.

### What I switched to

oxlint for lint, oxfmt for format, and tsc 7 for type checking. `bun check` was the fastest checker, but it's a canary build with open false-positive bugs, and it's under a second ahead of tsc 7. It gets another look once 1.4.3 is stable.

Dropping ESLint is what made tsc 7 easy. TypeScript 7 doesn't ship the compiler's JavaScript API until 7.1, and typescript-eslint is built on that API, so keeping ESLint would have meant installing TypeScript 6 next to 7. With ESLint gone, one package still needs the API: `packages/api` generates its clients with openapi-typescript, which builds its output with the TypeScript compiler's factory functions. That one package stays on TypeScript 6. Everything else is on 7.

The switch is six commits:

1. Replace ESLint with oxlint, using the config `@oxlint/migrate` produced. CI goes from a stale 421-warning ratchet to `--deny-warnings`.
2. Replace Prettier with oxfmt. `oxfmt --migrate=prettier` carried over every option except the import-sorting plugin, and `sortImports` with one custom group for `react` reproduces the old order: React, third party, `@/`, relative.
3. Reformat `apps/web/src`. 197 files, mechanical, nothing else in the commit.
4. Add that commit to `.git-blame-ignore-revs`, so `git blame` and GitHub's blame view skip it.
5. Move to TypeScript 7, except `packages/api`.
6. Fix two lint suppressions the reformat broke.

The last one is worth knowing about if you do this. Two components had a one-line effect with a suppression above it:

```tsx
// eslint-disable-next-line react-hooks/exhaustive-deps
useEffect(() => { refreshAll(); }, []);
```

The formatter split that into three lines. The warning is reported on the dependency array, which was now two lines below the comment, so the suppression stopped applying. Tests passed and type checks passed. `--deny-warnings` is the only thing that caught it, which is a decent argument for making lint warnings fail the build.

To check the new gates actually fail, I added a file with an `any` and an unused variable, and another with unsorted imports. oxlint and oxfmt each exited non-zero, and both went back to zero once the files were deleted.


## e2e: one stage removed, one cache borrowed

With #173 merged, e2e was the slowest check now. We could reuse some of what we learned and apply two changes got it to about 2.

### Dropping the scan stage

The last stage of the e2e pipeline ran Grype against the Docker image the test had just built, and failed on any high-severity finding. Grype is a vulnerability scanner that checks an image's packages against a CVE database. In the warm #173 run, that stage cost about 41 seconds: 31 to restore the 462 MB database cache, 4 to install Grype, and 6 for the scan itself.

I asked what the stage was actually testing, and the answer was "nothing in this PR." It called the Grype CLI directly from bash. No product code ran. The image it scanned is a generated MCP server on `oven/bun:1-alpine`, so the result depended on whether Alpine had an unpatched CVE that day. Our exceptions file had two entries for exactly that: OpenSSL and zlib, no fix available upstream, and `main` failing the same way as every PR.

The feature, which is Statio scanning images customers push and refusing to deploy one with critical findings, is tested elsewhere. Go unit tests feed the scan pipeline fake Grype reports and check that a critical blocks the deploy. Those run in seconds, need no database, and give the same answer every day.

So `scan` became an opt-in stage, like the two registry stages already were. The first run after that took **2m4s**.

### A cache no PR could ever restore

Then #174's e2e took **3m59s**. The log said:

```
Cache not found for input keys: go-Linux-e2e-a6a22cc...
```

The cause is how GitHub scopes caches. A workflow run can restore caches saved by its own branch or by the base branch, `main`. It can't see caches from other PRs. `ci.yml` runs on pushes to main, so its jobs always have a main-scoped cache for a new PR to start from. The e2e workflow only ran on `pull_request`. Each PR saved a cache that only that PR could use, so re-runs on the same branch were warm, but **the first run on every new PR compiled Go from scratch**. The deploy stage alone went from 46 seconds to 1m43s.

My first fix was to run e2e on pushes to main as well, so main would save a cache. It works, and it's the wrong trade. It adds a full 4-minute e2e run to every merge just to produce a cache, and that cache holds packages another job had already compiled.

`Go Build & Test` runs `go build ./...` and `go test ./...` on every push to main. The repo is a single Go module, so that job's cache already contains every package e2e compiles. It's 558 MB and keyed on the same `go.sum` hash. So e2e now restores it without saving anything:

```yaml
- name: Go cache (restore from ci.yml go job)
  uses: actions/cache/restore@v4
  with:
    path: |
      ~/.cache/go-build
      ~/go/pkg/mod
    key: go-${{ runner.os }}-go-${{ hashFiles('**/go.sum') }}-
    restore-keys: |
      go-${{ runner.os }}-go-${{ hashFiles('**/go.sum') }}-
      go-${{ runner.os }}-go-
```

That looks like it breaks the rule from bug two: give every job its own cache. It doesn't, and the reason is the rule exists because jobs that share a key also race to *save* it, and the fastest one overwrites everyone else's with 8 KB. `actions/cache/restore` is the restore half of the action with no post step. It can never write, so it can't race and it can't pollute anything. It also skips the 14 seconds e2e used to spend uploading its own cache at the end.

```chart
type: bars
title: e2e, cold vs borrowed cache
unit: s
series: #174 cold, #176 borrowed cache
gen stage: 24s, 1s
deploy stage: 1m43s, 43s
Pipeline step: 3m9s, 1m23s
**e2e total**: 3m59s, 2m9s
```

That was the first run of #176, on a branch that had never run e2e. Restoring 558 MB takes about 19 seconds. That's the biggest fixed cost left in the job, and it buys back about 1m40s of compiling (gen + deploy vs restore cache).

**When to borrow a cache read-only:** the other job builds a superset of what you build, with the same flags, and it runs on your default branch. **When not to:** if your job compiles things the other job doesn't, like a different module, different build tags, or `-race`. A restore-only job never saves, so anything it compiles that isn't in the borrowed cache gets compiled again on every run.

### A cache key that hashed nothing

While I was in that file, the Bun cache logged its key as `bun-Linux-`. That key was built from `hashFiles('**/bun.lock')`, and `hashFiles` returns an empty string when nothing matches. The Bun project in e2e is the MCP server the test generates at runtime, so there's no lockfile in the repo for it to hash. The key never changed, and the `bun install` it was meant to speed up takes 3 seconds. I removed it.

## Two days later: 2.5 minutes

After #174 and #176, the first PR run took 4.2 minutes. Two more PRs went in that evening, one for lint and one for the cache budget. Here's every PR run since, 125 of them, against each job's worst median from the week before:

```chart
type: stats
title: Medians, worst this week vs now
unit: m
series: worst, now
PR, open to all green: 5.3, 2.5
Go Lint: 6.0, 1.4
Go Build & Test: 5.2, 1.8
Frontend Check: 5.4, 2.4
e2e: 4.8, 2.3
note: "Worst" is each job's highest median during the week: after #166 for the Go jobs, the last main runs before #174 for Frontend Check, before #173 for e2e. "Now" is successful PR runs from October 7 evening to October 9.
```

The PR number is the time from the first job starting to the last one finishing, first attempts only, on the 44 commits that ran the full set of jobs. Median 2.48 minutes, p90 3.2, worst 3.6. 24 of the 44 finished in under 2.5. The target earlier in this post was 3 minutes with 2 as the stretch. The median is past the target and the p90 is close to it.

### Profiling the linter (#177)

Go Lint was still the job with `skip-cache: true` from #124, so every run did a cold analysis of the whole tree. Before deciding what to cut, I ran `golangci-lint` with `--cpu-profile-path` on a cold full-tree run and read the profile. golangci-lint runs every analyzer inside one pass, so `pprof -top` lumps them together. Summing the profile by analyzer gives this:

```chart
type: bars
title: golangci-lint CPU by analyzer, cold full tree, seconds
unit: s
series: flat CPU
SSA/IR build (shared): 56.1
wrapcheck: 16.6
gosec: 13.8
whitespace: 9.4
revive: 4.6
staticcheck: 3.5
unused: 2.8
note: The SSA/IR build is shared by staticcheck, unused, gosec and nilness. errcheck, goconst, dupl, musttag and exhaustruct were all near zero.
```

What I found was that the biggest cost isn't a linter. It's building the SSA form of the program, and four linters share it, so disabling any one of them saves nothing, it was determined that those checks were needed so we tried to take a quick look to see if we could speed this one up a bit as well. 

**The analysis cache is back on.** The reason it was off was `nolintlint` flipping between "this directive is required" and "this directive is unused." The config already had a comment about the same failure from a different cause: golangci-lint caps how many identical issues it reports, and when the issue a `//nolint` covers gets dropped by that cap, the directive looks unused. `max-same-issues: 0` was already in the config to turn that cap off, which is very likely why the flapping stopped reproducing. I checked it rather than assume it: a cold run, then five warm runs and three with a file touched, and all eight reported zero `nolintlint` issues. That's evidence, not proof, so the workflow comment says what to do if it comes back: put `skip-cache: true` back, one line. Warm, lint takes 3.1 seconds. One file touched, 7.4. Cold, 36 to 56.

**Generated code is no longer a lint target.** `linters.exclusions.paths` looked like it was skipping our generated code. It wasn't. It filters *results*, so every analyzer still walked 1.82 million generated lines, 48% of the tree, and the findings were thrown away at the end. The only way to skip that work is to leave those packages off the command line. A cold run at 2 CPUs, the runner's shape, went from 72.5 to 37.5 seconds with identical output. The generated packages still get type-checked, because the code we lint imports them. What goes away is running analyzers over them.

**`whitespace` is off.** 9.4 seconds to check for stray blank lines inside blocks, which is the formatter's job.

Removing that exclusion turned up something else. `db/ent` was an unanchored path pattern, so it matched more than the generated ent client. It also silenced the 27 hand-written schema files in `db/ent/schema/`, plus `db/ent_client.go` and `db/ent_test.go`, which just happen to contain `db/ent` in their paths. Both generated directories were already covered by `generated: strict`, which skips files with a `DO NOT EDIT` header, so the pattern added nothing for generated code and hid real code. Un-excluding those files surfaced two findings, both fixed in the PR.

Go Lint went from 3.2 to 3.6 minutes on the last runs before #177 to a median of 1.4 after it. The golangci-lint step itself is now 27 to 33 seconds.

### A cache that saves on every PR (#181)

The run-keyed cache from #170 has a cost I waved off earlier. Every run saves a new entry, and in this repo each Go cache is about 0.56 GiB. PRs included, that's a lot of half-gigabyte writes. The repo's cache usage was at 12.5 GiB against GitHub's 10 GiB limit, so eviction was running all the time, and it removes the least recently used entries first. Whatever runs least often loses its cache. For us that was the release pipeline, which is its own story and the next post. For PRs, the fact that matters is that every run was paying to upload half a gigabyte that only that branch could ever use.

The fix is to restore on every run and save only on pushes to main:

```yaml
- name: Go cache (restore)
  uses: actions/cache/restore@v4
  with:
    path: |
      ~/.cache/go-build
      ~/go/pkg/mod
    key: go-${{ runner.os }}-${{ github.job }}-${{ hashFiles('**/go.sum') }}-
    restore-keys: |
      go-${{ runner.os }}-${{ github.job }}-${{ hashFiles('**/go.sum') }}-
      go-${{ runner.os }}-${{ github.job }}-

# ... the job's Go work ...

- name: Go cache (save)
  if: github.event_name == 'push'
  uses: actions/cache/save@v4
  with:
    path: |
      ~/.cache/go-build
      ~/go/pkg/mod
    key: go-${{ runner.os }}-${{ github.job }}-${{ hashFiles('**/go.sum') }}-${{ github.run_id }}
```

PRs read and never write. Main writes on every merge, which is often enough to stay warm. That's 5 to 6 times fewer Go cache writes. The trade is that a PR starts from the last merge instead of its own last run, so a re-run on the same branch compiles a little more than it used to. With the branch scoping rules from the e2e section, the first run of a new PR was starting from main's cache anyway.

### Where the time goes now

```chart
type: race
title: PR-triggered jobs, median minutes
unit: m
series: Oct 1 to 6, After #170, Now
default: Now
Go Build & Test: 4.6, 1.8, 1.8
Go Lint: 5.0, 3.0, 1.4
Frontend Check: 3.7, 5.1, 2.4
E2E pipeline: 4.6, 4.8, 2.3
Grype (slower scan): 2.2, 2.3, 1.8
```

Every job now finishes within about a minute of the others. Frontend Check is still the last to finish most often, 25 of 44 runs, then Grype Node 7, e2e 4, Grype Go 3, Go Lint 3, Go Build 2. That spread is what a balanced pipeline looks like. It also means no single fix moves the PR much from here. Cutting a minute off Frontend Check would hand the finish line to e2e or Grype Node, which are within seconds of it.

The fixed costs are now bigger than the work. In a recent Go Lint run, the golangci-lint step took 27 to 33 seconds and restoring the Go cache took 34 to 39. Checkout and setup-go add another 16 to 24. Lint is cheaper to run than its cache is to download, and the next round of work is shrinking what each job has to restore.


## Self-hosted runners

I have a Proxmox box at home, and it's tempting. A persistent runner keeps its disk between jobs, so the Go build cache, the lint cache, the Grype DB, and `node_modules` are always warm, and nothing competes for a 10 GB cache budget. I used to think it was the only way to get the lint cache back. #177 did that on hosted runners.

I'm not doing it yet, for two reasons. One box means jobs queue behind each other, and this pipeline is fast because it's parallel. And a self-hosted runner executes whatever code is in the PR on hardware I own, on my network. For a security product that needs a real isolation story (ephemeral VMs per job, no access to the LAN) before it gets anywhere near a pull request.


## Notes for next time

- **Find the critical path before optimizing anything.** Make a table of which job finished last. That's a 20-minute script and it would have changed what I worked on first.
- **Check cache sizes, not cache hits.** "Cache restored" with an 8 KB payload is a miss.
- **A cache key that can hit never updates.** For build output, put something unique in the key and let `restore-keys` find the newest entry.
- **Jobs that share a cache key race to save it, and the fastest one wins.** Put `github.job` in the key.
- **A job that only runs on PRs never has a cache on main.** Every new PR starts cold. Borrow a main job's cache read-only rather than adding a main run to seed one.
- **Print your cache keys.** `hashFiles` on a glob that matches nothing returns an empty string, and the key quietly becomes a constant.
- **Ask what a CI stage tests.** A scan of today's CVE database against a base image tests the database, not the PR.
- **Profile the linter before tuning it.** The biggest cost was a shared SSA build that no single linter owns. Without the profile I'd have disabled staticcheck and saved nothing.
- **Exclusions filter results, they don't skip work.** To stop analyzing generated code, leave it off the lint targets.
- **Anchor your path patterns.** `db/ent` also matched 27 hand-written files and two that only had `db/ent` in their names.
- **A cache that saves on every PR spends the whole repo's budget.** Eviction hits the caches that run least often. Save on main, restore everywhere.
- **Watch the job you aren't working on.** Frontend Check doubled in a week from normal feature work. A weekly look at job medians would have caught it before it was the bottleneck.

And as a final note, if you go down optimizing your CI pipeline, which I think Github would GREATLY appreciate, don't expect quick wins like I was hoping for!


## Resources

- [actions/cache: how cache keys and restore-keys match](https://github.com/actions/cache#creating-a-cache-key)
- [actions/cache/restore: restore-only action](https://github.com/actions/cache/tree/main/restore)
- [actions/cache/save: save-only action](https://github.com/actions/cache/tree/main/save)
- [GitHub docs: restrictions for accessing a cache (branch scope)](https://docs.github.com/en/actions/using-workflows/caching-dependencies-to-speed-up-workflows#restrictions-for-accessing-a-cache)
- [GitHub docs: caching dependencies, usage limits and eviction](https://docs.github.com/en/actions/using-workflows/caching-dependencies-to-speed-up-workflows)
- [golangci-lint: `--new-from-rev` and related flags](https://golangci-lint.run/usage/configuration/)
- [Grype: database update behavior](https://github.com/anchore/grype#grypes-database)
- [Vitest: test sharding](https://vitest.dev/guide/improving-performance)
- [Vitest: fake timers](https://vitest.dev/guide/mocking#timers)