Predictive test selection.

BootyBayBroker runs more than 66,000 tests, and most changes touch a small part of the code. I built a selector that picks the test files a change needs, measured it on 110 failed CI runs before turning it on, and have let it choose the tests in CI since September 21, 2026.

My role
Design, implementation, measurement, rollout
Stack
TypeScript, GitHub Actions, a hosted judgment model (TypeSafe Jev)
Scope
66,000+ tests in about 2,800 test files
Status
Choosing BootyBayBroker’s tests in CI since September 21, 2026
13.7%
median share of the suite’s test files run per change in the replay
94 of 110
test files that failed because of their own change, and were selected
under 1¢
the model’s cost per run; the whole 110-run replay cost about $2.68
5–9 min
selected runs on launch day, where the full matrix took 20 to 29

What I built

Two layers decide what runs. A dependency graph covers what code can decide exactly, and a model answers one yes-or-no question for each test the graph reaches only indirectly. Anything unusual runs the full suite, and after each merge an audit job runs every test the selector skipped, so the whole suite still runs once per change.

The result in CI: a typical pull request runs one Jest shard and one Vitest shard instead of three of each, the selection step takes one and a half to two and a half minutes, and the model’s questions cost well under a cent per run.

Engineering decisions

A graph that over-approximates on purpose

Imports alone miss too much: 37% of the Jest files read repository paths from disk. So the graph follows imports, every repository file a test reads (migration SQL, workflow YAML, fixtures), shell scripts and hooks read as text, and build outputs traced back to their sources. A test within one hop of the change always runs, and no model is asked about it. In a static replay of 318 failed runs, 79.7% of the failing files sat within one hop.

One question per far test

Tests the graph reaches at two or more hops each get one question, forty to a request: does this test check the behavior, file, or configuration that this change modifies? The test side sends its titles, imports, the files it reads, and its reach chain; the change side sends a digest of paths, changed symbols, and trimmed hunks. The model answers with a probability. Tests at 0.15 or higher run, and an unanswered question counts as yes. A third tier for tests the graph never reaches shipped switched off: it asked three quarters of all questions and caught nothing.

Fail open

The full suite runs when there is no base commit, when the model is down, slow, or out of credit, when the runner, the dependencies, or the selector’s own inputs change, when a change touches more than 60 code files, when a selection would exceed 60% of a suite, or when the plan fails to build. A kill switch forces full runs, and a shadow mode publishes the plan while still running everything.

Audit every skipped test after merge

After each code push to main, an audit workflow runs exactly the tests the plan skipped. Selected plus audited equals the whole suite, once per merged change. The audit gates nothing: a miss opens an issue. In its first three days it ran fifty times, forty-nine of them clean and one stopped by container setup, and as of September 24 no miss had been opened.

Triage failures with the same model

Each failing file is sorted into a timeout, flaky network, a deadlock, a genuine assertion failure, or other. A rerun is paid for only when every failing file is transient at 0.95 or higher, none shows an assertion diff, and there are at most fifteen of them. The rerun’s own result decides the job.

Proven trees

A tree that has already passed a full run is proven, and its suites are skipped. A later pull-request head is planned as a delta from the previous head plus the files that head left red. The deploy verification after a merge now fires about 25 minutes sooner.

One change through the selector

From a diff to a plan, and the audit that closes the loop.

  1. ChangePaths, symbols, trimmed hunks
  2. GraphWithin one hop always runs
  3. QuestionsOne per far test, p ≥ 0.15 runs
  4. PlanOne shard each, or the full suite
  5. AuditEvery skipped test after merge

Measured before launch

A replay of real failures, with every miss checked by hand.

110
failed CI runs replayed, those with diffs of 120 files or fewer among 400 mined between August 4 and September 21, 2026
16
misses, each checked by hand: a timeout, a database deadlock, a flaky performance budget, or a test already failing
90 of 110
caught by the graph alone at a median 9.2% of files; the model’s questions add four more for 4.5 more points of share
63.8M
input tokens over 4,981 requests for the whole replay, about $2.68

The tradeoff

Failing open means an unusual change pays for a full run, and the post-merge audit means the suite’s total compute per merged change does not fall; what falls is the time a push waits for its verdict. Launch-day pushes finished in 5 to 9 minutes against 20 to 29 for the full matrix.

The product it protects