BootyBayBroker runs more than 66,000 tests, and most changes touch a small
part of the code. I built a selector that picks the test files a change
needs, measured it on 110 failed CI runs before turning it on, and have let
it choose the tests in CI since September 21, 2026.
My role
Design, implementation, measurement, rollout
Stack
TypeScript, GitHub Actions, a hosted judgment model (TypeSafe Jev)
Scope
66,000+ tests in about 2,800 test files
Status
Choosing BootyBayBroker’s tests in CI since September 21, 2026
13.7%
median share of the suite’s test files run per change in the replay
94 of 110
test files that failed because of their own change, and were selected
under 1¢
the model’s cost per run; the whole 110-run replay cost about $2.68
5–9 min
selected runs on launch day, where the full matrix took 20 to 29
What I built
Two layers decide what runs. A dependency graph covers what code can
decide exactly, and a model answers one yes-or-no question for each test
the graph reaches only indirectly. Anything unusual runs the full suite,
and after each merge an audit job runs every test the selector skipped,
so the whole suite still runs once per change.
The result in CI: a typical pull request runs one Jest shard and one
Vitest shard instead of three of each, the selection step takes one and a
half to two and a half minutes, and the model’s questions cost well under
a cent per run.
Engineering decisions
A graph that over-approximates on purpose
Imports alone miss too much: 37% of the Jest files read repository paths
from disk. So the graph follows imports, every repository file a test
reads (migration SQL, workflow YAML, fixtures), shell scripts and hooks
read as text, and build outputs traced back to their sources. A test
within one hop of the change always runs, and no model is asked about it.
In a static replay of 318 failed runs, 79.7% of the failing files sat
within one hop.
One question per far test
Tests the graph reaches at two or more hops each get one question, forty
to a request: does this test check the behavior, file, or configuration
that this change modifies? The test side sends its titles, imports, the
files it reads, and its reach chain; the change side sends a digest of
paths, changed symbols, and trimmed hunks. The model answers with a
probability. Tests at 0.15 or higher run, and an unanswered question
counts as yes. A third tier for tests the graph never reaches shipped
switched off: it asked three quarters of all questions and caught nothing.
Fail open
The full suite runs when there is no base commit, when the model is down,
slow, or out of credit, when the runner, the dependencies, or the
selector’s own inputs change, when a change touches more than 60 code
files, when a selection would exceed 60% of a suite, or when the plan
fails to build. A kill switch forces full runs, and a shadow mode
publishes the plan while still running everything.
Audit every skipped test after merge
After each code push to main, an audit workflow runs exactly the tests the
plan skipped. Selected plus audited equals the whole suite, once per
merged change. The audit gates nothing: a miss opens an issue. In its
first three days it ran fifty times, forty-nine of them clean and one
stopped by container setup, and as of September 24 no miss had been
opened.
Triage failures with the same model
Each failing file is sorted into a timeout, flaky network, a deadlock, a
genuine assertion failure, or other. A rerun is paid for only when every
failing file is transient at 0.95 or higher, none shows an assertion
diff, and there are at most fifteen of them. The rerun’s own result
decides the job.
Proven trees
A tree that has already passed a full run is proven, and its suites are
skipped. A later pull-request head is planned as a delta from the previous
head plus the files that head left red. The deploy verification after a
merge now fires about 25 minutes sooner.
One change through the selector
From a diff to a plan, and the audit that closes the loop.
ChangePaths, symbols, trimmed hunks
GraphWithin one hop always runs
QuestionsOne per far test, p ≥ 0.15 runs
PlanOne shard each, or the full suite
AuditEvery skipped test after merge
Measured before launch
A replay of real failures, with every miss checked by hand.
110
failed CI runs replayed, those with diffs of 120 files or fewer among 400 mined between August 4 and September 21, 2026
16
misses, each checked by hand: a timeout, a database deadlock, a flaky performance budget, or a test already failing
90 of 110
caught by the graph alone at a median 9.2% of files; the model’s questions add four more for 4.5 more points of share
63.8M
input tokens over 4,981 requests for the whole replay, about $2.68
The tradeoff
Failing open means an unusual change pays for a full run, and the
post-merge audit means the suite’s total compute per merged change does
not fall; what falls is the time a push waits for its verdict. Launch-day
pushes finished in 5 to 9 minutes against 20 to 29 for the full matrix.