Three questions shape them: what should run, what should the agent
know, and how do we check its work?
Predictive test selection
BootyBayBroker has more than 66,000 tests, and most changes touch a
small part of the code. My selector picks the test files a change
needs. A dependency graph covers what code can decide exactly: imports,
plus every repository file a test reads from disk, such as migration
SQL, workflow YAML, and fixtures. A test within one hop of the change
always runs.
For tests the graph reaches only indirectly, a model answers one
yes-or-no question per test, with a probability: does this test check
what this change modifies? Tests at 0.15 or higher run, and an
unanswered question counts as yes. Anything unusual, from a missing base
commit to a model outage, runs the full suite. After each merge, an
audit job runs every test the selector skipped.
It has chosen the tests in CI since September 21, 2026. On its first
day, pushes ran in 5 to 9 minutes, against 20 to 29 for the full suite.
As of September 24, no post-merge audit has found a missed failure.
Measured before launch
- Replay
- 110 failed CI runs from August and September 2026
- Caught
- 94 of the 110 test files that failed because of their change
- Ran
- A median 13.7% of the suite’s test files
- Graph alone
- 90 of 110, at a median 9.2%
- Misses
-
All 16 checked by hand: timeouts, a database deadlock, a flaky
performance budget, or tests that were already failing
Claude Code plugins and MCP servers
I’ve built 20 plugins and MCP servers for Claude Code. Eleven are in my
current catalog, eleven have Codex CLI ports, and nine of the current
eleven are public on GitHub. They cover code and UI review, CI/CD,
Railway operations, SEO, project onboarding, and planning.
My products run on them. Every Limerino pull request that changes code
gets an ios-code-review pass, and its interface work goes through
apple-ui-craft. I also retire what stops earning its place: five
plugins left the catalog in September 2026 after 78 days without use.
Persistent memory
I run retrieval memory for my coding agents on GoodMem, a self-hosted
retrieval server: project-scoped vector storage across 21 projects,
Voyage embeddings, and reranking. Agents query it when they hit an unfamiliar error, tool,
or API, and treat what comes back as a hint to verify, with its source
attached.
Keeping tens of thousands of memories useful is its own job. A
report-only service asks a model typed questions, yes-or-no and graded,
about each memory and its nearest neighbors, and flags the ones that
look stale, superseded, or duplicated. It never edits or deletes a
memory.
Read the July 2026 design notes
01Capture
Source, decision, verification path
02Retrieve
Project scope, embeddings, reranking
03Check
Revisit source truth; flag stale claims
04Use
Relevant context with provenance
Freshness is part of retrieval quality.
Repository-grounded evaluation
My local evaluation harness pins cases to a committed repository
revision and requires file-and-line evidence for claims. Model calls use
bounded, read-only archives, isolated from external memory and the
source checkout.
The goal is concrete: distinguish a plausible answer from one supported
by the repository. Cases span languages and projects, including checks
for context leaking across project boundaries.
A checkable engineering claim
- Source
- A specific repository revision
- Evidence
- The file and line supporting the claim
- Boundary
- Only the supplied read-only archive
- Result
- A claim that can be independently checked
Private tooling in my own workflow; the method is shown here.
Self-hosted CI/CD
63 GitHub Actions runners on three machines serve 17 repositories:
Linux for most jobs, Windows for Windows builds, and a Mac for iOS
builds and TestFlight uploads.