Engineering.

I build developer tools for AI coding agents: predictive test selection, 20 Claude Code plugins and MCP servers, retrieval memory, and evaluation harnesses. I measure each one on products I run.

AI developer tools

Three questions shape them: what should run, what should the agent know, and how do we check its work?

Predictive test selection

BootyBayBroker has more than 66,000 tests, and most changes touch a small part of the code. My selector picks the test files a change needs. A dependency graph covers what code can decide exactly: imports, plus every repository file a test reads from disk, such as migration SQL, workflow YAML, and fixtures. A test within one hop of the change always runs.

For tests the graph reaches only indirectly, a model answers one yes-or-no question per test, with a probability: does this test check what this change modifies? Tests at 0.15 or higher run, and an unanswered question counts as yes. Anything unusual, from a missing base commit to a model outage, runs the full suite. After each merge, an audit job runs every test the selector skipped.

It has chosen the tests in CI since September 21, 2026. On its first day, pushes ran in 5 to 9 minutes, against 20 to 29 for the full suite. As of September 24, no post-merge audit has found a missed failure.

Measured before launch

Replay
110 failed CI runs from August and September 2026
Caught
94 of the 110 test files that failed because of their change
Ran
A median 13.7% of the suite’s test files
Graph alone
90 of 110, at a median 9.2%
Misses
All 16 checked by hand: timeouts, a database deadlock, a flaky performance budget, or tests that were already failing

Claude Code plugins and MCP servers

I’ve built 20 plugins and MCP servers for Claude Code. Eleven are in my current catalog, eleven have Codex CLI ports, and nine of the current eleven are public on GitHub. They cover code and UI review, CI/CD, Railway operations, SEO, project onboarding, and planning.

My products run on them. Every Limerino pull request that changes code gets an ios-code-review pass, and its interface work goes through apple-ui-craft. I also retire what stops earning its place: five plugins left the catalog in September 2026 after 78 days without use.

anti-slop (external site)
Deterministic checks for code, UI, and prose
ui-craft (external site)
Web UI design and review agents
apple-ui-craft (external site)
SwiftUI design and review agents
ios-code-review (external site)
iOS review from App Review’s and a platform engineer’s view
cicd-expert (external site)
CI/CD architecture, review, and debugging
railway-operator (external site)
Railway deployment and operations
seo-optimization (external site)
SEO audits and fixes for any website
project-onboard (external site)
New-project setup and onboarding
scrum-master (external site)
Story-per-file planning for agent work
wow-addon-dev (external site)
An agent skill for World of Warcraft addon development

Persistent memory

I run retrieval memory for my coding agents on GoodMem, a self-hosted retrieval server: project-scoped vector storage across 21 projects, Voyage embeddings, and reranking. Agents query it when they hit an unfamiliar error, tool, or API, and treat what comes back as a hint to verify, with its source attached.

Keeping tens of thousands of memories useful is its own job. A report-only service asks a model typed questions, yes-or-no and graded, about each memory and its nearest neighbors, and flags the ones that look stale, superseded, or duplicated. It never edits or deletes a memory.

Read the July 2026 design notes
01Capture

Source, decision, verification path

02Retrieve

Project scope, embeddings, reranking

03Check

Revisit source truth; flag stale claims

04Use

Relevant context with provenance

Freshness is part of retrieval quality.

Repository-grounded evaluation

My local evaluation harness pins cases to a committed repository revision and requires file-and-line evidence for claims. Model calls use bounded, read-only archives, isolated from external memory and the source checkout.

The goal is concrete: distinguish a plausible answer from one supported by the repository. Cases span languages and projects, including checks for context leaking across project boundaries.

A checkable engineering claim

Source
A specific repository revision
Evidence
The file and line supporting the claim
Boundary
Only the supplied read-only archive
Result
A claim that can be independently checked

Private tooling in my own workflow; the method is shown here.

Self-hosted CI/CD

63 GitHub Actions runners on three machines serve 17 repositories: Linux for most jobs, Windows for Windows builds, and a Mac for iOS builds and TestFlight uploads.

Production engineering

One hard problem from each product, with the case study behind it.

Selected public repositories

GitHub (external site)

Tools from my engineering workflow, with source and tests you can inspect.

anti-slop (external site)

JavaScript, Node.js CLI, MIT license

Deterministic analysis

A zero-dependency scanner with 85 named rules for security, accessibility, and quality problems in code, UI, and prose. It runs independently of the coding agent.

Rules are checked against a labeled corpus with clean controls and documented coverage limits. Precision and recall gates make missed issues and noisy rules visible.

ui-craft (external site)

JavaScript evaluation tooling, MIT license

Review quality

A web UI design and review team with a regression harness for evaluating the reviewer itself. Findings are matched to labeled interface fixtures. I used it to review this site.

The scoring gate checks precision, recall, and clean controls. A finding must provide evidence beyond repeating a rule ID to count as a detection.

smart-compact (external site)

Python and shell, MIT license

Retired September 2026

Hooks that checkpointed a coding agent’s working state to disk before context compaction and restored it afterward, with a transcript-derived fallback and a 16KB restoration budget.

I ran it from July to September 2026, then retired it: on the 1M-token context window its ceiling estimate drifted, and sessions compacted at about half their context. The repository stays public with its 215 regression checks.

Interactive example: event pipeline

Interactive example: a simulated event pipeline Validation, deduplication, retries, and recovery, with sample data.

24 sample events
Simulated time, runs in your browser

Temporary failures recover within a bounded attempt budget.

Tick 0
IngestAdmit sample events0 received
ValidateCheck the schema0 waiting
QueueBackpressure0 / 4 slots
WorkerBounded concurrency0 / 2 workers
StoreCommit once per key0 committed
Rejected at validationRetry after 2, then 4 ticksDead letter after 3 attempts

This run

Committed
0
Duplicates ignored
0
Retry attempts scheduled
0
Invalid payloads
0
Dead letter
0

Ready. Run a scenario, or advance one tick at a time.

Follow an event

  • Committed
  • Waiting to retry
  • Dead letter
  • Invalid
  • Processing
  • Duplicate
Event 05Not received yet
  1. Run or step the simulation to see this event’s trace.

What this demonstrates

Capacity is a design decision. Two workers and a four-slot queue put a limit on in-flight work. When the queue fills, admission waits.

Delivery can repeat. Side effects should not. Idempotency keys prevent duplicate commits. Retries have increasing delays and a finite attempt budget.

A failure needs somewhere to go. Invalid inputs stop before work begins. Exhausted retries stay available in a dead-letter queue.

An educational model with deterministic sample data. It demonstrates patterns used across my work; it is not production telemetry or a copy of any one service.