I build and operate the path from radio audio to a searchable web feed.
Audio is converted to 16 kHz PCM, filtered by an energy-based speech
detector, transcribed with OpenAI’s gpt-transcribe, and posted to
a FastAPI and PostgreSQL service. The work around transcription is what
keeps that pipeline useful over time.
Engineering decisions
Bounded capture
Captured segments wait in a bounded queue of 20. If transcription falls
behind, the queue drops a segment instead of growing without limit.
Each segment gets two transcription attempts with a timeout, and a
watchdog restarts the pipeline after 15 minutes without progress.
Monitoring useful output
A running container does not prove new transcripts are arriving.
Railway’s health check only asks whether the service is up, because a
quiet night on the radio shouldn’t restart it. A separate freshness
endpoint reports stale after 15 minutes without a transcript, and I get
an email after an hour.
Feed failover
The pipeline listens to one scanner feed at a time. If the primary feed
stays silent for two hours, longer than any normal quiet stretch,
capture switches to a backup feed, so an upstream outage does not
quietly end the transcript.
Speech recognition as a dependency
Voice activity detection gates audio before it reaches the recognizer,
and the transcription provider is a configuration setting, which let me
switch providers in August 2026 with the old one kept as a rollback.
Architecture
How radio audio becomes a searchable transcript.
Radio audio16 kHz PCM capture
Speech gateEnergy VAD
Bounded queue20 segments, 2 attempts each
TranscriptionOpenAI gpt-transcribe
Public feedFastAPI and PostgreSQL
The tradeoff
A bounded queue keeps memory predictable during an outage, but it can
drop audio. Freshness checks, alerts, and feed failover make any gap
visible and short instead of silent. The interactive example on the
Engineering page shows the same retry and backpressure ideas with sample
data.