Everyone talks about models. Almost nobody talks about the plumbing that turns a model into a phone call that feels natural. Yet in voice AI the plumbing is most of the product: a brilliant model with 800 milliseconds of latency sounds worse than a modest one that answers instantly.
Latency is the product
A natural conversational gap is a few hundred milliseconds. Speech recognition, language model, speech synthesis and two network hops all have to fit inside it. We track a latency budget for every stage of the pipeline and treat a regression of fifty milliseconds as a bug.
Designing for failure
Calls cannot buffer. If a region degrades, traffic shifts automatically; if a model provider slows down, the assistant falls back to a faster path; if everything fails, the caller is offered a callback rather than silence. Graceful degradation is designed in from the start, not added after an outage.
Observability on every turn
Each exchange in a conversation is traced end to end — audio in, text, decision, audio out — with timings. When a customer asks why a call went wrong we can answer in minutes, and the same traces drive the regression tests that stop it happening again.
Boring choices, fast releases
Standard tooling, automated tests on real call recordings and small weekly deployments. It is the unglamorous combination that lets a small team move quickly while running production telephony for travel brands.
In voice, the infrastructure is what the customer hears. Every millisecond is audible.



