In voice AI, 300–500ms latency differences completely change conversation quality. Yet most teams still benchmark only tokens/sec instead of end-to-end conversational responsiveness. LLMs are evolving fast, and so is OpenAI. Every few months, the model layer changes in a meaningful way: better reasoning, better tool calling, better streaming
Most of us who built products before LLMs carry the same instincts into it. You spec it, you build it, you test it, you ship it. QA sits at the end. We carried that instinct into building with LLMs — and it broke. At Yuuki*, we built an AI coaching product
A startup CEO recently posted that his team generated a million lines of code with AI agents in a quarter. After twelve months shipping a production voice AI platform with AI coding agents, I've started paying attention to a different question: what happens after the code is written?
The constraint We served 8 enterprise clients with different compliance requirements, different business logic, and different quality standards. And, our Product, Design & Engineering (PDE) team was two people, Prashant and I. No QA team. No SRE. No engineering manager reviewing pull requests. Every line of code that reaches production