Chat Completions vs Responses vs WebSocket: TTFT and TRT in Rapida
In voice AI, 300–500ms latency differences completely change conversation quality. Yet most teams still benchmark only tokens/sec instead of end-to-end conversational responsiveness. LLMs are evolving fast, and so is OpenAI. Every few months, the model layer changes in a meaningful way: better reasoning, better tool calling, better streaming
Evals aren't testing, they're teaching
Most of us who built products before LLMs carry the same instincts into it. You spec it, you build it, you test it, you ship it. QA sits at the end. We carried that instinct into building with LLMs — and it broke. At Yuuki*, we built an AI coaching product
Agentic Engineering: Technically Correct, Contextually Naive
A startup CEO recently posted that his team generated a million lines of code with AI agents in a quarter. After twelve months shipping a production voice AI platform with AI coding agents, I've started paying attention to a different question: what happens after the code is written?
The Quality gap nobody talks about
The constraint We served 8 enterprise clients with different compliance requirements, different business logic, and different quality standards. And, our Product, Design & Engineering (PDE) team was two people, Prashant and I. No QA team. No SRE. No engineering manager reviewing pull requests. Every line of code that reaches production
Rapida Voice System Performance Benchmarks
This post documents latency and scaling behavior of Rapida under sustained concurrent voice traffic. The intent is to show how the system behaves under realistic production conditions rather than demo workloads. Rapida is an open source voice orchestration system designed to manage real time audio streaming, speech recognition, language model
Transforming Customer Service with Voice AI: From IVR to Intelligent Agents
Why legacy IVRs fall short Last week I called my bank and spent 4 minutes pressing buttons before a robot told me to call back during business hours. We can do better. Classic IVRs were designed to deflect calls with keypad trees. Today, customers expect fast, human-like help and seamless
Voice Agent Telemetry: Where Every Millisecond Counts
When building real time voice systems, there's a moment we've all felt. That pause after you finish speaking, waiting for the AI to respond. It's a fraction of a second, but it breaks the flow. That delay is what separates fast enough from real
Scaling Voice AI in India: The Hard Part
IRCTC proves voice AI can handle national-scale demand—but only if you design for the hard stuff first. Here’s the reality. Why India-Scale Voice Is Uniquely Hard 1) Language & Accents * Dozens of languages and dialects; heavy Hindi–English/Bengali–English switching. * Wide variance in accents, speech rates, and
Comparing Approaches for End-of-Speech Detection in Voice AI
While building RapidaAI, one of the biggest challenges we are navigating is figuring out when a user has actually finished speaking — especially across different languages, accents, and natural speaking styles. Get it wrong, and the conversation feels off: interrupt too soon, and the user has to repeat themselves; wait too
Voice is the growth engine in a $100B CPaaS future
The #CPaaS market isn’t just growing, it is exploding. The CPaaS Acceleration Alliance notes that the market is worth $32 billion in 2025 and could reach $100 billion by 2030. And the analysts agree that growth hinges on how well providers integrate #AI and new channels to create intelligent
Revolutionise Your Customer Service with AI: Discover the Power of RAG as a service
In today's customer service landscape, where accuracy is critical, a single misstep can lead to widespread dissatisfaction. The introduction of generative AI has marked a significant development in this sector, but it also introduces the challenge of AI "hallucinations" — misleading or incorrect AI-generated responses. Such errors