Chat Completions vs Responses vs WebSocket: TTFT and TRT in Rapida
In voice AI, 300–500ms latency differences completely change conversation quality. Yet most teams still benchmark only tokens/sec instead of end-to-end conversational responsiveness. LLMs are evolving fast, and so is OpenAI. Every few months, the model layer changes in a meaningful way: better reasoning, better tool calling, better streaming