Voice Paradigms
Three paradigms of voice architecture:
- Cascade — VAD→ASR→LLM→TTS pipeline: modularity, but accumulation of latency and loss of information at the joints;
- End-to-end Omni model — a single model listens-thinks-speaks: lower latency, prosody is preserved, but still “turn-based” by VAD;
- Full duplex — the model listens and speaks simultaneously, deciding to speak/listen/be silent/interrupt dozens of times per second, without assuming turn-taking.
The common thread of all three: how to get rid of guessing the boundaries of the queue by silence.
Related: Cascade voice pipeline, Full duplex, Streaming speech perception