Voice Paradigms

Three paradigms of voice architecture:

  1. Cascade — VAD→ASR→LLM→TTS pipeline: modularity, but accumulation of latency and loss of information at the joints;
  2. End-to-end Omni model — a single model listens-thinks-speaks: lower latency, prosody is preserved, but still “turn-based” by VAD;
  3. Full duplex — the model listens and speaks simultaneously, deciding to speak/listen/be silent/interrupt dozens of times per second, without assuming turn-taking.

The common thread of all three: how to get rid of guessing the boundaries of the queue by silence.

Related: Cascade voice pipeline, Full duplex, Streaming speech perception