Full Duplex

Full-duplex model listens and speaks simultaneously, completely removing the turn-taking assumption: overlapping speech and arbitrary interruption are natural behavior, requiring no detection.

  • Pioneer – Moshi – parallel user and model audio streams + textual “inner monologue”, ~200ms latency.
  • Thesis – interactivity should be built into the model, not bolted on (TML Interaction Models, micro-rounds ~200ms).
  • GPT-Live – brought to industrial scale: continuous processing of input and output, tens of decisions per second, plus separation of fast and slow thinking via background model delegation.

Related: Voice paradigms, Fast and slow thinking architecture