Human-like Speech Synthesis

The “perfection” of traditional TTS is the problem: overly smooth speech gives away the machine.

Human “imperfections” (pauses, “um’s,” repetitions) are an external manifestation of thought.

Solution: LLM outputs control tags ([THINKING] – pause with filler, [EMO:happy], [SPEED:0.8x]), TTS acts as a multimodal generator: plain text → speech, tag → corresponding non-speech sound.

Only the LLM knows where a pause for deliberation is needed.

Implementation: a custom model with native tag support or voice cloning with a library of reference recordings for different emotions and speeds.

Related: Voice paradigms, Cascaded voice pipeline