🗣️ WordVoice — Word-Level Controllable TTS

WordVoice performs zero-shot voice cloning with explicit, decoupled word-level control over five acoustic dimensions. Give it a reference clip + its transcript, then write the target text — attach inline control tags to any word to steer its prosody. Un-tagged words are planned automatically by the model.

Paper · Model · Code · built on CosyVoice3

Language

Inline control tags

Attach one or more tags immediately after a word, e.g. crazy[eng:0.9][dur:400] or never[pit:0][ton:ffall].

Tag Meaning Range
[dur:N] duration milliseconds (≈40 ms/token, clamped 40–1400)
[eng:x] energy / loudness 01
[pit:x] pitch (core F0) -11
[bnd:b] pause boundary after word b0 (none) … b4 (long)
[ton:t] tone / prosody shape flat, rise, rrise, fall, ffall, peak, valley

Anything you leave off is decided by the model's "acoustic thinking" planner.

Examples
Reference / prompt audio (≤ ~20s, sr ≥ 16 kHz) Prompt text (transcript of the reference audio) Text to synthesise (add inline control tags, see below) Language