🗣️ WordVoice — Word-Level Controllable TTS
WordVoice performs zero-shot voice cloning with explicit, decoupled word-level control over five acoustic dimensions. Give it a reference clip + its transcript, then write the target text — attach inline control tags to any word to steer its prosody. Un-tagged words are planned automatically by the model.
Inline control tags
Attach one or more tags immediately after a word, e.g.
crazy[eng:0.9][dur:400] or never[pit:0][ton:ffall].
| Tag | Meaning | Range |
|---|---|---|
[dur:N] |
duration | milliseconds (≈40 ms/token, clamped 40–1400) |
[eng:x] |
energy / loudness | 0–1 |
[pit:x] |
pitch (core F0) | -1–1 |
[bnd:b] |
pause boundary after word | b0 (none) … b4 (long) |
[ton:t] |
tone / prosody shape | flat, rise, rrise, fall, ffall, peak, valley |
Anything you leave off is decided by the model's "acoustic thinking" planner.
Examples
| Reference / prompt audio (≤ ~20s, sr ≥ 16 kHz) | Prompt text (transcript of the reference audio) | Text to synthesise (add inline control tags, see below) | Language |
|---|