CraftTTS is a three-stage framework for fine-grained word-level prosody control in zero-shot TTS. We first construct preference data via multi-round sampling, then establish basic tag controllability through SFT and DPO. Finally, we apply GRPO with multi-dimensional prosodic rewards to improve robustness and suppress acoustic artifacts. Below is an overview of our model's pipeline.
How we build the data: To overcome the scarcity of fine-grained prosody data without human annotation, we introduce a compute-driven zero-shot pipeline. First, we leverage the reasoning capability of LLMs to assign span-level prosodic tags to text based on semantic context. Then, we use IndexTTS2 Model to perform multi-round, tag-conditioned sampling.
Winning vs. Losing: We construct contrastive preference pairs by comparing tag-conditioned generation against neutral generation. Winning Samples are synthesized with explicit prosodic tags, successfully reflecting the targeted intensity and tempo. Losing Samples are the baseline outputs generated without these tags. These pairs serve as the foundation for our Stage 2 DPO training to enforce tag adherence.
| Text with Prosody Tags | Winning Samples (Chosen) | Losing Samples (Rejected) |
|---|---|---|
| Loading data... | ||
Tag-based Control: CraftTTS achieves precise local prosody control via explicit text instructions. We utilize four primary tags: <strong> and <weak> for modulating intensity (stress and loudness), and <fast> and <slow> for regulating speaking rate (tempo).
Notice how our model dynamically adjusts pitch, energy, and duration strictly at the specified text spans. Compared to behavioral cloning (SFT) and the CosyVoice 2 baseline, CraftTTS executes these fine-grained transitions smoothly without unnatural pauses or global emotional leakage.
| Reference Audio | Target Text (with tags) | CosyVoice 2 | CraftTTS (SFT only) | CraftTTS (Ours Full) |
|---|---|---|---|---|
| Loading data... | ||||
Purpose of this section: Aggressive localized acoustic conditioning can sometimes disrupt the global acoustic prior of autoregressive models. To prove that our three-stage alignment (SFT → DPO → GRPO) does not sacrifice base capabilities, these demos are generated without any explicit prosody tags.
These results evaluate the model's fundamental ability to perform high-fidelity voice cloning and maintain natural rhythmic flow. Thanks to our pause-aware ASR reward and region-aware tempo regularization during the RL stage, CraftTTS not only preserves the baseline's cloning ability but often generates more natural phrasing, accurate pauses, and stable global rhythm compared to the baselines.
| Reference Audio | Target Text (Plain) | IndexTTS 2 | CosyVoice 2 | CraftTTS (Ours) |
|---|---|---|---|---|
| Loading data... | ||||