CraftTTS: Fine-Grained Prosody Control for Text-to-Speech

ABSTRACT: While zero-shot Text-to-Speech (TTS) models perform well in global voice cloning, they struggle with fine-grained prosodic control, as strict word-level intensity and tempo manipulation can disrupt acoustic priors and introduce artifacts. We propose CraftTTS, a three-stage framework enabling stable word-level control without sacrificing fluency. First, a compute-driven zero-shot pipeline automatically constructs large-scale preference pairs without manual annotation. Second, joint Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) improve sensitivity to local prosodic tags. Third, Group Relative Policy Optimization (GRPO) with a multi-dimensional prosodic reward balances local controllability and global naturalness by regulating intelligibility, intensity contrast, and rhythm. Experiments demonstrate that CraftTTS achieves state-of-the-art fine-grained expressiveness while preserving zero-shot generation capability.

Model Architecture

CraftTTS is a three-stage framework for fine-grained word-level prosody control in zero-shot TTS. We first construct preference data via multi-round sampling, then establish basic tag controllability through SFT and DPO. Finally, we apply GRPO with multi-dimensional prosodic rewards to improve robustness and suppress acoustic artifacts. Below is an overview of our model's pipeline.

Model Architecture

Stage 1: Construction Datasets Demos

How we build the data: To overcome the scarcity of fine-grained prosody data without human annotation, we introduce a compute-driven zero-shot pipeline. First, we leverage the reasoning capability of LLMs to assign span-level prosodic tags to text based on semantic context. Then, we use IndexTTS2 Model to perform multi-round, tag-conditioned sampling.

Winning vs. Losing: We construct contrastive preference pairs by comparing tag-conditioned generation against neutral generation. Winning Samples are synthesized with explicit prosodic tags, successfully reflecting the targeted intensity and tempo. Losing Samples are the baseline outputs generated without these tags. These pairs serve as the foundation for our Stage 2 DPO training to enforce tag adherence.

Text with Prosody Tags Winning Samples (Chosen) Losing Samples (Rejected)
Loading data...

Word-Level Stress & Speaking Rate Control Demos

Tag-based Control: CraftTTS achieves precise local prosody control via explicit text instructions. We utilize four primary tags: <strong> and <weak> for modulating intensity (stress and loudness), and <fast> and <slow> for regulating speaking rate (tempo).

Notice how our model dynamically adjusts pitch, energy, and duration strictly at the specified text spans. Compared to behavioral cloning (SFT) and the CosyVoice 2 baseline, CraftTTS executes these fine-grained transitions smoothly without unnatural pauses or global emotional leakage.

Reference Audio Target Text (with tags) CosyVoice 2 CraftTTS (SFT only) CraftTTS (Ours Full)
Loading data...

General Zero-Shot Demos (No Control Tags)

Purpose of this section: Aggressive localized acoustic conditioning can sometimes disrupt the global acoustic prior of autoregressive models. To prove that our three-stage alignment (SFT → DPO → GRPO) does not sacrifice base capabilities, these demos are generated without any explicit prosody tags.

These results evaluate the model's fundamental ability to perform high-fidelity voice cloning and maintain natural rhythmic flow. Thanks to our pause-aware ASR reward and region-aware tempo regularization during the RL stage, CraftTTS not only preserves the baseline's cloning ability but often generates more natural phrasing, accurate pauses, and stable global rhythm compared to the baselines.

Reference Audio Target Text (Plain) IndexTTS 2 CosyVoice 2 CraftTTS (Ours)
Loading data...