风格字幕文本转语音系统使用自然语言来控制语音特征,,但单个单词如何影响声音输出仍不清楚。理解这一点对于诊断故障模式和提高表达 TTS 的可控性至关重要。我们提出了语音扩散模型的交叉注意归因,,首次将 DAAM 框架应用于语音领域, 并将其应用于 CapSpeech-TTS。我们的方法跨 25 层和 24 个 ODE 步骤提取每个代币的热图。我们分析了 3,600(stylecaption,texttranscript) 组合,其中包含 120 个 stylecaption,每个 , 生成 30 个文本转录本,揭示了字幕标记如何塑造波形。结果显示: (1) 风格标记的时间方差低于内容/功能标记, 确认全局调节; (2) 风格注意力与 F0 和能量相关; (3) 早期步骤和深层中的风格调节峰值; (4) 注意力熵在第 17 层达到最小值, 与风格重要性峰值同时出现, 表明在风格最关键的阶段实现最大的网络选择性。这是第一个关于自然语言如何影响语音扩散模型中交叉注意力的研究
Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure modes and improving controllability in expressive TTS. We propose cross-attention attribution for speech diffusion models, adapting the DAAM framework to the speech domain for the first time, and apply it to CapSpeech-TTS. Our method extracts per-token heatmaps across 25 layers and 24 ODE steps. We analyze 3,600 (style caption, text transcript) combinations comprising 120 style captions conditioning the generation of 30 text transcripts each, revealing how caption tokens shape waveforms. Results show: (1) style tokens have lower temporal variance than content/function tokens, confirming global conditioning; (2) style attention correlates with F0 and energy; (3) style conditioning peaks in early steps and deep layers; (4) attention entropy reaches its minimum at layer 17, co-occurring with the style importance peak, indicating maximal network selectivity at the most style-critical stage. This is the first study of how natural language influences cross-attention in speech diffusion models
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)