从源文档自动生成幻灯片是大型语言模型(LLMs) 的一个重要应用。现有基准主要评估幻灯片完整性和技术深度,,而忽略了目标受众作为关键的现实因素。例如,, 专家要求严格的证明,,而决策者则优先考虑可操作的结论。为了弥补这一差距,,我们引入了 X+Slides,,这是专门为观众条件幻灯片生成而设计的基准。基于涵盖 113 个主题和 7 个演示场景的多样化语料库, X+Slides 采用由 8,133 去重, 源接地探针构建的动态评估框架。通过将特定于受众的效用权重分配给相同的基于源的探针, X+Slides 报告四个互补指标: 受众覆盖率衡量传达了多少受众基本信息, 领域覆盖率显示覆盖了哪些信息类型, 效率衡量每单位注意力成本提供的效用, 和正确性验证幻灯片声明是否得到来源支持。 DeepPresenter, SlideTailor, 和 NotebookLM 上的实验表明,当前系统可以在 $\tau_A=0.7$, 处恢复大量但仍不完整的受众基本信息:,DeepPresenter 达到最佳受众覆盖率 0.714, SlideTailor 达到 0.594,,NotebookLM 消融达到0.853,同时显示出明显的基础差异。这些结果表明,如果没有基于来源的评估,视觉质量和广泛的主题覆盖不应被视为证据支持。
Automatically generating slide decks from source documents is an important application of large language models (LLMs). Existing benchmarks primarily assess slide completeness and technical depth, while overlooking the target audience as a critical real-world factor. For instance, specialists demand rigorous proofs, whereas decision-makers prioritize actionable conclusions. To bridge this gap, we introduce X+Slides, a benchmark specifically designed for audience-conditioned slide generation. Built on a diverse corpus spanning 113 topics and seven presentation scenes, X+Slides employs a dynamic evaluation framework constructed from 8,133 deduplicated, source-grounded probes. By assigning audience-specific utility weights to the same source-grounded probes, X+Slides reports four complementary metrics: Audience Coverage measures how much audience-essential information is conveyed, Domain-wise Coverage shows which information types are covered, Efficiency measures delivered utility per unit of attention cost, and Correctness verifies whether slide claims are supported by the source. Experiments on DeepPresenter, SlideTailor, and NotebookLM show that current systems can recover a substantial but still incomplete part of audience-essential information: at $\tau_A=0.7$, DeepPresenter reaches a best Audience Coverage of 0.714, SlideTailor reaches 0.594, and the NotebookLM ablation reaches 0.853 while showing clear grounding differences. These results indicate that visual quality and broad topic coverage should not be treated as evidence support without source-grounded evaluation.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)