评估 LLM 输出仍然是 NLP 的主要瓶颈: 人工评估昂贵且缓慢, 词汇指标与人类对开放式生成的判断相关性较差, 并且整体 LLM 法官经常产生难以调试的不透明分数。我们提出 BINEVAL, 一个框架,它将评估标准分解为原子二元问题,并将结果聚合为可解释的, 多维分数。给定任务提示,,元提示会生成细粒度的评估问题,,法学硕士会针对每个输出,独立回答这些问题,从而产生透明的问题级反馈以及校准的总体分数。这种分解使得评估更容易检查,,更容易诊断,,并可直接用于迅速改进。在 SummEval, Topical-Chat, 和 QAGS, 中,BINEVAL 匹配或优于强大的基线,包括 UniEval 和 G-Eval,,在 QAGS 等事实一致性基准上取得了特别好的结果。除了与人类判断, BINEVAL 的竞争相关性之外,BINEVAL 更好地匹配人类分数分布,并避免了先前 LLM 法官, 中常见的上限效应,从而更好地区分边缘输出和明显有缺陷的输出。我们进一步表明,相同的问题级反馈支持迭代提示优化,,在自我更新和跨模型更新设置下,改进了 IFBench 上汇总提示和生成提示的评估者提示。总体而言, BINEVAL 提供了一个与任务无关的, 免培训, 和可解释的评估框架,将强大的经验性能与实际诊断和优化价值相结合。
Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decomposes evaluation criteria into atomic binary questions and aggregates the resulting verdicts into interpretable, multi-dimensional scores. Given a task prompt, a meta-prompt generates fine-grained evaluation questions, and an LLM answers them independently for each output, yielding transparent question-level feedback together with calibrated overall scores. This decomposition makes evaluation easier to inspect, easier to diagnose, and directly usable for prompt improvement. Across SummEval, Topical-Chat, and QAGS, BINEVAL matches or outperforms strong baselines including UniEval and G-Eval, with especially strong results on factual consistency benchmarks such as QAGS. Beyond competitive correlation with human judgments, BINEVAL better matches human score distributions and avoids the ceiling effects common in prior LLM judges, leading to better discrimination between borderline and clearly flawed outputs. We further show that the same question-level feedback supports iterative prompt optimization, improving evaluator prompts on summarization and generation prompts on IFBench under both self-update and cross-model update settings. Overall, BINEVAL provides a task-agnostic, training-free, and interpretable evaluation framework that combines strong empirical performance with practical diagnostic and optimization value.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)