基于参考的自动评估方法在评估自然语言生成系统中发挥着至关重要的作用。现有的元评估主要衡量与人类判断或基准标签的一致性,,对受控条件下评估者行为的了解有限。我们引入了行为正确性假设,,这是一个用于评估基于参考的自动评估方法的补充框架。我们定义了正确性保留和正确性改变假设的分类法,并通过指定预期评分行为的受控响应转换来操作它们。我们评估不同的词汇,字符级,语义,基于LLM的,和混合评估器,并分析他们的假设级行为,稳定性,敏感性,重复运行变异性,配置敏感性,和再现性。我们的实验揭示了评估范式之间不同的行为权衡:没有评估者满足所有提出的正确性假设,并且具有相似总体表现的评估者可以表现出截然不同的行为概况。这些发现表明,行为正确性假设提供了传统聚合元评估所掩盖的诊断信息。
Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a taxonomy of correctness-preserving and correctness-altering assumptions and operationalize them through controlled response transformations that specify expected scoring behaviors. We evaluate diverse lexical, character-level, semantic, LLM-based, and hybrid evaluators and analyze their assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. Our experiments reveal distinct behavioral trade-offs across evaluation paradigms: no evaluator satisfies all proposed correctness assumptions, and evaluators with similar aggregate performance can exhibit substantially different behavioral profiles. These findings demonstrate that behavioral correctness assumptions provide diagnostic information obscured by conventional aggregate meta-evaluation.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)