AI 评估结果是大规模产生的,但在排行榜, 模型卡, 基准论文, 和公司博客中报告不一致。成本是解释性的: 读者无法可靠地比较不同来源的结果, 识别报告遗漏的内容, 或追踪总体声明的基础证据。最近的努力解决了孤立的组件,但留下了三个差距:它们只覆盖了评估生命周期的一小部分,并且不组成单个可解释的记录;它们指定的静态表示不能区分不同利益相关者对相同证据提出的问题;并且它们仍然是纸上的建议,缺乏大规模采用所需的提取基础设施。我们提供 \EvalCards{}, 一个操作报告层,它将基准元数据, 评估运行数据, 和模型元数据组成统一的记录。我们 (1) 从对 52 篇论文和 10 名利益相关者访谈的结构化审查中得出报告模式, (2) 实施四个解释信号 (可重复性, 文档完整性, 出处和风险, 和分数可比性), 通过针对研究和非研究受众进行校准的读者模式呈现, 和 (3) 部署适用的监控工具\EvalCards{} 跨 5,816 模型, 635 基准, 和 101,843 结果, 暴露当前报告实践中的系统性差距。
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sources, identify what a report omits, or trace an aggregate claim to its underlying evidence. Recent efforts address isolated components but leave three gaps: they cover only narrow slices of the evaluation lifecycle and do not compose into a single interpretable record; they specify static representations that do not differentiate the questions different stakeholders bring to the same evidence; and they remain proposals on paper, lacking the extraction infrastructure required for adoption at scale. We present \EvalCards{}, an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. We (1) derive a reporting schema from a structured review of 52 papers and 10 stakeholder interviews, (2) implement four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability), rendered through reader modes calibrated to research and non-research audiences, and (3) deploy a monitoring tool that applies \EvalCards{} across 5,816 models, 635 benchmarks, and 101,843 results, surfacing systematic gaps in current reporting practice.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)