视觉语言模型,大型语言模型,和多模态大型语言模型的最新进展改善了场景理解,决策,轨迹预测,和视觉问答等自动驾驶任务。然而, 评估这些模型是否能够可靠地推理安全关键事件仍然具有挑战性。为了解决这个差距,,我们提出了 AUTOPILOT-VQA, 一个以事件为中心的视觉问答基准,用于行车记录仪视频理解。该数据集通过围绕现实世界驾驶事件和接近事故设计的结构化问题来评估不同的系统。该基准涵盖多种安全相关类别,,包括天气和照明条件,交通环境,道路布局,路面状况,标牌,涉及实体,事故发生,影响位置,和可避免性相关推理。通过要求模型回答有关上下文场景属性和事件级事件细节的基础问题, AUTOPILOT-VQA 超越了对象识别,转向了基于时间的, 安全感知推理。该数据集作为 AUTOPILOT CVPR 2026 竞赛的一部分发布,为评估不同场景下自动驾驶系统的可靠性提供了标准化基准。我们的基准测试支持为现实世界的自动驾驶开发更具可解释性的,、稳健的, 和安全意识视觉语言系统。
Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world driving incidents and near-incidents. The benchmark covers diverse safety-relevant categories, including weather and lighting conditions, traffic environment, road layout, road surface state, signage, involved entities, accident occurrence, impact location, and avoidability-related reasoning. By requiring models to answer grounded questions about both contextual scene properties and event-level incident details, AUTOPILOT-VQA moves beyond object recognition toward temporally grounded, safety-aware reasoning. The dataset is released as part of the AUTOPILOT CVPR 2026 competition and provides a standardized benchmark for assessing the reliability of autonomous driving systems in different scenarios. Our benchmark support developments for more interpretable, robust, and safety-conscious vision-language systems for real-world autonomous driving.
科目: 人工智能 (cs.AI); 计算机视觉和模式识别 (cs.CV)
Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)