使用计算机的代理 (CUA) 正在数字世界中快速发展。 CUA 轨迹记录了代理的 操作, 状态, 和推理。验证它是否完成任务指令是 CUA 评估, 数据管理, 和强化学习的核心。人工编写的验证者和人类注释者都无法提供大规模, 的验证,因此该领域越来越多地转向视觉语言模型 (VLMs) 作为 CUA 轨迹的判断者。但一个基本问题长期以来一直未被检验: 这些 VLM 判断是否足够可靠? 为了系统地研究它,,我们引入 OSReward, 一个现实的, 高质量基准,用于根据 CUA 轨迹评估 VLM 判断。这些轨迹来自不同的代理骨干,在平台, 上执行经过人工验证的指令,然后通过多阶段人工注释严格标记真实判决。在此基础上,,我们推导出 OSReward-Hard, 挑战集,集中真正困难的情况, 和 OSReward-Multi,以实现细粒度效率和对齐评分。迄今为止对 VLM 法官最全面的评估发现,即使是最先进的模型也达不到理想的法官,,他们都存在系统性的宽大偏见,将失败的运行错误地标记为成功。少数可靠且值得信赖的模型由于成本太高而无法大规模运行,,而经济实惠的开放模型则远远落后。为了弥补这一差距,,我们为 CUA 社区构建并发布了 OS-Shepherd-100K, 开放式推理注释轨迹判断语料库。在 it, 上,我们训练 OS-Shepherd (9B 和 35B), 开放奖励模型,这些模型提供低成本,、稳定, 和可靠的奖励信号, 匹配商业法官,成本比前沿低 30-60 倍。广泛的分析进一步为大规模可靠 CUA 奖励的设计提供了信息。我们的代码, 基准测试, 数据集, 和模型检查点可在此 https URL 获取。
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent的 actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, and are then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60x lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at this https URL.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 计算机视觉和模式识别 (cs.CV)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)