大型语言模型代理越来越被设想为永远在线的个人助理,可以访问用户数字世界中的任何相关内容。然而,当前的系统仅在世界的一小部分, 上运行,限制了上下文相关的推理和有效的帮助。现有的基准测试同样仅提供部分用户状态,因此无法捕获如此广泛的, 始终在线设置中的性能。为了解决这一差距,,我们引入了 Claw-Anything, 基准,该基准可沿三个维度扩展代理上下文: 长期活动历史, 相互依赖的后端服务, 以及跨多个设备的集成 GUI 和 CLI 交互。为了实例化此设置,,我们通过多轮事件注入, 模拟数月的用户活动,产生复杂的世界状态和现实噪声,,包括不相关的事件和冲突信号。智能体必须在丰富的上下文环境中进行推理,同时对此类噪音保持鲁棒性。这种扩大的范围还可以评估主动援助,,要求代理预测用户需求并及时提供建议。实验表明,GPT-5.5 仅达到 34.5% pass@1,,大大低于之前的基准,,凸显了当前代理能力与始终在线个人协助的需求之间的差距。除了基准,之外,我们还发布了一个自动数据生成管道,该管道可产生2,000训练环境,并将基本模型改进23.7%,,展示了其可扩展数据基础设施的实用性。
Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user的 digital world. Yet current systems operate over only narrow slices of that world, limiting context-sensitive reasoning and effective assistance. Existing benchmarks similarly provide only partial user state and therefore fail to capture performance in such a broad, always-on setting. To address this gap, we introduce Claw-Anything, a benchmark that expands agent context along three dimensions: long-horizon activity histories, interdependent backend services, and integrated GUI and CLI interaction across multiple devices. To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise. This expanded scope also enables the evaluation of proactive assistance, requiring agents to anticipate user needs and deliver timely recommendations. Experiments show that GPT-5.5 achieves only 34.5% pass@1, substantially below prior benchmarks, underscoring a gap between current agent capabilities and the demands of always-on personal assistance. Alongside the benchmark, we release an automated data-generation pipeline that yields 2,000 training environments and improves the base model by 23.7%, demonstrating its utility of scalable data infrastructure.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)