现实的长期生产力工作很大程度上取决于特定于用户的计算机环境,,其中大部分工作上下文都是通过目录结构和内容丰富的工件来存储和组织的。为了扩展此类生产力场景,的合成数据创建,我们引入了规模,的合成计算机,这是一种可扩展的方法,用于创建具有现实文件夹层次结构和内容丰富的工件(例如,文档,电子表格,和演示文稿)的环境。在每台合成计算机,上,我们运行长期模拟:,一个代理创建特定于计算机的用户的生产力目标,需要多个专业交付成果和大约一个月的人工工作;,然后另一个代理充当该用户并继续在整个计算机上工作 - 例如,导航文件系统以接地,与模拟协作者,协调并生成专业工件 - 直到这些目标完成。在初步实验, 中,我们创建了 1,000 台合成计算机,并在它们上运行长期模拟; 每次运行需要超过 8 小时的代理运行时间,平均跨度超过 2,000 轮。这些模拟产生了丰富的体验式学习信号,,其有效性通过域内和域外生产力评估的代理性能的显着改进得到了验证。鉴于人物角色在十亿规模, 上非常丰富,这种方法原则上可以扩展到数百万甚至数十亿个合成用户世界,并具有足够的计算,,从而能够更广泛地覆盖不同的职业, 角色, 环境, 环境, 和生产力需求。我们认为,可扩展的合成计算机创建, 与大规模模拟, 相结合,作为长期生产力场景中智能体自我改进和智能体强化学习的基础,非常有前景。

Realistic long-horizon productivity work is strongly conditioned on user-specific computer environments, where much of the work context is stored and organized through directory structures and content-rich artifacts. To scale synthetic data creation for such productivity scenarios, we introduce Synthetic Computers at Scale, a scalable methodology for creating such environments with realistic folder hierarchies and content-rich artifacts (e.g., documents, spreadsheets, and presentations). Conditioned on each synthetic computer, we run long-horizon simulations: one agent creates productivity objectives that are specific to the computer的 user and require multiple professional deliverables and about a month of human work; another agent then acts as that user and keeps working across the computer -- for example, navigating the filesystem for grounding, coordinating with simulated collaborators, and producing professional artifacts -- until these objectives are completed. In preliminary experiments, we create 1,000 synthetic computers and run long-horizon simulations on them; each run requires over 8 hours of agent runtime and spans more than 2,000 turns on average. These simulations produce rich experiential learning signals, whose effectiveness is validated by significant improvements in agent performance on both in-domain and out-of-domain productivity evaluations. Given that personas are abundant at billion scale, this methodology can in principle scale to millions or even billions of synthetic user worlds with sufficient compute, enabling broader coverage of diverse professions, roles, contexts, environments, and productivity needs. We argue that scalable synthetic computer creation, together with at-scale simulations, is highly promising as a foundational substrate for agent self-improvement and agentic reinforcement learning in long-horizon productivity scenarios.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)