Xuhao Hu, Xi Zhang, Haiyang Xu, Kyle Qiao, Jingyi Yang, Xuanjing Huang, Jing Shao, Ming Yan, Jieping Ye · 2026-05-14 · 5 min AI

ToolCUA: 为计算机使用代理实现最佳 GUI 工具路径编排

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents

计算机使用代理 (CUAs) 可以通过原子 GUI 操作,(例如单击和键入,)和高级工具调用,(例如基于 API 的文件操作)进行操作...

Computer Use Agents (CUAs) can act through both atomic GUI actions, such as click and type, and high-level tool calls, such as API-based file operatio...

01
Carolyn Tiernan | MIT Open Learning · 2026-05-14 · 10 min AI

通用 AI 是 “a 通向 AI 流畅性的途径, 任何人都可以轻松访问, 任何地方”

Universal AI is “a pathway to AI fluency that’s accessible and approachable to anyone, anywhere”

“人工智能不再只是计算机科学家的专利;,它’将渗透到我们生活的方方面面,影响每一项业务,” ...

“Artificial intelligence is not just for computer scientists anymore; it’s going to permeate every aspect of our lives and influence every business,” ...

02
Carolyn Tiernan | MIT Open Learning · 2026-05-13 · 8 min AI

Q&A: 通过通用学习扩大 MIT的 的全球影响力

Q&A: Expanding MIT’s global reach through Universal Learning

MIT的 通用学习是麻省理工学院开放学习的一项新举措,旨在帮助世界各地的学习者做好准备,通过……应对复杂的全球挑战。

MIT's Universal Learning is a new initiative from MIT Open Learning designed to prepare learners everywhere to tackle complex global challenges throug...

03
Haoyang Su, Ying Wen · 2026-05-12 · 7 min AI

在选择性观察下学习具有结构化行动信用的 CLI 代理

Learning CLI Agents with Structured Action Credit under Selective Observation

命令行界面 (CLI) 代理正在成为代理与计算机交互的实用范例,通过不断发展的文件系统, 可执行命令...

Command line interface (CLI) agents are emerging as a practical paradigm for agent-computer interaction over evolving filesystems, executable command ...

04
Botos Csaba, Sreejan Kumar, Austin Tudor David Andrews, Laurence Hunt, Chris Summerfield, Joshua B. Tenenbaum, Rui Ponte Costa, Marcelo G. Mattar, Momchil Tomov · 2026-05-12 · 4 min AI

玩: 前沿 LRM 和人类游戏学习者之间的行为和大脑一致性的原因

Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners

人类在遇到新环境时快速学习抽象知识,并灵活运用这些知识来指导高效、智能的行为……

Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent act...

05
Wenxin Zhan · 2026-05-12 · 3 min AI

MPD$^2$-Router: 青光眼筛查和诊断中的掩模感知多专家先验正则化双头延迟路由器

MPD$^2$-Router: Mask-aware Multi-expert Prior-regularized Dual-head Deferral Router in Glaucoma Screening and Diagnosis

学习推迟 (L2D) 可以通过将困难的/uncertain病例转交给人类,,使青光眼筛查更安全,但标准公式忽略了专家的意见...

Learning-to-defer (L2D) can make glaucoma screening safer by routing difficult/uncertain cases to humans, yet standard formulations overlook expert av...

06
Manish Bhattarai, Ismael Boureima, Nishath Rajiv Ranasinghe, Scott Pakin, Dan O'Malley · 2026-05-12 · 8 min AI

基于评分标准的 RL: 针对可推广推理的结构化法官奖励

Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning

我们认为,将奖励分解为加权,可验证标准并使用LLM法官对其进行评分提供了部分学分优化信号......

We argue that decomposing reward into weighted, verifiable criteria and using an LLM judge to score them provides a partial-credit optimization signal...

07
James Petullo, Sonny George, Dylan Cashman, Nianwen Xue · 2026-05-12 · 8 min AI

VecCISC: 通过推理跟踪聚类和候选答案选择提高置信度知情的自我一致性

VecCISC: Improving Confidence-Informed Self-Consistency with Reasoning Trace Clustering and Candidate Answer Selection

扩展推理时间推理的标准技术是自我一致性,,其中多个候选答案是从法学硕士中抽取的,并且最...

A standard technique for scaling inference-time reasoning is Self-Consistency, whereby multiple candidate answers are sampled from an LLM and the most...

08
Ruiqi Lyu, Alistair Turcan, Bryan Wilder · 2026-05-11 · 9 min AI

SpatialEpiBench: 预测中的空间信息和流行病先验基准

SpatialEpiBench: Benchmarking Spatial Information and Epidemic Priors in Forecasting

准确的流行病预测对于公共卫生应对,资源分配,和疫情干预,至关重要,但由于资源稀少,仍然很困难...

Accurate epidemic forecasting is crucial for public health response, resource allocation, and outbreak intervention, but remains difficult with sparse...

09
Nafis Saami Azad, Raiyan Abdul Baten · 2026-05-11 · 3 min AI

人工智能引发的创意多样性崩溃的事前评估

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse

创造性的人工智能系统通常在个人效用,的水平上进行评估,但创造性的产出却被群体所消耗:,一个想法就会失去价值......

Creative AI systems are typically evaluated at the level of individual utility, yet creative outputs are consumed in populations: an idea loses value ...

10