Siyi Gu, Jialin Chen, Sophia Zhou, Arman Cohan, Rex Ying · 2026-06-19 · 7 min AI

重新思考奖励监督: 条件自蒸馏

Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation

推理语言模型的后训练通常由具有可验证奖励的监督蒸馏和强化学习驱动。蒸馏...

Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards. Distillat...

01
yxc0433 · 2026-06-19 · 3 min AI

媒体中的麻省理工学院: 对于科技的未来, "马萨诸塞州绝对可以领先"

MIT in the media: For the future of tech, "Massachusetts can absolutely lead"

6 月 9 日, 《波士顿环球报》发布了 2026 年 “Tech Power Players” 名单, 表彰了全马技术和商业领域 50 位有影响力的当地领导者...

On June 9, The Boston Globe released its 2026 “Tech Power Players” list, recognizing 50 influential local leaders in technology and business across Ma...

02
Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin, Alexander Theus, Thomas Hofmann, Valentina Boeva, T. Konstantin Rusch, Antonio Orvieto · 2026-06-18 · 5 min AI

定点 Reasoners: 稳定且自适应的深环变压器

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

循环架构为学习需要组合推理的任务的逐步过程提供了归纳偏差。电子数量...

Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of e...

03
Qi Chai, Wenhao Shen, Nanjie Yao, Yue Xia, Kaiyong Zhao, Jie Ma, Guosheng Lin, Hao Wang · 2026-06-18 · 9 min AI

EvolveNav: 用于零射击目标导航的主动预反射和自进化记忆

EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation

零射击对象目标导航 (ZS-OGN) 需要具体代理在没有任何事先训练的情况下探索和定位目标对象。为此,最近...

Zero-Shot Object-Goal Navigation (ZS-OGN) requires embodied agents to explore and locate target objects without any prior training. To this end, recen...

04
Adam Zewe | MIT News · 2026-06-18 · 10 min AI

AI 能告诉你钥匙落在哪里?

Could AI tell you where you left your keys?

一位汽车工厂的工人可以记得前一天晚上她把部分组装好的部件留在的储物箱,,并迅速返回那个地点去取东西。

An auto factory worker can remember the storage bin where she left a partly assembled component the night before, and quickly return to that spot to p...

05
Steve Nadis | MIT Laboratory for Information and Decision Systems · 2026-06-18 · 8 min AI

在博弈论中, 通才有时会战胜专家

In game theory, generalists sometimes win out over specialists

无论您’是与单一对手玩扑克,还是发现自己与另一位潜在买家陷入购房竞价战,,您都是......

Whether you’re playing poker against a single opponent or find yourself in a bidding war over a home purchase with another prospective buyer, you are ...

06
Nathan Gavenski, Juarez Monteiro, Francisco Galuppo, Adriano Veloso, Odinaldo Rodrigues · 2026-06-17 · 9 min AI

如有疑问, 计划好: 致力于反应式强化学习的小语言模型审议

When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning

强化学习 (RL) 策略在不熟悉的环境中通常会退化,因为它们缺乏明确的深思熟虑。我们建议计划, 对齐, 提交,...

Reinforcement Learning (RL) policies often degrade in unfamiliar environments because they lack explicit deliberation. We propose Plan, Align, Commit,...

07
Yanan Long · 2026-06-17 · 7 min AI

前沿人工智能评估公共档案的贝叶斯推理和决策审计

Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations

公共人工智能评估通常被解读为终端排行榜,,但潜在的证据是由报告规则, 基准形成的选择性时间序列...

Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting rules, bench...

08
Liam McDonnell | Office of Innovation and Strategy · 2026-06-17 · 4 min AI

MIT的 新制造业倡议蓄势待发

MIT’s Initiative for New Manufacturing builds momentum

5 月, 新制造业计划 (INM) 通过麻省理工学院制造周, 庆祝一周年,为期四天的活动吸引了更多...

In May, the Initiative for New Manufacturing (INM) marked its first anniversary with MIT Manufacturing Week, four days of events that attracted more t...

09
Xinyu Qiu, Yunzhu Zhang, Heng Jia, Shuheng Shen, Changhua Meng, Linchao Zhu · 2026-06-16 · 4 min AI

VISTA: GUI 基础的视图一致自我验证培训

VISTA: View-Consistent Self-Verified Training for GUI Grounding

当应用组相对策略优化 (GRPO) 进行 GUI 基础时, 部署是从单个屏幕截图视图中采样的; 组通常变得...

When applying Group Relative Policy Optimization (GRPO) for GUI Grounding, rollouts are sampled from a single screenshot view; groups often become eit...

10