Renning Pang, Tian Lan, Leyuan Liu, Piao Tong, Sheng Cao, Xiaosong Zhang · 2026-05-17 · 10 min AI

LLM 工具使用的自适应推理和执行的基于案例的校准

Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use

工具的使用将大型语言模型扩展到参数知识之外,,但可靠的执行需要平衡适当的推理深度与严格的......

Tool use extends large language models beyond parametric knowledge, but reliable execution requires balancing appropriate reasoning depth with strict ...

01
Rongman Xu, Yifei Li, Tianzhe Zhao, Yanrui Wu, Bo Li, Hang Yan · 2026-05-17 · 7 min AI

二维一致性: 在自适应推理时间缩放中平衡预算和质量

Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling

大型语言模型 (LLMs) 表现出了卓越的推理能力。然而, 通过推理时间缩放来最大限度地发挥其潜力...

Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-time scaling f...

02
Riccardo Terrenzi, Maximilian von Zastrow, Serkan Ayvaz · 2026-05-17 · 5 min AI

为什么邻域很重要: Agentic GraphRAG 中的遍历上下文和来源

Why Neighborhoods Matter: Traversal Context and Provenance in Agentic GraphRAG

检索增强生成可以通过将答案扎根于外部证据, 来提高事实性,但 Agentic GraphRAG 使其对公民的意义变得复杂化......

Retrieval-Augmented Generation can improve factuality by grounding answers in external evidence, but Agentic GraphRAG complicates what it means for ci...

03
Evan Rose, Tushin Mallick, Matthew D. Laws, Cristina Nita-Rotaru, Alina Oprea · 2026-05-17 · 8 min AI

APWA: 用于可并行代理工作流的分布式架构

APWA: A Distributed Architecture for Parallelizable Agentic Workflows

基于大型语言模型 (LLMs) 的自主多智能体系统在独立解决复杂任务方面表现出了卓越的能力......

Autonomous multi-agent systems based on large language models (LLMs) have demonstrated remarkable abilities in independently solving complex tasks in ...

04
Shang Zhou, Wenhao Chai, Kaiyuan Liu, Huanzhi Mao, Qiuyang Mang, Jingbo Shang · 2026-05-16 · 9 min AI

OpenDeepThink: 通过 Bradley-Terry 聚合进行并行推理

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation

测试时计算扩展是改进 LLM 推理的主轴。现有的方法主要通过扩展单个推理轨迹来扩展深度......

Test-time compute scaling is a primary axis for improving LLM reasoning. Existing methods primarily scale depth by extending a single reasoning trace....

05
Jiayi Zhang, Yongfeng Gu, Jianhao Ruan, Maojia Song, Yiran Peng, Zhiguang Han, Jinyu Xiang, Zhitao Wang, Caiyin Yang, Yixi Ouyang, Bang Liu, Chenglin Wu, Yuyu Luo · 2026-05-15 · 7 min AI

利用代理进化

Harnessing Agentic Evolution

代理进化已成为通过迭代生成候选,来改进程序,工作流程,和科学解决方案的强大范例。

Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions by iteratively generating candidates,...

06
Alberto G. Rodríguez Salgado · 2026-05-15 · 3 min AI

历史锚: 先前的行为如何引导法学硕士做出不安全行为的决定

History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions

前沿法学硕士越来越多地被部署为代理,在由相同或不同的管理人员生成的大量先前工具调用日志后选择下一步操作。

Frontier LLMs are increasingly deployed as agents that pick the next action after a long log of prior tool calls produced by the same or a different m...

07
Ajinkya Naik, Chaitanya Garg, S. Akshay, Ashutosh Gupta, Kuldeep S. Meel · 2026-05-15 · 3 min AI

量化树集成的敏感性: 符号和组合方法

Quantifying Sensitivity for Tree Ensembles: A symbolic and compositional approach

决策树集成 (DTE) 是广泛的 AI 分类任务的流行模型, 用于多个安全关键领域,,因此非常...

Decision tree ensembles (DTE) are a popular model for a wide range of AI classification tasks, used in multiple safety critical domains, and hence ver...

08
Julia Mongo | Office of Distinguished Fellowships · 2026-05-15 · 8 min AI

麻省理工学院的两名学生被评为 2026 年骑士

Two from MIT named 2026 Knight

MIT硕士的学生Sunshine Jiang’25和Rupert Li’24是今年的 Knight-Hennessy奖学金的获得者。现在已进入第九个年头, 高度...

MIT master’s student Sunshine Jiang ’25 and Rupert Li ’24 are recipients of this year’s Knight-Hennessy Scholarship. Now in its ninth year, the highly...

09
Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal, Bing Liu, Yunzhong He · 2026-05-14 · 7 min AI

基于评分标准的强化学习中的奖励黑客

Reward Hacking in Rubric-Based Reinforcement Learning

具有可验证奖励的强化学习在数学和编码等领域实现了强大的训练后收益,,尽管许多开放式设置...

Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended setti...

10