Arya Labroo, Mengjie Qian, Kate Knill · 2026-08-09 · 7 min AI

使用概念激活向量的 L2 口语评估系统的偏差分析

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

自动口语评估系统越来越多地部署在高风险环境中,以对第二语言(L2)学习者'口语测试,进行评分

Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making ...

01
Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, Yuan Xue · 2026-08-09 · 4 min AI

HarnessOpt-Bench: 在 Harness Optimization 中评估法学硕士

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

随着法学硕士越来越多地部署在代理系统中,,他们的能力不仅取决于模型权重,还取决于工具:,提示...

As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts...

02
Sagar Tamang, Ayush Vyas, Tabarakul Hazarika · 2026-08-09 · 5 min AI

超越 Top-K: 用可解释的代理操作取代黑盒检索

Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

长文档的检索增强生成主要由一种设计: 块文本, 嵌入块, 并显示 top-k 最近邻...

Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbo...

03
Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li · 2026-08-09 · 10 min AI

TRAJDEBUG: 跟踪错误生命周期以识别长期代理轨迹中的严重故障

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

基于 LLM 的代理系统在复杂领域, 中表现出了卓越的能力,但同时也面临着级联错误和调试困难的问题。铬...

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Cr...

04
Jerzy Stefanowski · 2026-08-08 · 7 min AI

评估静态和变化数据解释方法的挑战

Challenges in Evaluating Explanation Methods for Static and Evolving Data

本文解决了可解释人工智能 (XAI) 在评估不足方面的局限性。它们通过...进行说明。

This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated throug...

05
Sarvesh Baskar, Zikui Cai, Shayan Shabihi, Anirudh Satheesh, Muhammad R. Islam, Udari Madhushani Sehwag, Tom Goldstein, Furong Huang · 2026-08-08 · 7 min AI

低频陷阱: 视频语言模型无法完成简单的事件记账

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

现实世界的视频基准测试提供了广泛的覆盖范围,,但其固定剪辑纠缠了事件计数, 速率, 持续时间, 和视觉复杂性, 导致失败......

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure ...

06
Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi, Daniel Kang, Roy Ka-Wei Lee, Koustuv Saha, Christian Poellabauer, Christopher Lee, Sajeev Singh, Piyum Zonooz, Navin Kumar, Zeeshan Ahmed, Priyadarshini Kachroo · 2026-08-08 · 10 min AI

追踪心脏: 心力衰竭特征工程的循证管道

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

电子健康记录(EHR)特征工程是临床研究和AI,的主要瓶颈,占数据科学家'工作量的39-45%...

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' worklo...

07
Science Policy Initiative · 2026-08-08 · 4 min AI

将研究与国会山的政策联系起来

Connecting research to policy on Capitol Hill

今年春天, 25 名麻省理工学院的学生和博士后前往华盛顿与国会工作人员会面,并倡导联邦政府持续投资......

This spring, 25 MIT students and postdocs traveled to Washington to meet with congressional staffers and advocate for sustained federal investment in ...

08
Steve Nadis | Department of Nuclear Science and Engineering · 2026-08-08 · 10 min AI

解决溶剂问题

Solving the solvent problem

锂离子电池是当今的电动汽车和电池储能系统行业,的主要选择,但它们包含许多关键...

Lithium-ion batteries are the leading choice in today’s electric vehicle and battery energy storage system industries, but they contain a number of cr...

09
Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, Siheng Chen · 2026-08-07 · 8 min AI

ABSeeker: 通过答案回溯信用分配训练长期搜索代理

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

长视野搜索代理必须执行多个连续操作 (steps) 来搜索, 检索, 验证, 并整合证据以获得最终答案。 ...

Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. ...

10