Lauren Hinkel | MIT-IBM Computing Research Lab · 2026-09-08 · 7 min AI

从 MIT 到 IBM, 加速 AI 和量子部署

From MIT to IBM, expediting AI and quantum deployment

对于不同的研究人员来说,从基于理论的研究过渡到关注现实世界应用的经历可能会有很大差异。 ...

The experience of transitioning from research based in theory to focusing on real-world application can vary significantly for different researchers. ...

01
Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang · 2026-08-15 · 8 min AI

QuoteBench: 匹配分数如何隐藏命令路径故障

QuoteBench: How Matched Scores Can Hide Command-Path Failures

LLM 编码代理通过接口发出 Bash 命令,这些命令可以序列化, 包装, 并重新解析模型输出。仅匹配的执行分数并不能区分...

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot dis...

02
Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu · 2026-08-15 · 5 min AI

OmniScientist: 全模式全学科人工智能科学家

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

基础模型的最新进展使 AI 科学家能够自动化日益完整的研究工作流程,,从假设生成到分析……

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and c...

03
Rodrigo Guedes de Souza, Alison R. Panisson · 2026-08-14 · 3 min AI

谁的想法最好取决于你让他们多长时间: 法学硕士评估中的预算相关排名

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

大型语言模型的标准评估假设在推理条件下模型排名稳定。我们通过改变...来挑战这个假设。

Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the to...

04
Aleksandra Kalisz, Jack Simons, Krisztina Sinkovics, Noam Ghenassia, Shikha Surana, Henry Moss, Paul Duckworth · 2026-08-14 · 4 min AI

如何使用您的 Oracle 预算%3蛋白质结构预测模型实用指南

How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models

蛋白质结构预测的基础模型在某些目标上仍然不可靠。外部预言机可以标记并纠正这些故障,,但生物......

Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but bio...

05
Yuzhong Shen, Masha Sosonkina, Peng Xu, Mark S. Gordon · 2026-08-14 · 10 min AI

传统 HPC 现代化的代理工作流程: 转换 GAMESS 的两电子整体核心

An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS

对遗留 Fortran 进行现代化改造是一个很大的问题:,转换是单独的例行程序,,但代码库可能非常庞大,,并且跨越了大部分...

Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of...

06
Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor · 2026-08-14 · 8 min AI

VAKRA: 评估跨 API 的多跳推理和工具使用策略下的检索

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

部署在企业环境中的代理必须跨结构化 API 和文档集合进行推理,,但现有基准评估这些功能...

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilitie...

07
Saman Marandi, Yu-Shu Hu, Mohammad Modarres · 2026-08-14 · 5 min AI

使用检索增强大型语言模型将动态主逻辑模型构建为复杂系统诊断的知识图

Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models

动态主逻辑(DML)通过将功能目标链接到底层结构来提供表示系统行为的分层框架...

Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structura...

08
Adam Zewe | MIT News · 2026-08-13 · 4 min AI

医疗人工智能援助的好处因用户专业知识而异

The benefits of medical AI assistance vary based on user expertise

在设计协助用户进行疾病诊断的人工智能系统时,一刀切的方法可能不是的最佳策略。一个...

A one-size-fits-all approach likely isn’t the best strategy when designing artificial intelligence systems that assist users in disease diagnosis. A n...

09
Steve Nadis | Department of Nuclear Science and Engineering · 2026-08-13 · 3 min AI

解决溶剂问题

Solving the solvent problem

锂离子电池是当今的电动汽车和电池储能系统行业,的主要选择,但它们包含许多关键...

Lithium-ion batteries are the leading choice in today’s electric vehicle and battery energy storage system industries, but they contain a number of cr...

10