Qijia He, Jiayi Cheng, Chenqian Le, Rui Wang, Xunmei Liu, Yixian Chen, Jie Mei, Zhihao Wang, Xupeng Chen, Yuhuan Chen, Tao Wang · 2026-07-22 · 7 min AI

CodeRescue: 编码代理的预算校准恢复路由

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

编码代理越来越多地在可执行环境中运行,其中失败的尝试会产生可操作的反馈,而不仅仅是错误的答案......

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answ...

01
Lopez Jhon, Hinojosa Carlos, Ghanem Bernard · 2026-07-22 · 3 min AI

用于教育视频合成的 SGA: Plug&Play 几何验证

SGA: Plug&Play Geometric Verification for Educational Video Synthesis

最近的工作利用大型语言模型 (LLMs) 使用 Manim 等库生成教学动画的可执行代码。然而,随之而来...

Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim. However, ensu...

02
Brian K Chen · 2026-07-22 · 10 min AI

压力下的逻辑判断: 使用习得的软前缀诊断三段论稳定性

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

为了测试正确的逻辑判断如何响应学习的上下文,,我们在精确标记的三段论推理基准之前添加一个软前缀,同时......

To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while...

03
Yukuan Lu, Zaishuo Xia, Weyl Lu, Yubei Chen · 2026-07-20 · 4 min AI

改进 Atari Pong 中强代理背后的弱世界模型

Improving Weak World Models Behind Strong Agents in Atari Pong

强世界模型代理经常包含弱世界模型。我们通过在 A 中再现五个视觉世界模型代理来研究这种代理世界模型差距。

Strong world-model agents frequently contain weak world models. We study this agent-world-model gap by reproducing five visual world-model agents in A...

05
Emmanuel Jeannot · 2026-07-20 · 7 min AI

人工智能驱动科学研究 ; 的产业化及其后果

The Industrialization of Research ; On AI-Driven Science and Its Consequences

人工智能正在改变科学研究——不仅作为一种更强大的工具,,而且作为研究的自主参与者......

Artificial intelligence is transforming scientific research -- not merely as a more powerful instrument, but as an autonomous participant in the resea...

06
Goktug Ozkan · 2026-07-20 · 8 min AI

MedFailBench: 临床医生构建的医疗 AI 安全边界检查开源基准

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

大多数医疗人工智能基准测试都会衡量模型是否知道正确答案。 MedFailBench 提出了一个不同的问题: 哪个安全边界失败了? 我们...

Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We ...

07
Patrick Phuoc Do, Chau M. Ta, Chaoli Wang · 2026-07-20 · 10 min AI

科学可视化素养多模态大语言模型的基准测试

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

多模态大语言模型 (MLLMs) 越来越多地用于解释可视化,,但当前的评估仍然主要以图表为中心,并且...

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and p...

08
Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang · 2026-07-20 · 4 min AI

我们可以相信人工智能评估的项目反应理论?

Can We Trust Item Response Theory for AI Evaluation?

AI 基准越来越多地利用项目级统计模型, 特别是项目响应理论 (IRT), 来估计模型能力, 排名系统...

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank syste...

09
Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, Dongyu Liu · 2026-07-19 · 10 min AI

Plover: 通过以计划为中心的交互引导 GUI 代理

Plover: Steering GUI Agents through Plan-Centric Interaction

图形用户界面 (GUI) 自动化在现实环境中仍然具有挑战性,,其中动态布局, 意外对话框, 和不断发展的交互...

Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving inte...

10