Qijia He, Jiayi Cheng, Chenqian Le, Rui Wang, Xunmei Liu, Yixian Chen, Jie Mei, Zhihao Wang, Xupeng Chen, Yuhuan Chen, Tao Wang
·
2026-07-22
·
7 min
AI
CodeRescue: 编码代理的预算校准恢复路由
CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
编码代理越来越多地在可执行环境中运行,其中失败的尝试会产生可操作的反馈,而不仅仅是错误的答案......
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answ...
01
Lopez Jhon, Hinojosa Carlos, Ghanem Bernard
·
2026-07-22
·
3 min
AI
用于教育视频合成的 SGA: Plug&Play 几何验证
SGA: Plug&Play Geometric Verification for Educational Video Synthesis
最近的工作利用大型语言模型 (LLMs) 使用 Manim 等库生成教学动画的可执行代码。然而,随之而来...
Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim. However, ensu...
02
Brian K Chen
·
2026-07-22
·
10 min
AI
压力下的逻辑判断: 使用习得的软前缀诊断三段论稳定性
Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
为了测试正确的逻辑判断如何响应学习的上下文,,我们在精确标记的三段论推理基准之前添加一个软前缀,同时......
To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while...
03
Adam Zewe | MIT News
·
2026-07-21
·
8 min
AI
新方法旨在保护孩子免受非法人工智能的侵害
New method aims to keep kids safe from illegal AI
随着生成式人工智能的爆炸性普及,,许多开源模型现在可以在线供任何人适应他们的任务...
With the exploding popularity of generative artificial intelligence, many open-source models are now available online for anyone to adapt for their ta...
04
Yukuan Lu, Zaishuo Xia, Weyl Lu, Yubei Chen
·
2026-07-20
·
4 min
AI
改进 Atari Pong 中强代理背后的弱世界模型
Improving Weak World Models Behind Strong Agents in Atari Pong
强世界模型代理经常包含弱世界模型。我们通过在 A 中再现五个视觉世界模型代理来研究这种代理世界模型差距。
Strong world-model agents frequently contain weak world models. We study this agent-world-model gap by reproducing five visual world-model agents in A...
05
Emmanuel Jeannot
·
2026-07-20
·
7 min
AI
人工智能驱动科学研究 ; 的产业化及其后果
The Industrialization of Research ; On AI-Driven Science and Its Consequences
人工智能正在改变科学研究——不仅作为一种更强大的工具,,而且作为研究的自主参与者......
Artificial intelligence is transforming scientific research -- not merely as a more powerful instrument, but as an autonomous participant in the resea...
06
Goktug Ozkan
·
2026-07-20
·
8 min
AI
MedFailBench: 临床医生构建的医疗 AI 安全边界检查开源基准
MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection
大多数医疗人工智能基准测试都会衡量模型是否知道正确答案。 MedFailBench 提出了一个不同的问题: 哪个安全边界失败了? 我们...
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We ...
07
Patrick Phuoc Do, Chau M. Ta, Chaoli Wang
·
2026-07-20
·
10 min
AI
科学可视化素养多模态大语言模型的基准测试
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
多模态大语言模型 (MLLMs) 越来越多地用于解释可视化,,但当前的评估仍然主要以图表为中心,并且...
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and p...
08
Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang
·
2026-07-20
·
4 min
AI
我们可以相信人工智能评估的项目反应理论?
Can We Trust Item Response Theory for AI Evaluation?
AI 基准越来越多地利用项目级统计模型, 特别是项目响应理论 (IRT), 来估计模型能力, 排名系统...
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank syste...
09
Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, Dongyu Liu
·
2026-07-19
·
10 min
AI
Plover: 通过以计划为中心的交互引导 GUI 代理
Plover: Steering GUI Agents through Plan-Centric Interaction
图形用户界面 (GUI) 自动化在现实环境中仍然具有挑战性,,其中动态布局, 意外对话框, 和不断发展的交互...
Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving inte...
10