Goktug Ozkan · 2026-07-20 · 8 min AI

MedFailBench: 临床医生构建的医疗 AI 安全边界检查开源基准

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

大多数医疗人工智能基准测试都会衡量模型是否知道正确答案。 MedFailBench 提出了一个不同的问题: 哪个安全边界失败了? 我们...

Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We ...

01
Patrick Phuoc Do, Chau M. Ta, Chaoli Wang · 2026-07-20 · 10 min AI

科学可视化素养多模态大语言模型的基准测试

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

多模态大语言模型 (MLLMs) 越来越多地用于解释可视化,,但当前的评估仍然主要以图表为中心,并且...

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and p...

02
Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang · 2026-07-20 · 4 min AI

我们可以相信人工智能评估的项目反应理论?

Can We Trust Item Response Theory for AI Evaluation?

AI 基准越来越多地利用项目级统计模型, 特别是项目响应理论 (IRT), 来估计模型能力, 排名系统...

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank syste...

03
yxc0433 · 2026-07-19 · 8 min US-China Trade

由于投资下滑,中国季度GDP增速为2022年以来最慢

China posts slowest quarterly GDP growth since 2022 as investment slumps

中国第二季度经济增速为 2022 年第四季度以来最弱, 强化了政策刺激以加速经济增长的呼声……

China's economy in the second quarter expanded at its weakest pace since the fourth quarter of 2022, reinforcing calls for policy stimulus as an accel...

04
yxc0433 · 2026-07-19 · 10 min US-China Trade

中国汽车制造商在英国扩张—,许多英国人正在拥抱它们

Chinese automakers expand in UK — and many Brits are embracing them

MAIDSTONE, 英格兰 — Izzy Woodrow 是一名信徒。四个星期前,,他加入了少数但不断增加的购买中国制造汽车的英国人行列......

MAIDSTONE, England — Izzy Woodrow is a believer. Four weeks ago, he joined the small but growing number of Brits who have bought a Chinese-made vehicl...

05
Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, Dongyu Liu · 2026-07-19 · 10 min AI

Plover: 通过以计划为中心的交互引导 GUI 代理

Plover: Steering GUI Agents through Plan-Centric Interaction

图形用户界面 (GUI) 自动化在现实环境中仍然具有挑战性,,其中动态布局, 意外对话框, 和不断发展的交互...

Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving inte...

06
Hoang-Loc Cao, Van Pham, Truong Thanh Hung Nguyen, Phuc Truong Loc Nguyen, Phuc Ho, Veronica Whitford, Hung Cao · 2026-07-19 · 8 min AI

用于可解释抑郁症症状注释的自我进化的以人为中心的框架

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

注释质量是为心理健康研究构建可靠且可解释的人工智能 (XAI) 系统的主要瓶颈。在深度...

Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health research. In dep...

07
Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu · 2026-07-19 · 3 min AI

当言语安全但行动致命: 在隐藏状态风险空间中探索超越文本越狱的物理越狱

When Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space

大型语言模型 (LLMs) 越来越多地充当具体代理的高级规划器,,其中语言上良性的指令可能变得不安全......

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe...

08
Moein Taherinezhad, Sebastian Maier, Gerardo Vitagliano, Francesco Pierri, Stefan Feuerriegel · 2026-07-19 · 10 min AI

AutoSynthesis: 用于自动荟萃分析的代理系统

AutoSynthesis: An agentic system for automated meta-analysis

证据综合对于将初级研究转化为科学, 医学, 教育, 和政策的可靠知识至关重要。然而,定量评估...

Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quantitative evi...

09
Qiwei Li, Jorge Ortiz · 2026-07-19 · 8 min AI

告诉我为什么 (Ain't 除了堵塞): 对城市驾驶数据的探索性因果分析

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

交通机构现在可以访问大量视频数据来研究安全和拥堵情况。这些数据大部分是观察性的并且...

Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of these data are observational and c...

10