Thomas Winninger · 2026-07-06 · 7 min AI

通过 RFM-AGOP 的快速多维拒绝子空间

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

大型语言模型 (LLMs) 中的引导和监控激活越来越多地用于安全性和可解释性。早期的工作假设是...

Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed beh...

01
Xianhui Meng, Zirui Song, Yuchen Zhang, Li Zhang, Yongxuan Lv, Xiuying Chen, Kun Wang, Yan Luo, Kai Chen, Hangjun Ye, Long Chen, Jun Liu, Xiaoshuai Hao · 2026-07-06 · 3 min AI

非曼哈顿环境中文本驱动的 3D 室内场景合成

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

大型语言模型 (LLMs) 在曼哈顿环境的 3D 室内合成中表现出了卓越的能力。然而,现有方法...

Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods of...

02
Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J.Rodriguez-Martinez, Lorena Otero-Cerdeira · 2026-07-05 · 5 min AI

使用大型语言模型对 Linux/bash 考试进行自动评分: 四级认知分类方法

Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

命令行考试的可扩展且可靠的评分仍然是计算教育, 中的一个挑战,其中入学人数的增加使得手动评分变得困难...

Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking dif...

03
Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li, Qingyu Yin, Yulun Wu, Zhennan Shen, Tong Zhu, Yanshu Li, Guanjie Chen, Derek F. Wong, Yafu Li, Yu Cheng, Yang Yang · 2026-07-05 · 3 min AI

EvoPolicyGym: 评估交互式环境中的自主策略演化

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

人们越来越期望自主代理通过反馈来改进可执行政策,,但现有的评估经常将这一过程分解为……

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into...

04
Timo Bertram, Sidhant Bhavnani, Richard Freinschlag, Erich Kobler, Andreas Mayr, Günter Klambauer · 2026-07-05 · 9 min AI

G-RRM: 使用递归推理模型指导符号求解器

G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Models

在这项工作, 中,我们重点关注 SE-RRM,,这是 RRM 的符号等变实例,它对更大的问题规模表现出改进的外推能力。我们建议...

In this work, we focus on SE-RRMs, a symbol-equivariant instantiation of RRMs that exhibits improved extrapolation to larger problem sizes. We propose...

05
Arman Ghaffarizadeh, Danyal Mohaddes, Aliakbar Izadkhah, Shahriar Noroozizadeh · 2026-07-05 · 6 min AI

当没有人在看时,LLM 代理人会说什么: 社会结构和多代理人辩论中潜在客观出现

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

LLM 代理人将越来越多地在社会结构环境中行动,其中角色, 受众, 和关系背景可以决定什么是有利的或昂贵的......

LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly...

06
Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen, Ismini Lourentzou, Hanghang Tong, Jingrui He · 2026-07-05 · 5 min AI

ReContext: 递归证据重放作为 LLM 工具进行长上下文推理

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

对长上下文的理解和推理已成为在实际应用程序中部署大型语言模型 (LLMs) 的关键要求。阿尔斯...

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Alth...

07
Mona Schirmer, Metod Jazbec, Alexander Timans, Christian Naesseth, Maja Waldron, Eric Nalisnick · 2026-07-04 · 10 min AI

法学硕士在线安全监控

Online Safety Monitoring for LLMs

尽管进行了对齐培训,, LLM 仍然容易在部署时生成不安全的输出。在线监控输出并在安全时发出警报...

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when saf...

08
Josh Hills, Ida Caspary, Asa Cooper Stickland · 2026-07-04 · 3 min AI

持久状态人工智能控制中的分布式攻击

Distributed Attacks in Persistent-State AI Control

随着人工智能编码代理变得更加自主,,他们越来越多地迭代地交付代码,,并且代码库在会话之间持续存在。这种坚持让...

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence cr...

09
Michael Saldivar, Ben Slivinski · 2026-07-03 · 6 min AI

Theoria: 对非正式推理状态的重写可接受性验证

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

人工智能系统的 的答案何时应该可信? 形式证明助手提供确定性,但无法达到大部分问题分布; 标量法学硕士...

When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM ju...

10