Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li, Qingyu Yin, Yulun Wu, Zhennan Shen, Tong Zhu, Yanshu Li, Guanjie Chen, Derek F. Wong, Yafu Li, Yu Cheng, Yang Yang · 2026-07-05 · 3 min AI

EvoPolicyGym: 评估交互式环境中的自主策略演化

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

人们越来越期望自主代理通过反馈来改进可执行政策,,但现有的评估经常将这一过程分解为……

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into...

01
Timo Bertram, Sidhant Bhavnani, Richard Freinschlag, Erich Kobler, Andreas Mayr, Günter Klambauer · 2026-07-05 · 9 min AI

G-RRM: 使用递归推理模型指导符号求解器

G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Models

在这项工作, 中,我们重点关注 SE-RRM,,这是 RRM 的符号等变实例,它对更大的问题规模表现出改进的外推能力。我们建议...

In this work, we focus on SE-RRMs, a symbol-equivariant instantiation of RRMs that exhibits improved extrapolation to larger problem sizes. We propose...

02
Arman Ghaffarizadeh, Danyal Mohaddes, Aliakbar Izadkhah, Shahriar Noroozizadeh · 2026-07-05 · 6 min AI

当没有人在看时,LLM 代理人会说什么: 社会结构和多代理人辩论中潜在客观出现

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

LLM 代理人将越来越多地在社会结构环境中行动,其中角色, 受众, 和关系背景可以决定什么是有利的或昂贵的......

LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly...

03
Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen, Ismini Lourentzou, Hanghang Tong, Jingrui He · 2026-07-05 · 5 min AI

ReContext: 递归证据重放作为 LLM 工具进行长上下文推理

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

对长上下文的理解和推理已成为在实际应用程序中部署大型语言模型 (LLMs) 的关键要求。阿尔斯...

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Alth...

04
yxc0433 · 2026-07-04 · 8 min US-China Trade

Kling AI获得腾讯融资后,快手股价上涨

Kuaishou shares jump after securing Tencent funding for Kling AI

Kuaishou Technology shares rose nearly 7% Friday before trimming gains, after the company announced a capital injection of nearly $2.8 billion into it...

Kuaishou Technology shares rose nearly 7% Friday before trimming gains, after the company announced a capital injection of nearly $2.8 billion into it...

05
Mona Schirmer, Metod Jazbec, Alexander Timans, Christian Naesseth, Maja Waldron, Eric Nalisnick · 2026-07-04 · 10 min AI

法学硕士在线安全监控

Online Safety Monitoring for LLMs

尽管进行了对齐培训,, LLM 仍然容易在部署时生成不安全的输出。在线监控输出并在安全时发出警报...

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when saf...

06
Josh Hills, Ida Caspary, Asa Cooper Stickland · 2026-07-04 · 3 min AI

持久状态人工智能控制中的分布式攻击

Distributed Attacks in Persistent-State AI Control

随着人工智能编码代理变得更加自主,,他们越来越多地迭代地交付代码,,并且代码库在会话之间持续存在。这种坚持让...

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence cr...

07
yxc0433 · 2026-07-03 · 4 min US-China Trade

阿里巴巴

Alibaba

北京— 阿里巴巴旗下蚂蚁集团正在加紧进军人形机器人领域。蚂蚁金服领投哈尔滨5亿元($73.58万元)轮融资

BEIJING — Alibaba-affiliate Ant Group is ramping up its move into humanoid robots. Ant has led a 500 million yuan ($73.58 million) funding round in hu...

08
Michael Saldivar, Ben Slivinski · 2026-07-03 · 6 min AI

Theoria: 对非正式推理状态的重写可接受性验证

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

人工智能系统的 的答案何时应该可信? 形式证明助手提供确定性,但无法达到大部分问题分布; 标量法学硕士...

When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM ju...

09
Shengguang Wu, Hao Zhu, Yuhui Zhang, Xiaohan Wang, Serena Yeung-Levy · 2026-07-03 · 10 min AI

AutoMem: 自动学习记忆作为一种认知技能

AutoMem: Automated Learning of Memory as a Cognitive Skill

记忆专业知识是一种习得的技能:知道要编码什么,何时检索,以及如何组织知识——这种能力在认知科学中被称为……

Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capacity known in cognitive science as...

10