Wanyun Cui
·
2026-07-07
·
10 min
AI
线性注意力的海马体: 循环状态忘记的精确记忆
A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets
线性注意力和状态空间语言模型将前缀压缩为固定大小的循环状态,,以有损前的代价产生 O(1) 内存...
Linear-attention and state-space language models compress the prefix into a fixed-size recurrent state, yielding O(1) memory at the cost of a lossy ex...
01
Haonan Huang
·
2026-07-07
·
3 min
AI
扎根的自主研究: 前沿计算物理中从语料库到手稿的容错法学硕士管道
Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics
自主研究代理已经在机器学习沙箱中展示了端到端的法学硕士自动化,其中执行提供了校准。前沿PH...
Autonomous-research agents have demonstrated end-to-end LLM automation in machine-learning sandboxes where execution provides calibration. Frontier ph...
02
Xi Fang, Weijie Xu, Yingqiang Ge, Yuhui Xu, Stephanie Eckman, Chandan K. Reddy
·
2026-07-07
·
5 min
AI
DRIFTLENS: 测量个性化语言模型中记忆引起的推理漂移
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
个性化改变了模型对用户所说的内容; 我们表明,它还可以改变用于证明响应合理性的推理轨迹。现代法学硕士...
Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLM...
03
Uwe M. Borghoff, Paolo Bottoni, Remo Pareschi
·
2026-07-07
·
3 min
AI
安全关键型实时自治系统的硬件强制语义协调
Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems
代理人工智能的最新进展正在产生日益复杂的自主系统,这些系统集成了大型语言模型,世界模型,优化e...
Recent advances in agentic AI are producing increasingly complex autonomous systems that integrate large language models, world models, optimization e...
04
Amanda Diehl | MIT Schwarzman College of Computing
·
2026-07-07
·
3 min
AI
迈向一个让所有人都能受益于神经技术的未来
Toward a future that preserves benefits of neurotechnology for all
随着先进医疗技术越来越接近消费市场,,对受保护使用的护栏的需求应该会增加。可能会开始什么...
As advanced medical technology gets closer to hitting consumer markets, the need for guardrails on protected usage should increase. What might begin a...
05
Thomas Winninger
·
2026-07-06
·
10 min
AI
通过约束的可操纵性: 是编码代理可扩展监督的基础
Steerability via constraints: a substrate for scalable oversight of coding agents
编码代理有能力; 人类监督是瓶颈。不受约束的代理会带来安全风险, 侵蚀代码库可扩展性, 并使人...
Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, and make huma...
06
Thomas Winninger
·
2026-07-06
·
7 min
AI
通过 RFM-AGOP 的快速多维拒绝子空间
Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
大型语言模型 (LLMs) 中的引导和监控激活越来越多地用于安全性和可解释性。早期的工作假设是...
Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed beh...
07
Xianhui Meng, Zirui Song, Yuchen Zhang, Li Zhang, Yongxuan Lv, Xiuying Chen, Kun Wang, Yan Luo, Kai Chen, Hangjun Ye, Long Chen, Jun Liu, Xiaoshuai Hao
·
2026-07-06
·
3 min
AI
非曼哈顿环境中文本驱动的 3D 室内场景合成
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
大型语言模型 (LLMs) 在曼哈顿环境的 3D 室内合成中表现出了卓越的能力。然而,现有方法...
Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods of...
08
Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J.Rodriguez-Martinez, Lorena Otero-Cerdeira
·
2026-07-05
·
5 min
AI
使用大型语言模型对 Linux/bash 考试进行自动评分: 四级认知分类方法
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach
命令行考试的可扩展且可靠的评分仍然是计算教育, 中的一个挑战,其中入学人数的增加使得手动评分变得困难...
Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking dif...
09
Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li, Qingyu Yin, Yulun Wu, Zhennan Shen, Tong Zhu, Yanshu Li, Guanjie Chen, Derek F. Wong, Yafu Li, Yu Cheng, Yang Yang
·
2026-07-05
·
3 min
AI
EvoPolicyGym: 评估交互式环境中的自主策略演化
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
人们越来越期望自主代理通过反馈来改进可执行政策,,但现有的评估经常将这一过程分解为……
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into...
10