语言代理通过重用\emph{skills}(从过去的经验中提取的结构化程序工件)不断改进。特别是, \emph{域级}和\emph{模型生成的}技能特别有前途。它们通过编码特定于域的重复过程, 在域内提供快速适应,并且它们的规模超越了劳动密集型手工制作。然而,虽然提取方法继续激增,理解仍然有限,没有跨越整个技能生命周期的全面研究 - \textbf{经验生成}, \textbf{技能提取},和\textbf{技能消耗} - 询问这些技能在工作时是否真正有效,以及是什么让它们成功或失败。为了缩小这一差距,,我们构建了一个基于实用程序的评估框架,该框架提供跨提取器和目标代理, 的系统实验结果,涵盖五个不同的代理任务领域。我们发现模型生成的技能平均而言是有益的,但表现出不平凡的负迁移,,并且提取器和目标的行为都不一致。模型可以是强提取器,但可以是弱消费者,,反之亦然,,其技能实用性与模型规模或基线任务强度无关。为了解释这些模式,,我们随后深入剖析每个生命周期阶段, 分析经验构成如何塑造技能质量, 有用技能的特征, 哪些属性以及相同的技能如何在不同消费者之间转移。最后,,我们将这些发现转化为具体的\emph{meta-skill},指导技能提取到与实际效用相关的特征,,从而持续提高跨领域的技能质量并大幅减少负迁移。
Language agents increasingly improve by reusing \emph{skills} -- structured procedural artifacts distilled from past experience. In particular, \emph{domain-level} and \emph{model-generated} skills are especially promising. They offer fast adaptation within a domain by encoding domain-specific recurring procedures, and they scale beyond labor-intensive hand-crafting. However, while extraction methods continue to proliferate, understanding remains limited, with no comprehensive study spanning the full skill lifecycle -- \textbf{experience generation}, \textbf{skill extraction}, and \textbf{skill consumption} -- to ask whether such skills actually work, when they work, and what makes them succeed or fail. To close this gap, we build a utility-grounded evaluation framework that provides systematic experimental results across extractors and target agents, covering five diverse agentic task domains. We find that model-generated skills are beneficial on average but exhibit non-trivial negative transfer, and that neither extractors nor targets behave uniformly. A model can be a strong extractor yet a weak consumer, or vice versa, with skill utility independent of model scale or baseline task strength. To explain these patterns, we then dissect each lifecycle stage in depth, analyzing how experience composition shapes skill quality, what properties characterize useful skills, and how the same skill transfers across different consumers. Finally, we translate these findings into a concrete \emph{meta-skill} that guides skill extraction toward the features tied to actual utility, which consistently improves skill quality across domains and substantially reduces negative transfer.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)