显式技能库使使用计算机的代理更容易检查,,但目前尚不清楚是否可以从交互数据中挖掘此类库以改进下游策略。我们通过一个三阶段管道来研究这个问题,该管道将 GUI 轨迹, 集群片段分段为候选技能,,并根据结果注释训练技能感知策略。挖掘的集群在源基准上是可读的%3八个集群中有五个相对于 InteraSkill 工作流标签的纯度至少为 0.95。然而,可读性并不意味着转移。 GRPO 仅将 IW 技能步骤准确性从 18.5\% 提高到 20.5\%,,而 BrowseComp+ 基本保持不变,,并且在关键源域指标上表现不佳。因此,我们将该方法作为诊断研究:轨迹挖掘可以揭示可检查的技能结构,,但当前的边界检测器,无序段表示,和离线奖励模型不足以实现可靠的跨域策略改进。

Explicit skill libraries make computer-using agents easier to inspect, but it remains unclear whether such libraries can be mined from interaction data in a way that improves downstream policies. We study this question through a three-stage pipeline that segments GUI trajectories, clusters segments into candidate skills, and trains a skill-aware policy from the resulting annotations. The mined clusters are readable on the source benchmark: five of eight clusters have at least 0.95 purity against InteraSkill Workflows labels. However, readability does not imply transfer. GRPO improves IW skill-step accuracy only from 18.5\% to 20.5\%, leaves BrowseComp+ essentially unchanged, and underperforms trivial frequency priors on key source-domain metrics. We therefore present the method as a diagnostic study: trajectory mining can expose inspectable skill structure, but the current boundary detector, orderless segment representation, and offline reward model are insufficient for reliable cross-domain policy improvement.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)