诸如技能文件, 内存文件, 和行为配置文件之类的文本文件在定义现代代理的行为方式方面发挥着核心作用。通过人类或代理本身的编辑,,这些文件可能会随着时间的推移而演变, 直接指导代理 在未来交互中的行为。我们提出了一种通过将特征定义为文本嵌入模型的嵌入空间中的方向来测量代理$traits$的方法和框架。我们在标记的 "before" 与 "after" 技能文件差异上训练线性模型,以学习特征向量,,然后通过将其嵌入差异投影到该向量上来对任意技能编辑进行评分。在 68 个标记的技能差异对上评估寻求敏感数据, 的倾向特征,我们的方法在留一交叉验证下实现了 91.2% 的符号分类准确率和 $\rho = 0.82$ 的斯皮尔曼等级相关性。我们将此特征评估构建到更广泛的代理到代理协议中,该协议使一个代理能够通过可信中介评估另一个的 技能文件更新。

Text files such as skill files, memory files, and behavioral configuration files play a central role in defining how modern agents act. Through edits by humans or the agents themselves, these files may evolve over time, directly steering the agent的 behavior in future interactions. We present a methodology and framework for measuring agent $traits$ by defining traits as directions in the embedding space of a text embedding model. We train a linear model on labeled "before" versus "after" skill file diffs to learn a trait vector, then score arbitrary skill edits by projecting their embedding diffs onto this vector. Evaluated on 68 labeled skill diff pairs for the trait of propensity to seek sensitive data, our method achieves 91.2% sign classification accuracy and a Spearman rank correlation of $\rho = 0.82$ under leave-one-out cross-validation. We build this trait evaluation into a broader agent-to-agent protocol that enables one agent to evaluate another的 skill file updates through a trusted intermediary.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)