个人代理正在成为持久的用户拥有的中介:他们记住偏好,过滤平台中介的信息,使用工具,并与服务协商。现有基准评估工具使用,网络导航,桌面控制,个性化,推荐,和不断变化的环境,,但很少询问代理是否保留用户主权:在尊重隐私,同意,证据,用户负担,和对操纵激励的抵制的同时促进用户'的当前利益。我们引入了 SovereignPA-Bench, 一个可执行基准,用于在不断变化的意图, 平台调解, 隐私边界, 同意约束, 证据要求, 和负担权衡下评估用户拥有的个人代理。该基准将代理可见的 ObservableState 与仅评估者的 HiddenLabels, 报告任务成功, 对齐, 隐私, 同意, 证据, 操作, 负担, 和可审计性, 的组件指标分开,并保留模型和策略比较的配对场景排序。我们评估了 4 个模型系列和 8 个政策基线, 的 120 个主权压力情景,产生 3,840 个带有原始提示的冻结提示轨迹, 输出, 提供者形式响应, 解析操作, 可重新计算指标, 硬集分析, 定性案例, 以及超过 240 个项目的盲态 3 注释者审计。完全主权脚手架相对于直接,仅内存,仅同意,仅证据, ReAct/工具使用,安全提示,和法官保护基线提高了主权得分,同时减少隐私泄漏,同意违规,过度让步,和操纵捕获。人工审计显示出对隐私和同意的高度认可,而对操纵, 的认可度较低,确定了平台说服判断的主观边界。这些结果表明,个人代理评估必须超越任务完成,转向具有代表性的,同意意识,基于证据的行动。
Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negotiate with services. Existing benchmarks evaluate tool use, web navigation, desktop control, personalization, recommendation, and evolving context, but rarely ask whether an agent preserves user sovereignty: advancing the user的 current interests while respecting privacy, consent, evidence, user burden, and resistance to manipulative incentives. We introduce SovereignPA-Bench, an executable benchmark for evaluating user-owned personal agents under evolving intent, platform mediation, privacy boundaries, consent constraints, evidence requirements, and burden tradeoffs. The benchmark separates agent-visible ObservableState from evaluator-only HiddenLabels, reports component metrics for task success, alignment, privacy, consent, evidence, manipulation, burden, and auditability, and preserves paired scenario ordering for model and policy comparisons. We evaluate 120 sovereignty stress scenarios across 4 model families and 8 policy baselines, yielding 3,840 frozen-prompt trajectories with raw prompts, outputs, provider-form responses, parsed actions, recomputable metrics, hard-set analyses, qualitative cases, and a blinded 3-annotator audit over 240 items. Full-sovereign scaffolding improves sovereignty score over direct, memory-only, consent-only, evidence-only, ReAct/tool-use, safety-prompt, and judge-guard baselines while reducing privacy leakage, consent violation, over-concession, and manipulation capture. Human audit shows high agreement on privacy and consent and lower agreement on manipulation, identifying the subjective frontier of platform-persuasion judgments. These results show that personal-agent evaluation must move beyond task completion toward representative, consent-aware, evidence-grounded action.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)