LLM代理人将越来越多地在社会结构环境中行动,其中角色,受众,和关系背景可以决定什么是有利的或成本高昂的。我们研究在提示, 中没有任何明确目标的这种社会结构, 是否会改变代理人相对于相同条件下引发的非记录(OTR) 渠道公开表达的内容。我们引入了一个双通道辩论框架,其中代理产生公开言论,这些言论与被记录但从未向其他参与者显示的 OTR 响应一起进入共享历史。在 10 个模型, 3 个场景, 中,以及每个场景, 对齐诱导设置中的 5 个变化,在目标代理, 中产生系统性公共 OTR 分歧,其决策分歧从 $\sim$3% 基线上升到大约 40%。四项汇总分析:立场,语义相似性,自然语言推理,和调查回复的效果是一致的。在某些情况下,,OTR 回应明确将公共便利归因于关系压力,,例如职业风险或赞助义务。研究结果表明,代理评估应该超越明确的目标并检测紧急目标。我们提出了一个双渠道评估框架和补充行为措施来实施这一评估。
LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly to say. We study whether such social structure, without any explicit objective in the prompt, changes what an agent expresses publicly relative to an off-the-record (OTR) channel elicited under the same condition. We introduce a dual-channel debate framework in which agents produce public utterances that enter the shared history alongside OTR responses that are recorded but never shown to the other participant. Across 10 models, 3 scenarios, and 5 variations within each scenario, alignment-inducing settings produce systematic public-OTR divergence in the targeted agent, with its decision divergence rising from a $\sim$3% baseline to roughly 40%. The effect is consistent across four aggregate analyses: stance, semantic similarity, natural language inference, and survey responses. In some cases, the OTR response explicitly attributes public accommodation to relational pressures, such as career risk or sponsorship obligation. The findings suggest that agent evaluation should extend beyond explicit goals and detect emergent objectives. We present a dual-channel evaluation framework and complementary behavioral measures that operationalize this assessment.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 机器学习 (cs.LG); 多代理系统 (cs.MA)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA)