构建社交校准的大型语言模型,,该模型可以向他人学习,而不是简单地屈服于他们,,需要的不仅仅是减少作为一维故障模式的阿谀奉承。模型必须区分何时采纳他人的观点以及何时保持有根据的道德判断。我们研究了管理这种区别的更广泛的阻力-遵从过程。在三项研究,中,我们表明模型'判断修正是沿着与人类社会心理学中的经典现象平行的三个维度构建的:传入视图与模型'初始位置之间的距离,该视图的来源属性,以及支持它的联盟结构。模型通常更容易接受附近的立场,,更容易受到作为自己先前判断的观点的影响,,并且对群体压力的反应不同。这些发现将阿谀奉承重新塑造为由社会影响力塑造的更广泛的判断更新过程的一种表达。我们的框架为区分建设性信念修正与阿谀奉承, 提供了原则基础,从而支持道德后果互动中更好的协调。

Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models judgment revision is structured along three dimensions that parallel classic phenomena in human social psychology: the distance between an incoming view and the model的 initial position, the source attribution of that view, and the coalition structure supporting it. Models are generally more receptive to nearby positions, more influenced by views presented as their own prior judgments, and differently responsive to group pressure. These findings recast sycophancy as one expression of a broader judgment-updating process shaped by social influence. Our framework provides a principled basis for distinguishing constructive belief revision from sycophantic compliance, thereby supporting better alignment in morally consequential interactions.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)