即使使用温度 $T=0$, 大型语言模型 (LLMs) 进行解码,也可以为相同的输入产生不同的输出。 Thinking Machines Lab 最近的工作强调了非确定性, 的实现级来源,包括批量大小变化, 内核非不变性, 和浮点非关联性。在这篇简短的说明中,我们通过引入\emph{背景温度}$T_{\mathrm{bg}}$,的概念来形式化这种行为,即使在标称$T=0$时,也观察到由依赖于实现的扰动过程引起的有效温度。我们提供清晰的定义,显示$T_{\mathrm{bg}}$如何与推理环境$I$,控制的随机扰动相关,并提出一个经验协议通过理想参考的等效温度$T_n(I)$来估计$T_{bg}$系统。最后,我们在主要 LLM 提供商的代表性池上运行了一组试点实验,展示了这一想法并概述了对可重复性, 评估, 和部署的影响。
Even when decoding with temperature $T=0$, large language models (LLMs) can produce divergent outputs for identical inputs. Recent work by Thinking Machines Lab highlights implementation-level sources of nondeterminism, including batch-size variation, kernel non-invariance, and floating-point non-associativity. In this short note we formalize this behavior by introducing the notion of \emph{background temperature} $T_{\mathrm{bg}}$, the effective temperature induced by an implementation-dependent perturbation process observed even when nominal $T=0$. We provide clean definitions, show how $T_{\mathrm{bg}}$ relates to a stochastic perturbation governed by the inference environment $I$, and propose an empirical protocol to estimate $T_{bg}$ via the equivalent temperature $T_n(I)$ of an ideal reference system. We conclude with a set of pilot experiments run on a representative pool from the major LLM providers that demonstrate the idea and outline implications for reproducibility, evaluation, and deployment.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 机器学习 (cs.LG)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)