在策略蒸馏 (OPD) 通过使用密集的令牌级信号监督学生采样轨迹,提供卓越的容量传输。为了提供高质量的监督来源,从而提升蒸馏,的性能前沿,一个直观的方向是向教师或学生本身注入特权信息。然而,这种额外的输入会导致一种潜在的失败模式,我们称之为特权错觉:,这种模式将学生本应弥补的可转移能力差距,与只能模仿但永远无法复制的信息不对称差距混为一谈。代币级别监管固有的不一致性, 进一步放大了这个问题,其中只有一小部分代币携带关键的能力承载信号。为此,,我们提出 DOPD, 一种优势感知的双重蒸馏范式,该范式根据优势差距和相对概率在特权教师和特权学生策略之间动态路由令牌级监督。每个代币都接受来自教师或学生本身,的不同强度,目标,和策略的监督,传递可信能力,同时接收辅助信号,以减轻特权错觉。对大型语言模型 (LLM) 和视觉语言模型 (VLM) 设置的大量实验表明,DOPD 始终优于 Vanilla OPD 和其他同行。稳定性,鲁棒性,持续学习,和分布外任务的进一步结果验证了其优越性。
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure mode we dub privilege illusion: a pattern that conflates the transferable capability gap that students are meant to close, and the information asymmetry gap that can only be mimicked but never replicated. This issue is further amplified by the inherent non-uniformity of token-level supervision, where only a small subset of tokens carries pivotal capability-bearing signals. To this end, we propose DOPD, an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between privileged teacher and privileged student policies based on their advantage gap and relative probabilities. Each token receives supervision of different strength, objective, and strategy from either teacher or student itself, which transfers credible capability while simultaneously receiving auxiliary signals, to alleviate privilege illusion. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that DOPD consistently outperforms Vanilla OPD and other counterparts. Further results on stability, robustness, continual learning, and out-of-distribution tasks validate its superiority.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)