训练中期,是预训练和对齐,之间的阶段,模型'的每个域数据组成通常由数据可用性而不是原则设计来设置。我们询问这个决定会带来什么, 以及稍后的对齐过程是否可以撤销它。在受控逻辑推理设置 (Qwen3-8B-Base, 中,具有 4B 复制; 五个语义上规则不相交的 KOR-Bench 域),我们训练跨越五域单纯形, 的 30 个分配,24 个扫描配置加上从 fit, 中保留的 6 个分配,每个分配有 5 个种子。出现了三个发现。首先,每个域都有一个内部覆盖最佳:中等频带($10\%$-$40\%$)对于所有五个域来说都是最好的,并且二次内部性的校准排列测试给出$P\approx0.010$;仅拟合中间训练曲线,具有8B峰值$9.9\%$ 和 $35.1\%$, 之间再现曲线形状,但不再现峰值位置。第二, 间隙在固定预算对齐过程中幸存下来: 补偿性 SFT 提高了 116/120 个细胞 (mean $+4.32\%$) 但在 $5\%$ 阈值处桥接 $0/240$ 对,在阈值处桥接 $30/240$ $10\%$比率,等预算统一对照的行为几乎与,相同,并且排列零将桥接$13.8\pm3.3$和$77.9\pm8.5$对($P<0.001$)。第三, 零覆盖率在仅训练中崩溃了,,尽管仅 FineWeb-Edu 控制显示崩溃与通用漂移混合在一起。探索性 $\theta^*$ 分配获得了最大的全管道增益 ($+4.36\%$ 与 $+0.80\%$/$+0.64\%$\,pp) 相比,但在韦尔奇测试下是边缘的。

Mid-training, the stage between pre-training and alignment, is where a model的 per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit, at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band ($10\%$-$40\%$) is best for all five domains, and a calibrated permutation test for quadratic interiority gives $P\approx0.010$; the fitted mid-training-only curves, with 8B peaks between $9.9\%$ and $35.1\%$, reproduce for curve shape but not peak location. Second, the gaps survive a fixed-budget alignment pass: compensatory SFT raises 116/120 cells (mean $+4.32\%$) yet bridges $0/240$ pairs at a $5\%$ threshold and $30/240$ at a $10\%$ ratio, an equal-budget uniform control behaves almost identically, and a permutation null would bridge $13.8\pm3.3$ and $77.9\pm8.5$ pairs ($P<0.001$). Third, zero coverage collapses mid-training-only accuracy, though a FineWeb-Edu-only control shows the collapse is commingled with generic drift. An exploratory $\theta^*$ allocation attains the largest full-pipeline gain ($+4.36\%$ vs. $+0.80\%$/$+0.64\%$\,pp) but is marginal under Welch test.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)