蛋白质结构预测的基础模型在某些目标上仍然不可靠。外部预言机可以标记并纠正这些故障,,但生物预言机价格昂贵,,这使得预言机预算成为关键限制。现有的指导方法,(例如FK-steering, DPO, 和最佳K-of-N 采样,)在如何花费预算, 方面有所不同,但不存在系统比较来指导方法选择。为了弥补这一差距,,我们将这些方法与最近提出的输出优化(O3), 进行基准测试,该优化在生成模型的 潜在子空间中应用现成的优化器。我们将 O3 的用途扩展到蛋白质结构预测模型。总体而言,我们的工作为甲骨文预算意识指导提供了第一个实用参考。我们对两个蛋白质目标,钙调蛋白(1CLL)和大肠杆菌天冬氨酸转氨甲酰酶(9EEH),的评估表明,没有一种方法能够在所有预算和预言中始终占主导地位。具体来说,, O3 被证明在低预言机预算, 时最有效,而 FK 转向和 DPO 则随着预算的增加而表现出性能的提高。我们将这些发现提炼为在现实世界的预言机预算限制下运营的从业者的可行建议。

Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. Existing guidance methods, such as FK-steering, DPO, and Best K-of-N sampling, differ in how they spend this budget, yet no systematic comparison exists to guide method selection. To bridge this gap, we benchmark these methods alongside the recently proposed Optimisation Over Outputs (O3), which applies off-the-shelf optimisers within a generative model的 latent subspace. We extend the usage of O3 to protein structure prediction models. Overall, our work provides the first practical reference for oracle budget-aware guidance. Our evaluation on two protein targets, calmodulin (1CLL) and E. coli aspartate transcarbamoylase (9EEH), reveals that no single method consistently dominates across all budgets and oracles. Specifically, O3 proves most effective at low oracle budgets, while FK-steering and DPO demonstrate improved performance as the budget increases. We distil these findings into actionable recommendations for practitioners operating under real-world oracle-budget constraints.

科目: 人工智能 (cs.AI); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)