蛋白质’的功能由其结构,和结构—决定,蛋白质折叠—的方式由其氨基酸序列,(蛋白质的组成部分)决定。
A protein的 function is determined by its structure, and structure — the way a protein folds — is determined by its sequence of amino acids, the building blocks of proteins.
许多设计新型蛋白质, 的方法,包括可以与细胞中致病分子结合的示例, 都涉及两步过程: 首先是结构,,然后机器学习框架生成可能采用该结构的序列库。
Many methods for designing novel proteins, including examples that could bind to a disease-causing molecule in our cells, involve a two-step process: The structure comes first, and then a machine-learning framework generates a repertoire of sequences that could potentially adopt that structure.
在自然界,中,许多不同的氨基酸序列可以折叠成相同的结构。 At the same time, one amino acid sequence can potentially adopt different structures depending on the protein的 flexibility or a functional trigger. Therefore, when researchers use artificial intelligence to design new proteins, the challenge is to guide AI to “see” that there are many potentially useful answers — that many sequences can adopt the same fold
In nature, many different amino acid sequences can fold into the same structure. At the same time, one amino acid sequence can potentially adopt different structures depending on the protein的 flexibility or a functional trigger. Therefore, when researchers use artificial intelligence to design new proteins, the challenge is to guide AI to “see” that there are many potentially useful answers — that many sequences can adopt the same fold
将此框架添加到蛋白质设计流程中将使研究人员能够设计结构上可行的蛋白质,其序列与任何天然蛋白质的序列都不相似。
Adding this framework to a protein design pipeline will allow researchers to design structurally feasible proteins with sequences that don’t resemble those of any native protein.
“如果我们’正在考虑一个完全新颖的,设计结构,,那么就没有本地序列可以将其与,”进行比较,研究生兼主要作者福斯特·伯恩鲍姆(Foster Birnbaum)说。 “我们真正关心的是生成的序列折叠成所需结构的可能性,模型对序列能量景观的理解程度,以及它如何预测突变对蛋白质稳定性的影响。”
“If we’re thinking about a completely novel, designed structure, there would be no native sequence to compare it to,” says graduate student and lead author Foster Birnbaum. “What we actually care about is how likely the generated sequences are to fold into the desired structures, how well the model understands the sequence-energy landscape, and how well it can predict the effect of mutations on the stability of the protein.”
就像人工智能最近推动了一些巨大的社会变革一样,,机器学习也影响了基础生物学研究的速度和广度。直到最近才可以可靠地使用计算模型来生成蛋白质结构或序列。也许当今使用最广泛的型号是,,但是, 是在 2022 年发布的。
In the same way that AI has recently powered some dramatic social changes, so too has machine learning impacted the pace and breadth of fundamental biological research. Only recently has it become possible to reliably use a computational model to generate a protein structure or sequence. Perhaps the most widely used model today, however, was released in 2022.
“对于’这个领域,’的发展速度与生物学中的机器学习一样快,,该模型尚未被超越—,我们’一直试图理解为什么是,,以及该模型为何如此有用,” Birnbaum 说。
“For a field that的 moving as fast as machine learning in biology, that model has not been surpassed — we’ve been trying to understand why that is, and what it is about that model that makes it so useful,” Birnbaum says.
Birnbaum 首先对研究人员称为 “noise,” 的战略应用感兴趣,或者在训练期间向蛋白质结构添加变化。噪声降低了模型过度模仿天然序列, 的倾向,增加了 能够生成序列的结构的多样性。
Birnbaum was first interested in strategic applications of something researchers call “noise,” or adding variations to a protein structure during training. Noise decreases the tendency of the model to overly mimic native sequences, increasing the diversity of structures for which it的 able to generate sequences.
PottsMPNN 还使用成对分布来捕获氨基酸之间的相互作用。能够解释蛋白质中一对位置上所有 20 个可能的序列选项之间的物理相互作用,是 PottsMPNN 比其他方法更准确地模拟序列能量景观的关键原因。
PottsMPNN also uses a pairwise distribution to capture interactions between amino acids. The ability to account for the physical interactions between all 20 possible sequence options at a pair of positions in the protein is a key reason that PottsMPNN more accurately models the sequence-energy landscape than other methods.
最后, Birnbaum 说, 他们引入了一组进化相关的序列来训练 PottsMPNN 框架,以教导模型不同的序列如何采用相同的折叠结构。
Finally, Birnbaum says, they introduced sets of evolutionarily related sequences into training the PottsMPNN framework to teach the model how different sequences can adopt the same folded structure.
Birnbaum 承认,在试图摆脱对天然序列, 的束缚时,纳入进化信息的, 在某些方面, 仍然是对它们的依赖。但 PottsMPNN 成功证明,随着模型越来越少地依赖于天然序列,,结构兼容性和能量预测,(包括新蛋白质,)得到改善。
Birnbaum acknowledges that in trying to shift away from adhering to native sequences, incorporating evolutionary information is, in some ways, still a reliance on them. But PottsMPNN succeeded in demonstrating that as the model depends less and less on native sequences, structural compatibility and energy prediction, including for novel proteins, improve.
人工智能时代的蛋白质设计
Protein design in the age of AI
“一旦我们可以设计任何我们想要的蛋白质,,这使我们能够进行数量惊人的生物工程,” Birnbaum 说。 “It的 是一项艰巨的任务,,但我’m 对本世纪生物学的进展非常乐观。”
“Once we can design any protein we want, that enables us to do a potentially scary amount of biological engineering,” Birnbaum says. “It的 a difficult task, but I’m really optimistic about this century的 progress in biology.”
Birnbaum 希望该模型能够针对特定任务, 进行进一步改进和微调,这在过去导致了更好的预测,,例如, 对特定突变的结果或结果的预测。
Birnbaum hopes that the model could be further improved and fine-tuned for a specific task, which has in the past led to better predictions, for example, on the outcome or consequence of a particular mutation.
最终,, 根据 Keating, “我们的方法推动了该领域为各种应用设计有用的新型天然蛋白质,同时为未来的进步提供了更坚实的基础。”
Ultimately, according to Keating, “Our methods move the field toward designing useful new-to-nature proteins for diverse applications while providing a stronger foundation for future advances.”