Protein Prediction: How It Works and Why It Matters

Protein prediction refers to the computational task of figuring out a protein’s three-dimensional shape from its amino acid sequence, and recent deep-learning methods have made it remarkably accurate. The breakthrough came in 2020, when DeepMind’s AlphaFold2 scored 92.4 on the main accuracy metric at the field’s premier competition, effectively solving the single-chain structure problem that had stumped researchers for over fifty years.1PubMed. From CASP13 to the Nobel Prize: DeepMind’s AlphaFold Journey in Revolutionizing Protein Structure Prediction and Beyond But “prediction” now covers far more than just shape: it includes forecasting how proteins move, how they interact with partners, how mutations affect health, and even designing entirely new proteins from scratch.

Why Protein Shape Is So Hard to Predict

A protein starts life as a chain of amino acids strung together like beads. That chain then folds into a specific three-dimensional shape, and the shape determines what the protein does. The thermodynamic principle behind this, first articulated by Christian Anfinsen, is that a small protein settles into the arrangement that minimizes its free energy under normal biological conditions. The catch is a famous thought experiment known as Levinthal’s paradox: a chain of even modest length has so many possible configurations that randomly sampling them would take longer than the age of the universe.2Europe PMC / Chemical Reviews. Protein folding thermodynamics and dynamics: where physics, chemistry, and biology meet Real proteins fold in milliseconds to seconds because physics funnels them toward the right answer, but capturing that funnel computationally has been the core difficulty.

For decades, the only reliable way to determine a protein’s shape was experimental. X-ray crystallography, which involves coaxing proteins into crystals and bombarding them with X-rays, has long been the gold standard for small-to-medium proteins. Cryo-electron microscopy has become a powerful complement, especially for larger and more flexible assemblies that resist crystallization.3PubMed Central. X-rays in the Cryo-Electron Microscopy Era: Structural Biology’s Dynamic Future Both methods are slow, expensive, and require specialized equipment. By the early 2020s, experimental structures existed for only a small fraction of known proteins, which is what made computational prediction so urgent.

Classical Computational Methods

Before deep learning took over, two broad families of computational approaches dominated. Homology modeling worked by finding a protein whose sequence resembled the target’s and whose structure was already known, then threading the target sequence onto that known scaffold. This approach was useful when a close relative existed in the database but fell apart for proteins with no known structural cousins. Accurate energy functions were considered essential for refining these models, and researchers used a mix of physics-based force fields and knowledge-based scoring functions to evaluate candidate shapes.4PubMed. Force fields for homology modeling

The other family, sometimes called ab initio or “from scratch” prediction, tried to fold a protein without any template. These methods relied on molecular dynamics simulations, running physics calculations step by step to simulate how atoms move in water over time.5PubMed Central. Refinement of homology-based protein structures by molecular dynamics simulation techniques The results were often approximate, and the computational cost was enormous. Both families made steady but slow progress through community competitions held every two years, until deep learning changed the trajectory entirely.

How Deep Learning Changed Everything

The key insight behind modern prediction methods is that evolution has already run the experiment billions of times. When you line up related protein sequences from across the tree of life, the patterns of amino acids that change together reveal which positions in the chain sit close together in three-dimensional space. Early statistical methods called direct coupling analysis (DCA) extracted these co-evolutionary signals, but they were noisy and struggled to separate genuine structural contacts from artifacts of shared ancestry. Protein language models trained on multiple sequence alignments turned out to be far more robust: in controlled tests, they disentangled structural constraints from phylogenetic noise two to three times better than the older statistical methods.6Nature Communications. Protein language models trained on multiple sequence alignments learn phylogenetic relationships

AlphaFold2, the system that dominated the 2020 competition, combined these evolutionary signals with a neural network architecture that iteratively refined its understanding of both the sequence relationships and the spatial geometry. The result was structures that, for most single-chain proteins, were nearly indistinguishable from experimentally determined ones. DeepMind subsequently released predicted structures for virtually every known protein, an achievement recognized with a Nobel Prize in Chemistry.

A parallel line of research has pushed even further by dropping the requirement for evolutionary alignments altogether. Researchers at Meta AI scaled a protein language model to 15 billion parameters and showed that atomic-level structure can emerge directly from a single protein sequence, without ever looking at related sequences.7PubMed. Evolutionary-scale prediction of atomic-level protein structure with a language model This matters because some proteins, like antibodies engineered in the lab, have few or no natural relatives in sequence databases. Single-sequence methods such as RaptorX-Single have shown they can outperform alignment-based methods on antibodies and proteins with very few known homologs.8PubMed Central. Single-sequence protein structure prediction by integrating protein language models

Predicting How Proteins Interact

Most proteins do not work alone. They bind to other proteins, forming complexes that carry out everything from DNA replication to immune defense. Predicting the structure of these complexes is a harder problem than predicting a single chain, because you need to figure out not just each protein’s shape but how the two surfaces fit together. AlphaFold-Multimer, an extension of the original system, enabled structural modeling of protein complexes with unprecedented accuracy.9Bioinformatics. AlphaPulldown—a python package for protein–protein interaction screens using AlphaFold-Multimer Tools built on top of it now allow researchers to screen large numbers of candidate interactions computationally, essentially running “pull-down” experiments on a computer instead of at the bench.

There is an important limitation, though. AlphaFold-Multimer can predict what a complex looks like if two proteins do interact, but it cannot reliably tell you whether they interact in the first place.10PubMed Central. SpatialPPI: Three-dimensional space protein-protein interaction prediction with AlphaFold Multimer Give it two proteins that never meet in a cell, and it may still produce a confident-looking model. Newer methods like SpatialPPI are trying to add that discriminative layer, but distinguishing real interactions from plausible-looking fictions remains an active challenge.

Where Prediction Still Struggles

The single-structure paradigm works beautifully for proteins that fold into one stable shape and stay there. But a large fraction of the proteome does not cooperate. Intrinsically disordered regions, stretches of protein chain that remain flexible and lack a fixed three-dimensional structure, make up over 30% of the human proteome and play critical roles in cell signaling and disease.11Preprints. AF-RECAL: Multi-Modal Structural Recalibration of AlphaFold 3 Hallucinations in Intrinsically Disordered Human Disease Targets Current prediction tools tend to collapse these floppy regions into artificially rigid structures, sometimes even assigning them high confidence scores that make the hallucinated structure look trustworthy. This compaction bias is a recognized problem with AlphaFold3’s diffusion-based architecture and an area of active correction.

Even proteins that do fold into a stable shape are not truly static. They breathe, flex, and sample alternative conformations that can be critical for function. A binding pocket might only open transiently, creating what researchers call a cryptic site. Classical molecular dynamics can model these motions but is computationally brutal. A new generation of generative deep-learning tools aims to produce realistic ensembles of protein conformations at a fraction of the cost. One such tool, BioEmu, samples functionally relevant motions including cryptic pocket formation, local unfolding events, and large-scale domain rearrangements.12bioRxiv. Scalable emulation of protein equilibrium ensembles with generative deep learning If these methods mature, they could bridge the gap between static prediction and the dynamic reality of proteins in cells.

Designing Proteins That Have Never Existed

Prediction and design are two sides of the same coin. If you can predict shape from sequence reliably, you can invert the process: start with a desired shape and compute a sequence that would fold into it. This “inverse folding” approach has been transformed by deep learning. ProteinMPNN, a neural network trained on the relationship between protein backbones and their amino acid sequences, recovers the correct sequence about 52% of the time on natural protein backbones, a substantial jump over the roughly 33% achieved by the previous best physics-based method.13PubMed Central. Robust deep learning-based protein sequence design using ProteinMPNN

Newer architectures are pushing those numbers higher still. DualMPNN, which incorporates information from structurally similar templates, has achieved sequence recovery rates around 65 to 71% across standard benchmarks, outperforming ProteinMPNN by 12 to 17 percentage points depending on the test set.14NeurIPS Proceedings. Protein Prediction: How It Matters and Why It Matters Other groups are incorporating protein surface geometry alongside backbone structure, recognizing that the shape of a protein’s outer surface constrains which amino acids can sit on exposed positions.15bioRxiv. Protein inverse folding through joint modeling of surface and backbone geometry

Going beyond sequence design, tools like RFdiffusion can generate entirely new protein backbones from scratch. By training a structure prediction network on denoising tasks, the researchers produced a generative model capable of designing protein binders, symmetric assemblies, enzyme active sites, and metal-binding proteins. Hundreds of these designs were experimentally validated.16PubMed Central. De novo design of protein structure and function with RFdiffusion The practical implication is that biologists are no longer limited to the proteins nature invented. If you need a protein that binds a specific target or catalyzes a specific reaction, you can increasingly design one to order.

Drug Discovery and Virtual Screening

Knowing a protein’s shape tells you what small molecules might fit into its pockets, which is the starting point for most drug development. Structure-based virtual screening uses predicted or experimentally determined protein structures to computationally test billions of chemical compounds for potential binding. A platform called RosettaVS, which incorporates the ability to model receptor flexibility, recently screened multi-billion compound libraries against two unrelated disease targets. Against a ubiquitin ligase called KLHDC2, it identified seven hits at a 14% hit rate, while against a pain-related sodium channel it found four hits at a 44% hit rate, all with binding affinities in the single-digit micromolar range. Both screens were completed in under a week, and an X-ray crystal structure confirmed the predicted binding pose for one of the KLHDC2 hits.17PubMed Central. An artificial intelligence accelerated virtual screening platform for drug discovery

Those timelines are striking. Traditional high-throughput physical screening of that many compounds would take months and cost vastly more. Predicted structures from AlphaFold are increasingly being used as starting points for these virtual screens when no experimental structure exists, though the accuracy of the predicted binding pockets can vary, especially if the protein flexes significantly at the binding site.

Predicting How Mutations Cause Disease

Every person’s genome contains thousands of missense variants: single-letter changes in DNA that swap one amino acid for another in a protein. Most are harmless, but some cause disease by destabilizing the protein’s structure or disrupting its function. Classifying which variants matter and which do not is one of the central problems in clinical genetics. AlphaMissense, built on AlphaFold’s architecture and fine-tuned on human and primate population frequency data, predicts the pathogenicity of missense variants across the entire human proteome by combining structural context with evolutionary conservation.18PubMed. Accurate proteome-wide missense variant effect prediction with AlphaMissense

For well-studied genes, these predictions can be combined with other structural information for even better performance. A study of the tumor suppressor p53 found that integrating AlphaMissense scores with protein folding stability calculations improved the ability to classify the impact of missense variants beyond what either method achieved alone.19PubMed Central. Integration of protein stability and AlphaMissense scores improves bioinformatic impact prediction for p53 missense and in-frame amino acid deletion variants This kind of layered approach, using structural prediction as one signal among several, is likely how variant interpretation will work in clinical settings going forward.

Enzyme Engineering and Plastic Degradation

Protein prediction tools are also accelerating the engineering of enzymes for industrial and environmental applications. A high-profile example is the effort to improve enzymes that break down PET, the plastic used in beverage bottles. Naturally occurring PET-degrading enzymes work, but they are too slow and too fragile for industrial use. Directed evolution, a method that mimics natural selection in the lab, has been used to improve the catalytic activity and heat tolerance of these enzymes.20PubMed Central. Improving plastic degrading enzymes via directed evolution Structure prediction plays a supporting role here by revealing which parts of the enzyme contact the plastic surface, where mutations might improve binding or stability, and how engineered variants are likely to fold.

The broader challenge is that plastic waste is chemically diverse, and enzymes that degrade one type of plastic typically have little effect on others. Structural biology helps explain why: the active site of a PET-degrading enzyme is shaped to grip the specific chemical bonds in PET, not the different bonds found in polyethylene or polystyrene.21PubMed Central. Engineering Plastic Eating Enzymes Using Structural Biology Designing enzymes for new plastic substrates, or broadening the specificity of existing ones, is an area where de novo protein design tools like RFdiffusion could eventually make a large impact.

Post-Translational Modifications Add Complexity

The story does not end once a protein folds. Cells routinely attach chemical groups to proteins after they are made, through processes called post-translational modifications. Phosphorylation, glycosylation, and acetylation are among the most common. These modifications can shift a protein’s shape, alter its interactions, and toggle its activity on or off. Studies comparing modified and unmodified versions of the same protein have found that the structural changes are real but usually subtle: only about 7% of glycosylated and 13% of phosphorylated proteins show global structural shifts greater than 2 angstroms. Phosphorylation appears to reduce overall conformational variability by about 25%, effectively stabilizing the structure.22Bioinformatics. Post-translational modifications induce significant yet not extreme changes to protein structure

These modifications tend to cluster in intrinsically disordered regions, and sites that are targeted by multiple types of modification show an especially strong preference for disorder.23PubMed Central. The structural and functional signatures of proteins that undergo multiple events of post-translational modification This makes sense: flexible regions can accommodate the structural adjustments that different modifications demand. But it also means that predicting how modifications alter protein behavior requires tools that handle disorder well, which, as noted earlier, is precisely where current prediction methods are weakest. Computational approaches for predicting modification-induced conformational changes exist, but they lag behind the structure prediction tools in accuracy and generality.24PubMed. In silico prediction and characterization of protein post-translational modifications

Beyond Proteins to RNA and Multi-Molecule Systems

The success of deep learning in protein structure prediction has inspired parallel efforts for other biological molecules. AlphaFold3, the latest version of DeepMind’s system, expanded to model not just proteins but also DNA, RNA, and small molecules, moving toward the goal of predicting entire molecular assemblies. Several deep-learning tools for RNA 3D structure prediction have emerged, including DRfold, DeepFoldRNA, RhoFold, RoseTTAFoldNA, trRosettaRNA, and AlphaFold 3 itself.25bioRxiv. A Comparative Review of Deep Learning Methods for RNA Tertiary Structure Prediction RNA structure prediction is generally considered harder than protein prediction because RNA molecules are more flexible and their structures depend more heavily on environmental conditions. The field is roughly where protein prediction was a decade ago: promising but far from solved.

Domain segmentation is another area where prediction tools are expanding their reach. Large proteins often consist of multiple independently folding modules called domains. Identifying these domains automatically is useful for classifying proteins, understanding their evolutionary history, and breaking complex structures into manageable pieces. Merizo, a deep-learning method for domain segmentation, was applied to the human proteome and identified over 40,000 putative domains that could be matched to known structural families.26Nature Communications. Merizo: a rapid and accurate protein domain segmentation method using invariant point attention

The Cost of Training and Running These Models

The accuracy of modern protein prediction comes at a computational price. Training the largest protein language models involves hundreds of runs across model sizes ranging from millions to billions of parameters, fed with hundreds of billions of sequence tokens drawn from datasets of nearly a billion protein sequences.27NeurIPS Proceedings. Optimally Training Protein Language Models This is comparable in scale to training large language models for text, and it requires specialized GPU clusters that are beyond the budget of most academic labs.

Inference, the cost of actually using a trained model to predict a structure, is much cheaper but still non-trivial for large proteins or massive screening campaigns. The trend is toward more efficient architectures and community-maintained servers that run predictions for free, such as the AlphaFold Protein Structure Database. But for cutting-edge applications like screening billions of protein-drug interactions or generating ensembles of conformations, compute remains a bottleneck that shapes which questions researchers can realistically ask.

Biosecurity Risks of Protein Design

The ability to design proteins from scratch raises genuine biosecurity concerns. AI-generated proteins can be functionally equivalent to known toxins while sharing little sequence similarity, which means current screening tools that work by comparing new sequences to databases of known dangerous proteins would not flag them.28PubMed Central. Protein design, generative AI and biological security The widespread availability of open-source design tools further lowers the barrier. Mitigation strategies being discussed include function-based screening that looks at what a designed protein does rather than what it looks like in sequence space, synthesis screening by DNA providers, and tiered access to the most capable models. The tension between open science, which has driven the field’s rapid progress, and the need for guardrails is one that the community is still working through, with no consensus solution yet in place.