Predicting how proteins interact with one another has become one of the most active and practically consequential problems in computational biology. The methods range from evolutionary covariance analyses that extract contact signals buried in sequence alignments, through template-based homology modeling, to deep learning architectures that learn interaction fingerprints directly from molecular surfaces or raw amino acid sequences. No single approach dominates across all settings, and “best practice” depends heavily on what you already know about the proteins in question, the scale of the problem, and the type of interaction you are trying to capture.
Why Experimental Data Alone Falls Short
Most computational prediction methods exist because experimental techniques for detecting protein-protein interactions are incomplete and error-prone. The yeast two-hybrid (Y2H) system, one of the workhorse assays for mapping interactions at large scale, illustrates the problem well. Estimated false-positive rates per unique interaction run roughly 25% in yeast and climb to 40–45% in worm and fly, with membrane-associated proteins driving much of the noise.
1PLoS Computational Biology. Where Have All the Interactions Gone? Estimating the Coverage of Two-Hybrid Protein Interaction MapsThe root cause is not fully understood in physical terms, but “sticky” hydrophobic proteins or patches are thought to produce weak, nonspecific binding that registers as a hit. Overexpression of bait and prey proteins can push these marginal contacts into observable range, inflating the count of apparent interactions. Using multiple reporter genes in a single screen can weed out some false positives, since a spurious interaction is unlikely to activate all reporters, though this also sacrifices weaker genuine interactions.
2PubMed Central. Making the right choice: Critical parameters of the Y2H systemsCross-linking mass spectrometry offers a complementary experimental lens, capturing interactions in a more native cellular context and providing distance constraints that can validate or refute predicted structures at proteome scale.
3bioRxiv. Cross-linking mass spectrometry discovers, evaluates, and validates the experimental and predicted structural proteomeThe upshot is that any computational prediction pipeline that trains on or validates against experimental interaction data inherits these limitations. Treating Y2H hits as ground truth without filtering is a recipe for learning artifacts.
Evolutionary and Sequence-Based Methods
One of the oldest computational strategies for identifying interacting residues exploits the fact that residues in physical contact tend to coevolve. If a mutation at one site destabilizes a binding interface, compensatory mutations at the partner site restore it, and these correlated changes accumulate across thousands of homologous sequences. Direct coupling analysis (DCA) formalized this idea by disentangling direct correlations from the indirect statistical noise that pervades large alignments.
DCA has been shown to correctly recapitulate the global contact map for the majority of protein domains tested, working purely from sequence information. Beyond single proteins, DCA can pick up signals from interdomain contacts in oligomers and even from alternative conformational states.
4PubMed Central. Direct-coupling analysis of residue coevolution captures native contacts across many protein families The approach scales to homo-oligomeric interfaces as well, where coevolution signals can be translated into spatial constraints for structure prediction.5PubMed Central. Large-scale identification of coevolution signals across homo-oligomeric protein interfaces by direct coupling analysis
A persistent challenge is that the evolutionary signal at a protein-protein interface is much weaker than the signal within a single protein’s fold. Intraprotein contacts coevolve under constant selective pressure; interprotein contacts do so only when the complex is functionally obligate, and even then the signal can be faint if the number of available homologous sequence pairs is small.
6The Journal of Physical Chemistry B. Expanding Direct Coupling Analysis to Identify Heterodimeric Interfaces from Limited Protein Sequence DataProtein Language Models
The more recent wave of sequence-based prediction leans on protein language models (PLMs), which learn contextual representations of amino acids from massive sequence databases in much the same way that large language models learn word relationships from text corpora. Early approaches simply extracted features from a pre-trained PLM and fed them into a classifier, treating each protein independently. The limitation is obvious: the model never sees the two proteins as an interacting pair.
PLM-interact addressed this by jointly encoding protein pairs, borrowing the “next-sentence prediction” concept from natural language processing. Trained on human interaction data and tested across mouse, fly, worm, E. coli, and yeast, it achieved state-of-the-art cross-species performance, demonstrating that language model representations contain enough biochemical information to generalize across evolutionary distances.
7Nature Communications. PLM-interact: extending protein language models to predict protein-protein interactions A paired-sequence variant, PPLM-PPI, pushed accuracy further on the same five-species benchmark, improving over the previous best method by anywhere from 4% to nearly 18% depending on the organism.
8Nature Communications. A paired sequence language model for protein-protein interaction modelingThese results matter practically because sequence is the cheapest data type you can get for a protein. For newly sequenced organisms or poorly characterized proteins with no structural data, a PLM-based predictor may be the only viable starting point.
Template-Based and Homology Modeling
When structural information exists for related proteins, template-based modeling is a powerful shortcut. The idea is straightforward: if two proteins share enough sequence similarity with proteins whose complex structure has already been solved, you can build a model of the new complex by analogy. SWISS-MODEL, one of the most widely used platforms, infers stoichiometry and overall complex architecture from homologous interacting pairs (sometimes called “interologs”).
9PubMed Central. SWISS-MODEL: homology modelling of protein structures and complexesTemplates can be identified either by sequence alignment or, if you already have monomer structures, by structural comparison.
10PubMed Central. Template-based structure modeling of protein-protein interactions The approach works well when the homology is clear. Above roughly 35% sequence identity, complexes tend to share similar structures and interaction modes, and an energy-based scoring function can rank the correct structure above about 92% of alternative decoy configurations.
11PubMed Central. Homology modelling of protein-protein complexes: a simple method and its possibilities and limitationsThe method breaks down when no suitable template exists. Novel interactions, or interactions that occur through disordered regions with no fixed fold, are effectively invisible to template-based modeling. It is also unreliable for weakly stable or transient complexes, which may use interaction modes not captured by the structural archive.
Deep Learning for Structure Prediction
AlphaFold2-Multimer reshaped the field by demonstrating that deep learning could predict the structures of protein complexes directly from sequence, without explicit templates. Yet the method is not infallible. On standard benchmarks, only about 60% of dimers are accurately predicted by the base multimer pipeline.
12PubMed Central. Improved protein complex prediction with AlphaFold-multimer by denoising the MSA profile Subsequent refinements have pushed accuracy higher. DeepSCFold, which exploits sequence-derived structural complementarity, achieved an average DockQ score of 0.59 on a benchmark set, compared to 0.49 for AlphaFold3, with comparable gains in overall structural similarity.13Nature Communications. High-accuracy protein complex structure modeling based on sequence-derived structure complementarity
A separate line of work focuses not on predicting the full complex structure but on identifying which surface residues mediate the interaction. Geometric deep learning methods read molecular surfaces as point clouds and learn patterns of chemical and geometric complementarity that fingerprint interaction sites. MaSIF, an early framework of this type, hypothesized that proteins with similar interaction modes share common surface fingerprints regardless of evolutionary relationship.
14Nature Methods. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning PInet built on this by consuming pairs of point clouds representing both partner surfaces simultaneously, learning complementarity directly from geometry and physicochemistry.
15Bioinformatics. Protein interaction interface region prediction by geometric deep learningPeSTo took a more minimalist approach, working as a parameter-free geometric transformer that acts on atoms described only by their elemental identity and spatial coordinates, with no explicitly encoded chemical properties like charge or hydrophobicity.
16Nature Communications. PeSTo: parameter-free geometric deep learning for accurate prediction of protein binding interfaces That such a stripped-down input representation can learn to predict binding interfaces suggests the relevant physicochemical information is implicitly encoded in atomic geometry.
The Negative Data Problem
Every machine learning model for interaction prediction needs both positive examples (proteins that interact) and negative examples (proteins that do not). Getting reliable negatives turns out to be surprisingly hard. In a real interaction network, fewer than 2% of all possible protein pairs actually interact, creating a massive imbalance.
17Briefings in Bioinformatics. Machine learning on protein–protein interaction prediction: models, challenges and trends Many datasets simply contain no negative examples at all, reporting only confirmed interactions. The common workaround of picking random protein pairs as negatives produces “easy” negatives that inflate benchmark performance without testing the model’s ability to distinguish close calls.
More principled approaches exist. One method harvests negatives from Y2H data by identifying protein pairs that were plausibly tested but not detected as interacting, assigning confidence scores based on shortest-path distance in the observed interaction network.
18PubMed. Negative protein-protein interaction datasets derived from large-scale two-hybrid experiments A newer strategy called TPPNI uses network topology to construct “hard” negatives: protein pairs that look like they should interact, based on their network neighborhood, but apparently do not. Models trained on these harder negatives generalize better to real-world prediction tasks.
19Bioinformatics. Topology-driven negative sampling enhances generalizability in protein–protein interaction predictionData Leakage and Evaluation Traps
Even with well-curated training data, how you split that data for evaluation matters enormously. A recent analysis of commonly used protein complex benchmarks found that standard splitting strategies based on sequence similarity or metadata introduce major data leakage, producing overoptimistic estimates of how well a model generalizes to truly novel complexes.
20arXiv. Revealing data leakage in protein interaction benchmarks In effect, many published benchmarks measure a model’s ability to memorize training examples rather than its practical predictive power.
The community has developed several resources to combat this. The Critical Assessment of Predicted Interactions (CAPRI) experiment, running since 2000, tests docking algorithms in blind predictions where the experimental structure is unknown to participants at the time of submission.
21PubMed Central. Assessing predictions of protein-protein interaction: the CAPRI experiment CAPRI has since expanded its scope and maintains benchmarking databases, including the Protein-Protein Docking Benchmark and the CAPRI Scoreset, that use standardized assessment metrics tested for robustness over two decades.
22PubMed Central. CAPRI-Q: The CAPRI resource evaluating the quality of predicted structures of protein complexesIf you are evaluating a new model, the practical recommendation is to avoid relying solely on sequence-identity-based splits. Test on truly held-out targets, ideally from CAPRI or from complexes deposited after your training data cutoff. The gap between reported accuracy and real-world accuracy in this field is frequently large enough to change your downstream decisions.
Transient Interactions, Disorder, and Short Linear Motifs
Not all protein interactions look the same. Permanent (obligate) complexes, where the components are rarely found apart, have different interface properties than transient (non-obligate) ones. Amino acid substitution patterns at these two types of interfaces differ with statistical significance, and specialized predictors exist that exploit those differences.
23PubMed Central. Predicting permanent and transient protein-protein interfaces The dynamics of the complex structure also offer a discriminating signal: obligate complexes tend to have more tightly coupled motions across their interfaces than non-obligate ones.24PLOS Computational Biology. DynaFace: Discrimination between Obligatory and Non-obligatory Protein-Protein Interactions Based on the Complex’s Dynamics
Transient interactions mediated by intrinsically disordered regions present a particularly tough prediction problem. Short linear motifs (SLiMs), typically 3 to 12 residues long, are key mediators of interactions between disordered protein regions and their structured partners.
25PubMed. Prediction of short linear protein binding regions Because these motifs are short, degenerate, and embedded in flexible regions with little fixed structure, they are easy to miss. Standard structure prediction tools like AlphaFold-Multimer often struggle here because the disordered region may not adopt a stable conformation until it binds its partner, and that induced-fit process is poorly captured by static structure predictors. Dedicated SLiM discovery tools that search for over-represented short sequence patterns across interaction datasets remain a necessary complement.26PubMed. Computational Prediction of Disordered Protein Motifs Using SLiMSuite
Scaling to Large Assemblies
Predicting the structure of a two-protein complex is hard enough. Scaling to large assemblies with tens of subunits adds combinatorial complexity that can overwhelm even deep learning methods. CombFold, a combinatorial assembly algorithm paired with AlphaFold2, accurately predicted 72% of large asymmetric assemblies (up to 30 chains and 18,000 amino acids) among its top 10 predictions.
27bioRxiv. Predicting structures of large protein assemblies using combinatorial assembly algorithm and AlphaFold2An important finding from the Monte Carlo tree search approach to assembly prediction is that symmetry matters a great deal. Complexes with dihedral symmetry are assembled successfully far more often than asymmetric ones. In one benchmark, asymmetric complexes achieved a median structural similarity score of only 0.49, compared to 0.80 for complexes with any kind of symmetry.
28Nature Communications. Predicting the structure of large protein complexes using AlphaFold and Monte Carlo tree search If you are working with a large asymmetric complex, treat any predicted structure with extra skepticism and look for experimental constraints to validate it.
Mutation Effects and Disease Context
One of the most consequential applications of interaction prediction is understanding how genetic mutations rewire cellular networks. Predicting the change in binding free energy caused by a mutation tells you whether a variant strengthens, weakens, or abolishes an interaction, which in turn helps explain disease mechanisms and drug resistance.
29PubMed Central. Decoding the effects of mutation on protein interactions using machine learning Computational tools like e-MutPath model how missense mutations propagate through interaction networks, identifying which specific protein-protein links are disrupted. In one example, a single mutation in KEAP1 was predicted to perturb its interaction with DPP3 in liver cancer, a prediction later confirmed experimentally by Y2H.
30Nucleic Acids Research. e-MutPath: computational modeling reveals the functional landscape of genetic mutations rewiring interactome networksThe interolog concept, transferring known interactions from one organism to another based on sequence conservation, has also proven useful for disease gene discovery. Protein complexes are conserved preferentially compared to transient interactions, and mapping human interactions onto yeast reveals a strong positive correlation between a protein’s evolutionary conservation and the number of partners it has.
31PubMed Central. Unequal evolutionary conservation of human protein interactions in interologous networks Cross-species network comparison has identified thousands of previously unknown protein functions and interactions, lending biological credibility to computationally transferred edges.32PubMed Central. Conserved patterns of protein interaction in multiple species
Subcellular localization adds another practically useful layer. Proteins targeting the same compartment are more likely to interact, and incorporating localization constraints into network-based disease models has revealed thousands of disease-pair associations not identifiable from shared genes or direct interactions alone.
33PubMed Central. Protein localization as a principal feature of the etiology and comorbidity of genetic diseasesDesigned Interactions and Therapeutic Applications
Interaction prediction is not just about understanding biology as it exists. The same structural and energetic models underpin de novo protein design, where the goal is to create new proteins that bind a chosen target with high affinity. In prospective tests on biologically important targets including ALK, LTK, and interleukin receptors, filtering designed protein binders with AlphaFold2 substantially increased the experimental success rate compared to designs filtered by physics-based energy scoring alone.
34Nature Communications. Improving de novo protein binder design with deep learningA frontier area is the design of heterochiral interactions, where a protein made of D-amino acids (mirror-image building blocks) is engineered to bind a natural L-protein target. Because D-proteins resist natural proteases, successful heterochiral binders could serve as exceptionally stable therapeutics and diagnostics.
35Cell Research. Accurate de novo design of heterochiral protein–protein interactionsPost-translational modifications add yet another wrinkle to therapeutic applications. Phosphorylation, glycosylation, and other modifications frequently alter protein structures and drug-binding affinities. Interaction-affecting modifications tend to occur closer in space to the ligand binding site than modifications that have no effect, which provides a geometric prior that specialized predictors can exploit.
36Journal of Medicinal Chemistry. Predicting the Post-translational Modification Effects on Protein−Ligand Interactions via End-Point Binding Free Energy Calculation: Database Creation and Strategy Optimization For drug development pipelines that rely on interaction modeling, ignoring the modification state of a target protein can lead to predictions that look right on paper but fail in the cellular context where modifications are present.