Deep learning has reshaped nearly every corner of biology, from predicting how proteins fold to reading meaning out of raw DNA sequences to designing molecules that might become drugs. The changes are not incremental upgrades to older computational tools. They represent a shift in what biologists can ask their data, with neural networks now solving problems that were considered intractable a decade ago. What makes the current moment especially interesting is that these models are not just getting better at narrow tasks; they are converging on general-purpose “foundation models” trained on enormous biological datasets, and the implications of that trend touch everything from medicine to biosafety.
Predicting and Designing Protein Structures
Protein structure prediction is probably the most visible success story. AlphaFold, developed by DeepMind, cracked a problem that had frustrated structural biologists for half a century: given only a protein’s amino acid sequence, predict its three-dimensional shape. A comparative study evaluating AlphaFold2 and AlphaFold3 against a set of over 1,600 known monomeric structures found that both versions achieved roughly 87–88% accuracy, with about two-thirds of predictions matching experimental structures closely and another quarter showing correct overall topology with some domain-level variation.1NAR Genomics and Bioinformatics. Comparative evaluation of the prediction accuracy of AlphaFold and ESMFold for monomeric and dimeric proteins ESMFold, a faster alternative that skips the time-consuming search for related sequences, managed about 50% high-accuracy predictions on the same benchmark, with roughly a quarter classified as incorrect. Speed versus accuracy is a real tradeoff here, and which model researchers choose depends on whether they need a quick approximation or a publication-grade structure.
Structure prediction was the opening act. The newer frontier is protein design: using generative models to create entirely new proteins that do not exist in nature but perform specified functions. Language models and diffusion-based architectures have shown they can generate novel, realistic protein sequences with desirable properties.2PubMed. Generative artificial intelligence for de novo protein design This is not just an academic exercise. Designed proteins could serve as custom enzymes for industrial chemistry, as scaffolds for vaccines, or as biosensors. The field is still young enough that most designs require experimental validation before anyone gets too excited, but the pace of improvement has been striking.
Reading the Genome at Single-Nucleotide Resolution
Understanding how DNA sequence variation affects gene regulation is one of biology’s hardest puzzles. Most disease-linked genetic variants sit outside the protein-coding portions of the genome, in regulatory regions whose function is poorly understood. Deep learning has become the primary tool for cracking this code. DeepSEA, one of the earlier models in this space, demonstrated that a neural network trained on large-scale chromatin-profiling data could predict how single-nucleotide changes alter gene regulation, improving the ability to prioritize variants linked to disease.3PubMed Central. Predicting effects of noncoding variants with deep learning–based sequence model
That was roughly a decade ago. The current generation of models operates at a different scale entirely. AlphaGenome, released in 2025, takes in a million base pairs of DNA sequence and predicts thousands of functional readouts at single-base-pair resolution, including gene expression, chromatin accessibility, histone modifications, transcription factor binding, and splice site usage. In head-to-head evaluations, it matched or exceeded the strongest competing models in 25 of 26 variant effect prediction benchmarks.4Nature. Advancing regulatory variant effect prediction with AlphaGenome
Meanwhile, a different class of models borrows the architecture behind large language models and applies it to DNA. HyenaDNA, for instance, was pretrained on the human reference genome with context lengths of up to one million nucleotide tokens, representing a 500-fold increase over previous attention-based genomic models.5NeurIPS Proceedings. Deep Learning in Biology: Innovations and Future Directions OmniNA takes a broader approach, training on over 91 million nucleotide sequences totaling more than a trillion bases across diverse species, and jointly learning from both sequences and their functional annotations. It achieved state-of-the-art or competitive results across 23 benchmarks.6Nucleic Acids Research. A foundation model for nucleotide sequences A recent comprehensive benchmark of five DNA foundation models across tasks like sequence classification, gene expression prediction, and variant effect quantification found that no single model dominates every task, suggesting the field still has room for architectural innovation.7Nature Communications. Benchmarking DNA foundation models for genomic and genetic tasks
Making Sense of Single-Cell and Spatial Data
Modern experiments can measure gene activity in hundreds of thousands of individual cells at once, but the resulting datasets are noisy and riddled with technical artifacts. When cells are processed in different batches, systematic differences between batches can swamp the real biological signal. DeepBID, a method built on a specialized autoencoder, tackles this by aligning cells from different batches into a shared space and iteratively cleaning up the noise. Applied to single-cell RNA data from patients with Alzheimer’s disease, it improved cell clustering and helped annotate previously unidentified cell types.8PubMed Central. Deep Batch Integration and Denoise of Single-Cell RNA-Seq Data
Spatial transcriptomics, which maps gene expression to physical locations in a tissue, is another area where deep learning has opened new doors. One early effort trained a model on over 30,000 spatially resolved gene expression measurements matched to standard tissue staining images from 23 breast cancer patients and identified over 100 genes whose expression could be predicted directly from the stained images at a resolution of 100 micrometers. These included known biomarkers of tumor heterogeneity and immune activation.9Nature Biomedical Engineering. Integrating spatial gene expression and breast tumour morphology via deep learning A more recent tool called GIST uses foundation models pretrained on millions of histology images to improve the integration of tissue images with transcriptomics data, substantially improving the accuracy of segmenting tumor microenvironments in lung, breast, and colorectal cancers.10PubMed Central. Deep Learning-Enabled Integration of Histology and Transcriptomics for Tissue Spatial Profile Analysis
Drug Discovery and Molecular Design
Drug discovery has historically been slow and expensive. Deep learning is being applied at multiple stages, from identifying promising molecular targets to designing the drugs themselves. Generative models can now explore vast chemical spaces and propose novel molecules with desired biological properties, though evaluating and prioritizing those candidates remains a challenge.11Journal of Chemical Information and Modeling. Generative Deep Learning for de Novo Drug Design: A Chemical Space Odyssey
Predicting how tightly a candidate drug binds to its protein target is a core task. Graph neural networks have become the go-to architecture here, because they can naturally represent the three-dimensional structure of proteins and small molecules as networks of atoms and bonds. GraphscoreDTA, for example, uses separate graph neural network modules to learn features from protein structures, ligand structures, and the interaction between them.12Bioinformatics. GraphscoreDTA: optimized graph neural network for protein–ligand binding affinity prediction GIGN takes a similar approach but explicitly encodes the physical interactions between protein and ligand atoms, including both covalent and noncovalent forces, into its message-passing framework.13PubMed. Geometric Interaction Graph Neural Network for Predicting Protein-Ligand Binding Affinities from 3D Structures (GIGN) These models are not replacing lab experiments, but they are dramatically narrowing the list of candidates that need to be tested physically.
Biological Imaging and Super-Resolution
Microscopy sits at the heart of cell biology, and deep learning has pushed the boundaries of what imaging can reveal. One of the cleaner demonstrations involved training a generative adversarial network to transform standard, diffraction-limited fluorescence microscopy images into super-resolved ones, without requiring a physical model of the imaging process or knowledge of the microscope’s point-spread function.14PubMed Central. Deep learning enables cross-modality super-resolution in fluorescence microscopy In practical terms, this means researchers can extract higher-resolution information from cheaper, faster imaging setups. The approach works across different microscopy modalities, which makes it broadly useful rather than tied to a single instrument.
Tracing Evolution and Population History
Protein language models, trained on millions of sequences, do something unexpected: they learn patterns that reflect evolutionary fitness. A method called evo-velocity uses the likelihoods assigned by these language models to infer the direction of evolutionary trajectories. By constructing a network of related protein sequences and letting the model assign a “velocity” to each connection based on changes in sequence likelihood, researchers can reconstruct evolutionary vector fields that reveal how protein families have diversified over time.15Cell Systems. Unsupervised Protein Language Model Infers Evolutionary Dynamics and Fitness Landscapes These same models estimate the protein fitness landscape more broadly, which is useful for predicting the effects of mutations and guiding protein design.16Nature Computational Science. Understanding language model scaling for protein fitness prediction
At the population level, deep learning has been applied to joint inference of natural selection and demographic history from genomic data. Traditional statistical methods struggle to separate local signals of selection from the genome-wide fingerprints of population size changes, but neural networks can learn to disentangle the two simultaneously.17PLOS Computational Biology. Deep Learning for Population Genetic Inference Understanding which genomic features drive these inferences remains an active area of work, with tools like ConfuseNN designed to probe what convolutional neural networks actually learn from population genomic data and where their architectures fall short.18PubMed Central. ConfuseNN: Interpreting convolutional neural network inferences in population genomics with data shuffling
Gene Regulatory Networks and Disease
Genes do not act alone. They are organized into regulatory networks where transcription factors switch other genes on and off, and understanding these networks in specific cell types is critical for understanding disease. Deep learning methods are now being used to reconstruct cell-type-specific gene regulatory networks from single-cell data. One tool, scMultiomeGRN, was applied to brain tissue from Alzheimer’s patients and identified key transcription factors with altered regulatory networks in microglia, the brain’s immune cells. Two of these, SPI1 and RUNX1, showed significantly increased regulatory activity in Alzheimer’s samples and have been independently linked to known Alzheimer’s genetic risk loci.19Nucleic Acids Research. Deep learning-based cell-specific gene regulatory networks inferred from single-cell multiome data
A related approach, DeepDRIM, was used to compare regulatory networks in immune cells from COVID-19 patients with mild versus severe disease. The differentially regulated gene targets were enriched in processes like lysosome function and response to low oxygen, processes already known to be relevant to coronavirus infection.20Briefings in Bioinformatics. DeepDRIM: a deep neural network to reconstruct cell-type-specific gene regulatory network using single-cell RNA-seq data These methods are also being extended to model disease similarity by learning representations of genes and then using those to characterize the diseases they contribute to.21Bioinformatics. Evaluating disease similarity based on gene network reconstruction and representation
Synthetic Biology and Metabolic Pathway Design
Synthetic biology aims to engineer living cells to perform useful tasks, whether producing biofuels, manufacturing pharmaceuticals, or breaking down pollutants. Designing the metabolic pathways inside these engineered cells is a bottleneck, and deep learning is increasingly being used to predict viable pathway configurations and optimize them computationally before anyone picks up a pipette.22PubMed. Deep learning for metabolic pathway design The long-term goal is to merge pathway design, host organism selection, and growth condition optimization into a single computational pipeline, though achieving this end-to-end integration remains more aspiration than reality.23PubMed Central. Machine Learning and Deep Learning in Synthetic Biology: Key Architectures, Applications, and Challenges
The Interpretability Problem
One persistent criticism of deep learning in biology is that the models are black boxes. A network might accurately predict which DNA sequence variants affect gene regulation, but if no one can explain why, the prediction is less scientifically useful. Several lines of work are tackling this. ExplaiNN, for example, is designed to be interpretable from the ground up: each unit in the network learns a filter that can be directly converted into a sequence motif and matched against known transcription factor binding profiles in databases like JASPAR.24PubMed Central. ExplaiNN: interpretable and transparent neural networks for genomics
Post-hoc explanation methods offer a complementary approach. Techniques like Grad-CAM and Integrated Gradients can be applied to trained models to highlight which parts of an input sequence drove a particular prediction. When applied to protein sequence classification, these methods identified motifs enriched in amino acids commonly found at catalytic and metal-binding sites in enzymes, suggesting the models are learning biochemically meaningful features rather than statistical artifacts.25arXiv. XAI-Driven Deep Learning for Protein Sequence Functional Group Classification Interpretability is not a solved problem, but the tools are maturing.
Data Leakage and Benchmarking Pitfalls
The field has a credibility problem that does not get enough attention outside specialist circles: data leakage. Because biological datasets have complex internal relationships, it is easy for information about the test set to leak into training, inflating apparent performance far beyond what the model would achieve in the real world.26PubMed Central. Guiding questions to avoid data leakage in biological machine learning applications This is not a theoretical concern. PLABench, a recently introduced benchmarking framework for protein-ligand binding affinity prediction, was designed specifically to address widespread leakage and inconsistent evaluation across the field. It standardizes structural inputs and enforces rigorous data splits to enable fair comparison across models.27bioRxiv. Leakage-controlled benchmarking reveals generalization limits of deep learning for protein-ligand binding affinity prediction
A related issue is out-of-distribution generalization. Models trained on one distribution of biological data often fail when the data shifts even modestly from what they have seen before.28PubMed Central. Out of distribution learning in bioinformatics: advancements and challenges For drug discovery, this means a model that performs well on known protein families might fail completely on a novel target. For genomics, it means a model trained on one population’s genetic variation might not generalize to another. These are not bugs that will be casually patched; they reflect fundamental challenges in how neural networks learn.
Self-Driving Labs
One of the more ambitious visions for deep learning in biology is the self-driving laboratory, where AI not only analyzes experimental results but decides what experiments to run next. These systems pair fully automated robotic experiments with artificial intelligence that designs each new round of tests.29PubMed. Perspectives for self-driving labs in synthetic biology The SAMPLE platform offers a concrete example: an intelligent agent designed and optimized protein variants by issuing commands to a remote cloud laboratory, which handled gene assembly, protein expression, and biochemical assays without human intervention in the loop.30PubMed Central. Autonomous ‘self-driving’ laboratories: a review of technology and policy implications These platforms are still expensive and limited to well-defined experimental workflows, but they point toward a future where the design-build-test cycle in biology speeds up by orders of magnitude.
Dual-Use Risks and Biosafety
As biological foundation models become more powerful and more accessible, the dual-use problem grows harder to ignore. Open-weight models that can design proteins, predict pathogen fitness, or optimize metabolic pathways for useful purposes could, in principle, be repurposed by bad actors to develop more dangerous biological agents. Current mitigation strategies focus on filtering potentially hazardous data during model training, but the robustness of that approach is unclear, especially against someone willing to fine-tune a model on new data after release. BioRiskEval, a recently proposed framework, is designed to evaluate how well these safety procedures hold up under adversarial conditions.31arXiv. Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models The tension between open science, which accelerates discovery, and restricted access, which reduces risk, does not have an obvious resolution, and the biology community is still working out where the lines should be drawn.