AI in Biology: New Horizons for Data-Driven Discovery

Artificial intelligence has moved from a peripheral tool in biology to something closer to a core method, reshaping how researchers predict molecular structures, interpret genomes, design drugs, and even reconstruct evolutionary history. The shift accelerated dramatically after deep learning models began matching or outperforming decades of experimental work in tasks like protein folding, and it has since spread into nearly every subdiscipline of the life sciences. What makes this moment different from earlier waves of computational biology is the scale: foundation models trained on tens of millions of cells or billions of amino acid sequences are learning biological grammar that no human could manually encode.

Protein Structure and the AlphaFold Watershed

The clearest demonstration of AI’s power in biology came from protein structure prediction. For decades, determining the three-dimensional shape of a protein required laborious experimental techniques that could take months or years per structure. AlphaFold changed that calculus. The system, built on deep neural networks, became the first computational method to regularly predict protein structures with atomic-level accuracy, even for proteins with no known similar structure. At the CASP14 competition, a rigorous blind test of structure prediction methods, AlphaFold’s accuracy was competitive with experimental structures in the majority of cases, far outstripping any previous approach.1PubMed Central. Highly accurate protein structure prediction with AlphaFold

The practical fallout has been enormous. Researchers who once waited months for a crystal structure can now generate a reliable prediction in minutes. This does not replace experimental validation for every purpose, but it has opened the door to structural analyses at a scale that was previously unimaginable. The public database of AlphaFold predictions now covers hundreds of millions of protein structures, providing a resource that biologists across disciplines draw on daily.

Designing Proteins That Nature Never Made

Predicting existing structures was the first act. The second is designing entirely new ones. Generative AI architectures, including language models adapted from the text-processing world and diffusion models borrowed from image generation, have proven surprisingly effective at creating novel proteins from scratch. These models generate amino acid sequences that fold into stable structures and perform specified functions, proteins that evolution never produced. Current state-of-the-art design protocols achieve experimental success rates approaching 20%, meaning roughly one in five computationally designed proteins actually works as intended when synthesized in the lab.2PubMed. Generative artificial intelligence for de novo protein design

A 20% hit rate might sound modest, but it represents a radical improvement over earlier approaches, where the vast majority of designed proteins failed to fold properly. This is enough to make de novo protein design practical for real applications: custom enzymes for industrial chemistry, therapeutic proteins tailored to specific diseases, and biosensors engineered to detect particular molecules. The bottleneck has shifted from “can we design a protein that works?” to “which of the many plausible designs should we build first?”

Reading the Genome With New Eyes

AI models trained on genomic data are changing how researchers interpret genetic variation. Traditional approaches for evaluating whether a DNA mutation is harmful relied on comparing it to known pathogenic variants or assessing its location in a well-characterized gene. Deep learning models now go further, predicting the functional consequences of mutations across protein structure, RNA splicing, and noncoding regulatory regions. But this capability has revealed something surprising: the same mutation can be predicted as harmful in one person’s genetic background and harmless in another’s. Across multiple deep learning models, many clinical variants show heterogeneous predicted effects depending on which version of the surrounding DNA they sit in.3PubMed Central. Genetic background shapes AI-predicted variant effects

This finding has real implications for clinical genetics. It suggests that variant classification, the process of deciding whether a mutation is pathogenic, may need to account for the full genetic context rather than treating each variant in isolation. The same substitution in a gene could behave differently in two patients because of differences in linked regulatory elements or nearby variants that compensate for (or exacerbate) its effect. It is a more nuanced picture than the binary “pathogenic or benign” framework that most genetic testing currently uses.

Foundation Models for Single Cells

One of the most data-rich frontiers in biology is single-cell genomics, where technologies now routinely measure gene activity in hundreds of thousands or millions of individual cells from a single experiment. The sheer volume of data has made it a natural fit for foundation models, large AI systems pre-trained on massive datasets and then fine-tuned for specific tasks. scGPT, for instance, was trained on a repository of over 33 million cells and can be adapted for tasks ranging from classifying cell types to predicting how cells respond to genetic or chemical perturbations.4Nature Methods. scGPT: toward building a foundation model for single-cell multi-omics using generative AI CellFM, another foundation model pre-trained on roughly 100 million human cells, has outperformed existing models across cell annotation, perturbation prediction, gene function prediction, and mapping relationships between genes.5Nature Communications. CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells

These models effectively learn a compressed representation of cellular biology, capturing patterns in gene expression that encode information about cell identity, state, and behavior. The result is a kind of biological intuition embedded in software: given a new cell’s gene expression profile, the model can infer its type, predict its likely response to a drug, or identify which genes are working together in regulatory networks. This matters because the alternative, manually curating rules for each of these tasks, scales poorly when experiments produce millions of data points.

Accelerating Drug Discovery

Drug development is famously slow and expensive, and AI has been targeted at nearly every stage of the pipeline. On the target identification side, AI models predict protein structures, rank disease-relevant genes, and assess whether a particular protein can actually be bound by a small molecule drug. In virtual screening, they search through vast chemical libraries, predict how drugs interact with their targets, and optimize candidates for properties like solubility and metabolic stability.6npj precision oncology. AI accelerate the identification of druggable targets by 3D structures of proteins and compounds

On the chemistry side, AI-driven retrosynthesis prediction has become a real tool for medicinal chemists. These models work backward from a desired molecule, proposing routes to synthesize it from available starting materials. The accuracy and diversity of proposed synthetic routes have improved significantly, giving chemists more options and flagging pathways they might not have considered.7ACS Publications. Artificial Intelligence in Retrosynthesis Prediction and its Applications in Medicinal Chemistry The practical effect is that a chemist can spend less time figuring out how to make a molecule and more time deciding which molecules are worth making.

Seeing Molecules More Clearly

Cryo-electron microscopy has become a workhorse technique for determining the structures of large biological molecules, but the raw data it produces are noisy and require heavy computational processing. AI has penetrated this workflow at multiple steps. In particle picking, the laborious task of identifying individual protein images in a micrograph, deep learning models now handle what used to require extensive manual curation, with measurable improvements in the resolution of the resulting 3D maps.8Briefings in Bioinformatics. Artificial intelligence in cryo-EM protein particle picking: recent advances and remaining challenges

Beyond just picking particles, deep learning is helping researchers understand the flexibility of biological molecules. Proteins are not static; they shift between conformations, and a cryo-EM dataset often contains a mixture of shapes. Neural network architectures can now automatically sort out this structural heterogeneity and map particles onto a small set of conformational and compositional states, revealing the range of motions a protein complex actually undergoes.9Nature Methods. Deep learning-based mixed-dimensional Gaussian mixture model for characterizing variability in cryo-EM This moves cryo-EM from a technique that captures a single snapshot to one that tells a dynamic story.

Engineering Metabolism With Machine Learning

Synthetic biology aims to engineer living cells to produce useful molecules, from biofuels to pharmaceuticals. The challenge is that cellular metabolism is enormously complex: tweaking one gene often has cascading effects on dozens of pathways, and optimizing production of a target molecule requires balancing many competing cellular demands. Machine learning is increasingly used to navigate this complexity, guiding the design of metabolic pathways, predicting how changes in gene expression will affect output, and identifying rate-limiting bottlenecks.10PubMed Central. Machine learning for metabolic pathway optimization: A review

A concrete example: researchers applied machine learning to a combinatorial library of ribosome binding site sequences controlling a monoterpenoid production pathway in bacteria. By screening under 3% of the possible combinations, the algorithm identified optimal configurations that boosted production by over 60%.11PubMed. Machine Learning of Designed Translational Control Allows Predictive Pathway Optimization in Escherichia coli Tools like the Automated Recommendation Tool (ART) go a step further, combining machine learning with experimental design to recommend which experiments to run next, making the optimization cycle faster and more efficient even when starting data are limited.12PubMed Central. Machine Learning and Deep Learning in Synthetic Biology: Key Architectures, Applications, and Challenges

Self-Driving Laboratories

The logical endpoint of combining AI with automated experiments is the self-driving lab: a facility where robotic systems perform experiments, AI analyzes the results, and the AI then decides what to test next, all without human intervention between cycles. These systems are already operating in synthetic biology contexts, where the design-build-test-learn loop can be run continuously.13PubMed. Perspectives for self-driving labs in synthetic biology The appeal is speed. A human researcher might run one experiment per week and spend days interpreting results before planning the next round. A self-driving lab can complete that cycle in hours, accumulating data and refining its models far faster than any manual process.

The concept is still maturing. Most current self-driving labs operate within fairly narrow experimental domains, and they still require human oversight for quality control and to catch situations where the AI’s model of the system drifts away from reality. But the trajectory points toward increasingly autonomous experimentation, particularly in fields where the search space is too large for human intuition to navigate effectively.

Decoding the Dark Matter of Microbial Life

The vast majority of microbial species on Earth have never been grown in a lab. Metagenomics, which involves sequencing all the DNA from an environmental sample and computationally assembling genomes, is the primary window into this hidden diversity. The problem is that assembling genomes from a complex mixture of fragmented DNA is error-prone and computationally demanding. AI has begun to transform this workflow. Machine learning and deep learning enhance quality control, error correction, assembly, and the process of sorting DNA fragments into bins representing individual genomes. The result is more complete, more accurate reconstructions of microbial genomes from environments ranging from ocean sediments to the human gut.14PubMed. Artificial intelligence in metagenome-assembled genome reconstruction: Tools, pipelines, and future directions

Ecology, Agriculture, and the Physical World

AI is not confined to molecular biology. In ecology, acoustic sensors deployed in forests, grasslands, and marine environments generate massive streams of audio data that need to be classified by species. Machine learning models trained to recognize bird calls, bat echolocation, or whale song have become critical for processing these data, which would be impossible for human listeners to review manually.15Methods in Ecology and Evolution. Integrating AI models into ecological research workflows: The case of terrestrial bioacoustics In agriculture, machine learning connects genotype data to phenotype outcomes at various levels, from biochemical traits to crop yield, helping breeders identify promising genetic combinations without growing every possible cross in the field.16PubMed Central. Machine learning in plant science and plant breeding

Deep learning has also reshaped cellular image analysis. Models now classify cell types, segment structures within images, track objects over time, and even enhance microscopy images computationally, filling in details that the physical optics could not resolve.17PubMed Central. Deep learning for cellular image analysis These capabilities collectively mean that biologists working at scales from the subcellular to the ecosystem are generating data faster than they can interpret it, and AI is becoming the bottleneck-breaker.

How DNA Folds in Three Dimensions

Understanding how the genome is physically organized inside the nucleus turns out to matter as much as knowing the DNA sequence itself. Distant stretches of DNA can loop together, bringing regulatory elements into contact with genes they control, and disruptions to this three-dimensional architecture are linked to disease. Deep learning models trained on chromosome-conformation data have achieved remarkably high accuracy in predicting these long-range interactions directly from DNA sequence, suggesting that the sequence itself encodes much of the grammar governing how chromatin folds.18PubMed Central. Predicting 3D chromatin interactions from DNA sequence using Deep Learning More recent models integrate both DNA sequence and epigenetic signals, chemical modifications that affect how tightly DNA is packaged, to predict the three-dimensional interaction patterns with increasing precision.19Briefings in Bioinformatics. A review of deep learning models for the prediction of chromatin interactions with DNA and epigenomic profiles

This line of work is still relatively young compared to protein-structure prediction, but it addresses a gap that experimental methods struggle to fill. Measuring 3D genome contacts experimentally is expensive and typically gives an average across millions of cells. AI models that can predict these contacts from sequence alone open the possibility of studying how individual mutations might rearrange genome architecture, and by extension, how they contribute to diseases driven by misregulation of gene expression.

Reconstructing Evolutionary History

AI is also reaching backward in time. A new class of deep generative models can reconstruct the likely sequences of ancestral proteins by explicitly modeling the evolutionary process. One approach uses the branching structure of a phylogenetic tree as a mathematical prior, guiding the model to produce ancestral sequences that are consistent with what we know about how proteins evolve over time.20PubMed Central. Ancestral protein sequence reconstruction using a tree-structured Ornstein-Uhlenbeck variational autoencoder Reconstructed ancestral proteins can then be synthesized in the lab and tested, offering a window into the biochemistry of organisms that went extinct millions of years ago. Some of these resurrected proteins turn out to be more heat-stable or more broadly functional than their modern descendants, making them useful starting points for enzyme engineering.

The Biosecurity Question

The same generative AI tools that design therapeutic proteins or optimize metabolic pathways could, in principle, be used to design harmful biological agents. This dual-use risk is not hypothetical. Generative AI lowers the barrier to misuse by making it possible, at least in theory, to generate synthetic viral proteins or toxin-like sequences, and existing safety guardrails can often be circumvented through deceptive prompting or jailbreak techniques.21arXiv. Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity Large language models that assist with protein engineering can equally be prompted to generate predicted toxin-like sequences, a capability that most alignment efforts have not specifically addressed.22arXiv. A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

This is an area where the science is moving faster than governance. DNA synthesis companies already screen orders against databases of known pathogens, but AI-generated sequences might not match anything in those databases while still being dangerous. The research community is debating how to balance openness, which accelerates beneficial discovery, against the need for safeguards that prevent misuse. There are no easy answers, and the current consensus is that technical screening tools, responsible disclosure norms, and regulatory frameworks all need to evolve in parallel with the AI capabilities themselves.

Integrating It All With Multi-Omics

Biology does not come in neat single-data-type packages. A tumor, for example, can be characterized by its DNA mutations, its gene expression profile, its protein levels, its metabolite concentrations, and its appearance under a microscope. Each of these data types captures a different slice of the underlying biology, and combining them offers a more complete picture than any single measurement. Deep learning models designed for multi-omics integration can fuse molecular data like genomics, transcriptomics, and proteomics with imaging data to improve performance on tasks like predicting patient outcomes or classifying disease subtypes.23BioData Mining. Deep learning-based approaches for multi-omics data integration and analysis

Foundation models in bioinformatics are increasingly designed with this integration in mind. Rather than being specialists in one data type, they are trained across genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis simultaneously, learning representations that bridge these domains.24PubMed Central. Foundation models in bioinformatics The long-term vision is a unified computational framework where a researcher can ask questions that span molecular scales, from the effect of a single mutation on protein folding to its downstream consequences for cell behavior and tissue function, and get answers grounded in data rather than guesswork. That vision remains aspirational, but the individual pieces are falling into place faster than most biologists expected even five years ago.

Leave a Reply

Your email address will not be published. Required fields are marked *