Data in biology is any recorded observation, measurement, or sequence of information that scientists use to describe, analyze, and understand living systems. That covers an enormous range: the four-letter code of a genome, the three-dimensional coordinates of a protein, the GPS track of a migrating bird, the blood-glucose readings from a wearable sensor, even traces of DNA floating in a river. What makes biological data important is that living systems are staggeringly complex, and no single experiment or field notebook can capture enough of that complexity to reveal how organisms develop, adapt, get sick, or interact with their environments. The growing ability to collect, store, and analyze massive volumes of biological data has reshaped nearly every branch of the life sciences over the past two decades.
What Counts as Biological Data
If you picture a biology lab from fifty years ago, the data might have been handwritten tables of plant heights, sketches of cell structures under a microscope, or counts of beetles in a forest plot. All of that still qualifies. But today the term also includes DNA and RNA sequences, protein structures, medical images, electronic health records, satellite imagery of ecosystems, and continuous streams of sensor readings from devices attached to animals or patients. What ties these together is that each one represents a structured record of something happening in a living system, captured in a form that can be stored, shared, and reanalyzed.
One useful way to think about it is in layers. At the molecular level, you have genomic data (the sequence of bases in DNA), transcriptomic data (which genes are active in a given cell at a given moment), proteomic data (which proteins are present and in what quantities), and metabolomic data (the small molecules involved in metabolism). Zoom out and you reach the level of cells and tissues, where imaging data and histology slides live. Zoom out further and you get organism-level clinical data, population-level survey data, and ecosystem-level environmental monitoring. Each layer generates its own kind of data, and increasingly, the real insights come from connecting data across layers.
How DNA Sequencing Changed the Scale
Nothing expanded the meaning of “biological data” quite like the revolution in DNA sequencing. Modern sequencing machines can read millions of DNA or RNA fragments at the same time, producing terabytes of raw data in a single run.1PubMed. Next-generation sequencing: big data meets high performance computing That throughput, combined with a dramatic drop in cost to roughly a thousand dollars per human genome, has made population-scale genomic projects realistic for the first time.1PubMed. Next-generation sequencing: big data meets high performance computing
The practical payoff is broad. In cancer research, sequencing a tumor’s DNA helps doctors choose therapies matched to the specific mutations driving that patient’s disease. In rare-disease diagnosis, a single sequencing run can pinpoint a causative mutation that would have taken years of traditional testing to find. In infectious-disease surveillance, sequencing viral genomes lets public-health agencies track how a pathogen is evolving in near real time.2Pathology – Research and Practice. Next generation DNA sequencing data analysis and its application in clinical genomics None of these applications would work without the ability to generate, store, and computationally analyze huge volumes of sequence data.
Protein Structures and Molecular Shape
Knowing a gene’s sequence tells you the recipe for a protein, but it does not tell you the protein’s three-dimensional shape, and shape is what determines function. The Protein Data Bank, established in 1971 as the first open-access digital resource in biology, serves as the single global archive for experimentally determined structures of biological molecules.3PubMed Central. Protein Data Bank (PDB): The Single Global Macromolecular Structure Archive It now receives roughly 11,000 new entries per year, each one the product of painstaking laboratory work using techniques like X-ray crystallography or cryo-electron microscopy.3PubMed Central. Protein Data Bank (PDB): The Single Global Macromolecular Structure Archive
Even at that rate, experimental methods cannot keep up with the flood of newly discovered protein sequences. That gap is where artificial intelligence entered the picture. AlphaFold, a deep-learning system developed by DeepMind, uses physical and biological knowledge about how proteins fold to predict structures with high accuracy.4PubMed Central. Highly accurate protein structure prediction with AlphaFold The AlphaFold database now provides predicted structures for over 214 million protein sequences, and those predictions have been integrated into primary data resources like UniProt and the PDB itself.5Nucleic Acids Research. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences Drug designers, for example, rely on structural data to figure out where a small molecule might bind to a target protein, and having predicted structures for virtually every known protein massively accelerates that work.
Tracking Life in the Wild
Biological data is not confined to laboratories. Ecologists generate enormous datasets by monitoring organisms in their natural habitats, and the tools they use have become far more powerful in the past decade.
Environmental DNA, or eDNA, is one of the newer approaches. Every organism sheds genetic material into its surroundings through skin cells, mucus, feces, or decaying tissue. By collecting water or soil samples and sequencing the DNA fragments in them, researchers can detect which species are present without ever seeing or capturing an animal. The method is especially sensitive for identifying rare, endangered, and invasive species, and it works across aquatic, terrestrial, and even atmospheric environments.6PubMed Central. Environmental DNA (eDNA) Technology in Biodiversity and Ecosystem Health Research: Advances and Prospects For conservation biologists trying to survey remote or hard-to-access ecosystems, eDNA data can replace months of fieldwork with a few liters of filtered water.
Satellite telemetry provides a different kind of ecological data. GPS collars and tags transmit an animal’s location at regular intervals, allowing researchers to reconstruct movement paths, identify critical habitats, and study migration. A large survey of satellite-telemetry projects found that, on average, users obtained about 78% of the GPS fixes they expected from deployed units, though roughly a quarter of all deployments ended early because of technical failure and another large fraction ended due to animal-related issues like mortality or the animal removing the device.7PLOS ONE. Right on track? Performance of satellite telemetry in terrestrial wildlife research Those numbers matter because they shape how much usable data any tracking study actually produces, and they remind us that data collection in field biology is messier than in a genome sequencer.
Researchers also fuse tracking data with satellite imagery. One study combined GPS tracks of maned wolves with vegetation indices derived from multiple satellite sources and found that higher spatial resolution of the environmental layer mattered more than higher temporal resolution for predicting habitat selection.8Ecological Informatics. Multi-source data fusion of optical satellite imagery to characterize habitat selection from wildlife tracking data That finding has practical implications: ecologists often default to coarser but more frequently updated satellite products when, for some species, a sharper snapshot may be more informative.
Data at the Bedside
Clinical and biomedical data represent another vast category. Electronic health records contain a patient’s diagnoses, lab results, medications, imaging reports, and clinical notes. When combined with genomic data, those records can help clinicians improve the accuracy of diagnosis, optimize prescribing, and target risk-reduction strategies.9PubMed Central. It Is in Our DNA: Bringing Electronic Health Records and Genomic Data Together for Precision Medicine In practice, though, genomic data is still mostly treated as a one-off diagnostic tool and is not routinely woven into everyday clinical workflows.9PubMed Central. It Is in Our DNA: Bringing Electronic Health Records and Genomic Data Together for Precision Medicine
On a more personal scale, continuous glucose monitors have turned individual patients into walking data generators. These wearable sensors measure blood-sugar levels in near real time over several consecutive days, giving people with diabetes and their doctors a far richer picture of glucose fluctuations than a single finger-prick test ever could.10PubMed Central. Continuous Glucose Monitoring Sensors for Diabetes Management: A Review of Technologies and Applications The data stream from a CGM can reveal patterns, like a blood-sugar spike after a particular meal or a dangerous dip during sleep, that would otherwise go unnoticed.
Why Combining Different Data Types Matters
One of the most active frontiers in biology is multi-omics integration: taking genomic, transcriptomic, proteomic, and metabolomic data from the same biological system and analyzing them together. The logic is straightforward. A genome tells you what could happen in a cell, the transcriptome tells you what is being read off the genome right now, the proteome tells you what molecular machinery is actually present, and the metabolome tells you what the cell is producing and consuming. No single layer gives the full picture; combining them highlights the interrelationships among biomolecules and their functions.11PubMed Central. Multi-omics Data Integration, Interpretation, and Its Application
Multi-omics integration is a rapidly advancing area in computational biology, but it is also technically demanding. Each data type has its own formats, noise characteristics, and scales, and stitching them together requires sophisticated computational methods.12Briefings in Bioinformatics. A technical review of multi-omics data integration methods: from classical statistical to deep generative approaches When it works, though, the combined view can reveal disease mechanisms that would be invisible in any single data layer, which is why so much effort is going into making it routine.
Making Biological Data Findable and Reusable
Generating data is only half the battle. If a dataset sits on a single researcher’s hard drive in an undocumented format, it might as well not exist for the rest of the scientific community. The FAIR principles, short for Findable, Accessible, Interoperable, and Reusable, were developed as guidelines to ensure that published data can actually be found by other researchers, accessed without unnecessary barriers, combined with other datasets, and reused in new analyses.13PubMed Central. Enhancing Reuse of Data and Biological Material in Medical Research: From FAIR to FAIR-Health
In practice, adoption has been uneven. Many fields in biology, particularly those that rely heavily on manual data entry, have not fully embraced FAIR principles, and missing metadata remains a common problem that limits how much existing data can be reused.14PubMed Central. An approach to making life sciences FAIR-FAIR-DS as a tool for Aspergillus fumigatus Ontologies, which are standardized vocabularies that define biological terms and their relationships, play a key role in solving this. When two labs use the same controlled vocabulary to describe their experiments, their datasets become far easier to merge and compare.15PubMed. Beyond the data deluge: data integration and bio-ontologies Without that common language, even datasets on the same topic can be nearly impossible to combine.
The Computational Challenge
Biological data has a storage problem. A single sequencing run can produce several terabytes, and entire genomic projects generate petabytes. Storing, analyzing, and sharing datasets at that scale is a serious challenge for research communities that often lack dedicated computing infrastructure.16PubMed Central. Cloud Computing Enabled Big Multi-Omics Data Analytics Cloud computing has helped, giving smaller labs access to scalable storage and processing power without building their own data centers.17Journal of Industrial Information Integration. Cloud computing for storing and analyzing petabytes of genomic data
But the bottleneck is not just hardware. The analysis itself requires people who understand both biology and computational methods, and there are not enough of them. A survey of life science educators identified a lack of instructor and student background in data-science skills as one of the three biggest barriers to integrating data-science training into biology curricula.18BioScience. Data Science in Undergraduate Life Science Education: A Need for Instructor Skills Training A global review of bioinformatics training needs reached a similar conclusion: bioinformatics and biostatistics should be introduced at the undergraduate level and fully integrated into life science degree programs, rather than treated as optional add-ons for specialists.19PubMed Central. A global perspective on evolving bioinformatics and data science training needs Until that gap closes, many biologists will continue to produce datasets they are not fully equipped to analyze.
Data Sharing and Reproducibility
Science depends on the ability to check and reproduce findings, and that requires access to the underlying data. Journal data-sharing policies are the most direct lever, yet a review of 318 journals found that only about 12% required data sharing as a condition of publication. Another 9% required it without stating consequences for noncompliance, and roughly a third did not mention data sharing at all.20PeerJ. Reproducible and reusable research: are journal data sharing policies meeting the mark? Even where policies exist, compliance can be low. An analysis of data-availability statements in tens of thousands of papers published in a major open-access journal found that only about 20% indicated the data had actually been deposited in a repository.21PubMed Central. Incentivising research data sharing: a scoping review
The consequences of poor data sharing are real. When other researchers cannot access the data behind a published finding, they cannot verify it, build on it, or combine it with their own work. This slows science down and wastes resources. Funding agencies and publishers are slowly tightening requirements, but the culture shift is far from complete.
Ethics and Indigenous Data Sovereignty
As biological data collection accelerates, so do ethical questions about who controls that data and who benefits from it. Genomic data is especially sensitive because it can reveal information about health, ancestry, and familial relationships, not just for the individual whose DNA was sampled but for their relatives and community.
Indigenous data sovereignty, or IDSov, is a framework through which Indigenous Peoples assert the right to control data about their communities and lands.22PubMed. Indigenous Data Sovereignty in Genomics and Human Genetics: Genomic Equity and Justice for Indigenous Peoples Genomics has historically operated under an “openness” model where data is shared as freely as possible, and that model has driven enormous scientific progress. But open sharing can also enable exploitation. Genetic data from Indigenous communities has sometimes been used in ways those communities never consented to, or shared without benefit flowing back to them. A growing consensus holds that the definition of “open data” needs revision to ensure that openness does not become a barrier to engagement for communities that have been historically marginalized by research.23PubMed Central. Balancing openness with Indigenous data sovereignty: An opportunity to leave no one behind in the journey to sequence all of life Integrating IDSov principles into research design, rather than bolting them on as an afterthought, supports meaningful partnerships and helps ensure that the benefits of genomic research are more equitably distributed.22PubMed. Indigenous Data Sovereignty in Genomics and Human Genetics: Genomic Equity and Justice for Indigenous Peoples
Ancient DNA and the Deep Past
Biological data can also reach backward in time. Ancient DNA extracted from fossils, preserved bones, permafrost soil, and even cave sediments lets researchers reconstruct the genomes of organisms that lived thousands or even hundreds of thousands of years ago. Early results from ancient DNA studies revealed surprisingly complex population histories and showed that modern genetic-geography studies can give misleading impressions about even the recent evolutionary past.24PubMed Central. Ancient DNA
Ancient DNA data has been used to trace human migrations, identify past episodes of interbreeding between species, reveal how pathogens evolved alongside their hosts, and reconstruct the diets and ecosystems of extinct animals. The data is fragmentary and degraded compared to modern genomic sequences, which makes analysis harder, but the information it provides is genuinely irreplaceable. No living organism carries a complete record of its species’ past, and fossils alone cannot capture the genetic dimension. Ancient DNA fills that gap in ways that were unimaginable before sequencing became cheap and sensitive enough to work with badly damaged samples.
Digital Sequence Information and Synthetic Biology
Biology is increasingly a digital science. Scientists working in synthetic biology now routinely design organisms using digital sequence information rather than physical samples of genetic material.25PubMed Central. Digital Sequence Information and the Access and Benefit-Sharing Obligation of the Convention on Biological Diversity You can download a gene sequence from a public database, synthesize the corresponding DNA in a lab, and insert it into a new organism without ever handling the species the gene originally came from. In crop science, synthetic biology approaches guided by AI-driven analysis of large datasets aim to redesign biological components to enhance yields, nutrient absorption, and resilience to stress.
This shift raises new governance questions. The United Nations Convention on Biological Diversity has been working to develop standard policies on access and benefit-sharing for digital sequence information, because the old rules were built around the physical transfer of biological samples between countries.23PubMed Central. Balancing openness with Indigenous data sovereignty: An opportunity to leave no one behind in the journey to sequence all of life When a gene can be emailed as a text file and synthesized anywhere in the world, the line between “accessing a resource” and “accessing data about a resource” blurs in ways that existing legal frameworks were not designed to handle. How the international community resolves that tension will shape the future of biological data sharing for decades.