Genomic Data Storage: How It Works and Why It Matters

Every time a human genome is sequenced, the process generates roughly 100 to 200 gigabytes of raw data, and the number of genomes being sequenced worldwide is growing fast enough that genomics may soon rival astronomy as the largest generator of scientific data on Earth. Storing, organizing, securing, and making sense of all that information is a quietly enormous challenge that touches everything from the price of a clinical test to the privacy of your medical identity. The problem is not just “where do we put all these files” but how to keep them useful, affordable, and safe for decades.

Why Genomic Data Gets So Big So Fast

Modern sequencing machines work by chopping DNA into billions of short fragments and reading each one. A single sequencing run can produce several terabytes of raw output, most of it in the form of text-like files recording each fragment and its quality score.1PubMed. Next-generation sequencing: big data meets high performance computing A whole-genome sequence for one person, once aligned and processed, typically lands around 100 to 120 gigabytes, but the intermediate files created along the way can be much larger. Multiply that by the thousands or millions of genomes now being sequenced for research, clinical diagnosis, and population health programs, and the numbers become staggering.

A widely cited comparison projected that by 2025, genomics would rival or exceed astronomy, YouTube, and Twitter in data acquisition, storage, distribution, and analysis demands.2PubMed Central. Big Data: Astronomical or Genomical? – Section: Abstract That projection has largely held up. National biobanks in the UK, US, China, and elsewhere are each aiming to sequence hundreds of thousands to millions of participants. Clinical labs running targeted gene panels produce less data per patient, but they run constantly and cumulatively contribute enormous volumes. The data never stops arriving, and unlike a video that can simply be deleted after its relevance fades, genomic data often needs to be kept for the lifetime of the patient or longer.

File Formats and Compression

Genomic data lives in a handful of specialized file formats, each suited to a different stage of analysis. The raw output from a sequencer typically arrives as FASTQ files, which store the sequence of each DNA fragment alongside a quality score for every single letter. Once those fragments are mapped to a reference genome, they become alignment files stored as BAM (a binary format) or its more compressed cousin, CRAM. Variant call files (VCF) then record only the places where an individual’s DNA differs from the reference. Each step shrinks the file size, but researchers often want to keep every stage in case they need to go back and reprocess from scratch.

Compression matters enormously when you are storing millions of these files. The CRAM 3.1 format achieves files that are roughly 50 to 70 percent smaller than the equivalent BAM file for standard short-read sequencing data, though the savings are more modest for long-read technologies that carry noisier signals.3PubMed Central. CRAM 3.1: advances in the CRAM file format – Section: Abstract Beyond format-level improvements, referential compression algorithms can shrink genome files further by encoding each genome as a set of differences from a shared reference, rather than storing every letter independently.4Bioinformatics. High efficiency referential genome compression algorithm Even the choice of general-purpose compression tool matters: one lab found that switching from gzip to the zstandard algorithm for their raw FASTQ files saved more than 20 percent of storage space overall.5PubMed Central. The small genomics lab experience optimizing data cold storage

These percentages sound modest in isolation, but applied across a biobank holding petabytes of data, a 20 percent reduction translates to millions saved on hard drives, cloud fees, and electricity. Format choice is one of the most consequential decisions a genomics lab makes, and it is often made once, early on, without revisiting it for years.

Cloud Storage and the Economics of Hot Versus Cold

Most genomic data today ends up in cloud storage, whether through a major provider or a national research infrastructure. The cost of keeping that data accessible varies wildly depending on how quickly you need to retrieve it. Standard cloud storage tiers, sometimes called “hot” storage, let you pull a file within seconds. Deep-archive or “cold” tiers cost a fraction of the price but require hours or even a full day to retrieve a file. For genomic data, where most files sit untouched for months or years between analyses, cold storage is often the rational choice.

The cost difference is dramatic. Storing a single whole genome of about 120 gigabytes in a deep-archive tier costs roughly $14 over ten years. The same file in a standard cloud storage tier would cost over $300 across the same period, more than twenty times as much.6PubMed Central. Practical estimation of cloud storage costs for clinical genomic data – Section: Results Smaller files from targeted gene-panel tests or exomes are even cheaper to archive, costing just pennies per year in deep-archive tiers. In the context of what it costs to actually run the sequencing, including reagents, lab staff, and equipment, the marginal storage cost at archival rates becomes almost negligible.

The practical challenge is designing a migration strategy. Data starts in hot storage because bioinformaticians need fast access during the initial analysis. But once that analysis is done, a smart approach moves the files into cold storage within weeks or months, keeping retrieval possible for future reanalysis while cutting ongoing costs. One policy analysis for a proposed UK newborn screening program estimated that automatically migrating files to deep-archive tiers three months after screening could cut lifetime storage costs by about 91 percent, to roughly £18 per child, with 12 to 24 hour retrieval delays that are clinically acceptable for the kinds of reanalysis that come up years later.7Journal of Medical Genetics. Sequencing every UK newborn: why cold storage economics should shape policy Labs that leave everything in hot storage by default often discover a ballooning bill that feels surprising only because nobody modeled the trajectory.

Getting Genomic Data Into Medical Records

Sequencing a patient’s genome is useful only if the results reach the clinician making treatment decisions. In practice, integrating genomic data with electronic health records has been one of the slowest-moving parts of the field. The fundamental tension is that a genomic dataset is not like a blood-test result. It is enormous, it is interpretive (variants need to be classified), and its meaning can change over time as science advances.

A widely referenced framework for this integration laid out seven functional requirements, including the need to keep raw molecular observations separate from clinical interpretations, support lossless compression, anticipate changes in how human variation is understood, and simultaneously serve both human-readable and machine-readable formats so that clinical decision-support tools can fire automatically.8Journal of Biomedical Informatics. Technical desiderata for the integration of genomic data into Electronic Health Records – Section: Results That last point is especially important: a clinician seeing a patient with a newly discovered drug-gene interaction needs the EHR to flag it in real time, not just store a static PDF of an old report.

Efforts to connect genomic knowledge bases with EHR systems through standards like HL7’s “infobutton” protocol have shown promise but also exposed gaps, particularly in the adoption of shared terminology standards between genomic databases and clinical systems.9PubMed Central. Integrating Genomic Resources with Electronic Health Records using the HL7 Infobutton Standard – Section: RESULTS Without a common vocabulary, a genomic variant recorded as one identifier in a research database may not match the nomenclature used by the hospital’s EHR, making automated alerts impossible. This is a mundane-sounding problem that has genuinely life-or-death implications: if a pharmacogenomic result sits unrecognized in a patient file, a drug that should have been avoided gets prescribed anyway.

Privacy Risks That Come With Stored Genomes

Genomic data is uniquely sensitive because it is both personally identifying and immutable. You can change a compromised password; you cannot change your DNA. And the re-identification risk is real, not hypothetical. Researchers have demonstrated that uploading a set of common genetic markers from an anonymous study participant to a public genetic genealogy website was enough to identify the participant’s surname through relative matches.10PubMed Central. Assessing Privacy Vulnerabilities in Genetic Data Sets: Scoping Review – Section: SNPs The explosion of consumer genetic testing databases makes this kind of attack easier with each passing year, because the chances of finding a relative in those databases keep growing.

Traditional de-identification methods, like stripping names and dates, are insufficient for genomic data because the genome itself is an identifier. This has pushed researchers toward cryptographic approaches, particularly homomorphic encryption, which allows computations to be performed on encrypted data without ever decrypting it. One practical implementation showed that a full genome-wide association study on a real dataset of more than 25,000 individuals could be conducted with all individual data remaining encrypted throughout the analysis.11PubMed Central. Secure large-scale genome-wide association studies using homomorphic encryption – Section: Abstract Other groups have applied the same principle to allow private genome analysis in untrusted cloud environments, where the cloud provider never gains access to the decryption key.12PubMed Central. Private genome analysis through homomorphic encryption – Section: Abstract

These approaches are computationally expensive, which means they add cost and processing time. But for large-scale research where genomic data from thousands of people must be analyzed across institutions, they may be the only realistic way to share data without exposing it. Federated approaches, where the analysis travels to the data rather than the data traveling to the analyst, are also being combined with homomorphic encryption for added protection.13IACR Cryptology ePrint Archive. Privacy-Preserving Federated Inference for Genomic Analysis with Homomorphic Encryption

Regulatory Gray Zones

In the United States, HIPAA treats genetic information as protected health information, meaning it is subject to the same privacy rules as other medical data. But applying HIPAA’s “minimum necessary” standard to genomic files creates confusion that regulators have not fully resolved. The minimum necessary rule says that any use or disclosure of health data should involve only the smallest amount reasonably needed for the purpose. For a lab report showing a single blood-sugar reading, that standard is easy to apply. For a BAM file or a VCF file containing a patient’s entire set of genetic variants, it is far less clear: is it ever “reasonably necessary” to share the whole file, or should only specific variants be disclosed?14PubMed Central. Impact of HIPAA’s Minimum Necessary Standard on Genomic Data Sharing – Section: Introduction Regulators have not issued guidance to help genomic testing laboratories answer that question, leaving labs to make their own judgment calls and opening potential legal exposure regardless of which direction they choose.

Outside the US, the landscape is patchwork. The EU’s GDPR classifies genetic data as a “special category” requiring explicit consent for processing, but the interaction between GDPR and large-scale research biobanks is still being worked out country by country. National newborn screening programs add another layer: if a government sequences every baby’s genome at birth, the retention period, access rules, and consent framework for that data become policy decisions with generational implications.

Indigenous Data Sovereignty

One of the more consequential governance discussions around genomic data storage involves Indigenous communities. Historically, genetic data from Indigenous Peoples has been collected, stored, and analyzed by outside researchers with little or no community control over how it is used, shared, or interpreted. The concept of Indigenous data sovereignty provides a framework through which communities can assert their right to govern data about themselves and their lands.15PubMed. Indigenous Data Sovereignty in Genomics and Human Genetics: Genomic Equity and Justice for Indigenous Peoples

The practical implications for genomic data storage are significant. Principles like the CARE guidelines (Collective benefit, Authority to control, Responsibility, Ethics) call for data to be stored in ways that give communities meaningful authority over access, rather than defaulting to open-access repositories that treat all data as a global commons. A paper examining large biodiversity sequencing initiatives argued that the traditional scientific ideal of fully open data has unintentionally created barriers to engagement for Indigenous Peoples, because “openness” as commonly practiced can mean losing control of culturally significant genetic information.16PubMed Central. Balancing openness with Indigenous data sovereignty: An opportunity to leave no one behind in the journey to sequence all of life Some projects have responded by implementing tiered access systems or community-governed data repositories, but adoption is uneven and the technical infrastructure for this kind of granular control is still maturing.

The Carbon Cost of Storing and Analyzing Genomes

Large-scale genomic computation has a measurable environmental footprint, and it is one that researchers are just beginning to quantify. Running a single genome-wide association study on a dataset the size of the UK Biobank (about 500,000 individuals) was estimated to emit roughly 17 kilograms of COâ‚‚ equivalent using an older version of a common analysis tool, dropping to about 5 kilograms with a newer, more efficient version of the same software — a 73 percent reduction achieved purely through algorithmic improvement.17Molecular Biology and Evolution. The Carbon Footprint of Bioinformatics – Section: Results A single study’s emissions may sound modest, but thousands of such analyses run every year across the field, and the carbon cost of long-term storage compounds over time as more data accumulates.

This is part of why the hot-versus-cold storage debate matters beyond just budget. Hot storage tiers run on servers that are powered around the clock. Archival tiers, which keep data on media that can be partially powered down, consume substantially less energy. The newborn-screening policy paper mentioned earlier specifically cited net-zero targets alongside budget constraints as a reason to push data into cold storage quickly. For a national program sequencing every baby born in a country, the aggregate energy difference between hot and cold storage over decades is not trivial.

Biobank IT Architecture

Large-scale biobanks, the repositories that pair stored biological samples with genomic and clinical data, need an IT backbone capable of handling complexity that most database systems are not designed for. A biobank is not just a warehouse for genomes. It must link genetic data to clinical records, track the physical samples themselves, manage consent records, handle queries from multiple research groups with different access permissions, and do all of this while maintaining participant privacy. A federated architecture, where clinical annotations, anonymization services, and sample-management systems remain separate but coordinated, has emerged as a common model, with flexibility and scalability identified as critical requirements given that a single hospital may need to support many overlapping studies simultaneously.18PubMed Central. IT Infrastructure Components for Biobanking – Section: Conclusion

The real engineering difficulty is in the metadata: knowing which sample came from which patient, at what time, under what consent terms, processed by which lab protocol, and sequenced on which instrument, then making all of that queryable for researchers who may want to filter by disease status, ancestry, tissue type, or age at collection. A genome file without its metadata is practically useless for research, and poor metadata is one of the most commonly cited barriers to reusing genomic data across studies.

Reuse, Interoperability, and the Metadata Problem

Genomic data has enormous potential for reuse. A genome sequenced to diagnose a rare disease in one child might later contribute to a study on drug metabolism, cancer risk, or population history. But reuse depends on the data being findable, accessible, and annotated well enough that a new research group can work with it. A review focused on agricultural genomics, where many of the same problems exist, identified deficiencies in metadata standards, data interoperability, availability, and equity as persistent barriers to getting real value out of stored data.19PubMed Central. Data reuse in agricultural genomics research: challenges and recommendations – Section: Abstract These challenges are not unique to agriculture; they pervade human genomics as well, especially when data originates from many labs using different protocols, reference genomes, and annotation practices.

Efforts like the Global Alliance for Genomics and Health (GA4GH) have developed standards and application programming interfaces aimed at making genomic datasets interoperable across institutions and borders. But standardization is a social and institutional challenge as much as a technical one. Researchers who already have working pipelines are reluctant to overhaul them for the sake of compatibility with someone else’s system, and funding agencies have only recently begun requiring that deposited data meet specific formatting and metadata standards. The result is that much of the world’s stored genomic data is technically accessible but practically difficult to use without substantial curation effort.

DNA as a Storage Medium

There is an ironic twist in the genomic storage story: some researchers are exploring the use of synthetic DNA itself as a medium for storing digital information — not just genomic data, but any data at all. The idea exploits DNA’s extraordinary information density (a gram of DNA can theoretically encode hundreds of petabytes) and its durability (DNA recovered from ancient remains has been readable after tens of thousands of years under the right conditions). Various encoding algorithms translate digital ones and zeros into sequences of the four DNA bases, and the resulting molecules can be stored in a tiny volume at room temperature or below.20PubMed Central. Design considerations for advancing data storage with synthetic DNA for long-term archiving – Section: Abstract

The practical obstacles remain substantial. Writing DNA synthetically is slow and expensive compared to magnetic or optical media. Reading it back requires sequencing, which introduces errors. One photolithographic synthesis approach measured error rates of about 2.6 percent for substitutions, 6.2 percent for deletions, and 5.7 percent for insertions, considerably higher than the roughly 0.5 percent substitution rate achieved by earlier, more expensive methods.21PubMed Central. Low cost DNA data storage using photolithographic synthesis and advanced information reconstruction and error correction – Section: Maskless array DNA synthesis performance Error-correction codes can compensate, but they add overhead that reduces the effective storage density. For now, DNA data storage is a proof of concept rather than a competitive technology for everyday use, though its potential for ultra-dense, ultra-long-term archiving keeps attracting research investment.

Blockchain and Decentralized Approaches

A growing number of projects are experimenting with decentralized architectures for genomic data management. One recent proposal combines blockchain technology with decentralized identity standards and the InterPlanetary File System (IPFS) to create a model where the patient, not the hospital or the research institution, controls who can access their genomic data.22PubMed. A Novel Blockchain-Based Model for Secure Genomic Data Management In this kind of system, the genomic file itself is stored in a distributed way, and the blockchain acts as an immutable access log, recording every time someone is granted or revoked permission to see the data.

The appeal is real: traditional centralized databases create single points of failure and single points of trust, and many patients are understandably uneasy about a hospital or a tech company holding their most intimate biological information. But decentralized systems bring their own challenges. IPFS and similar distributed storage networks rely on nodes being willing to host the data, which creates availability problems if nodes go offline. Blockchain transactions have their own energy costs. And regulatory compliance becomes complicated when data is not physically located in a single jurisdiction. These systems are still in early prototype stages for genomic data, and whether they can scale to handle the volume of a national sequencing program remains to be seen.

Cloud Computing for Large-Scale Genomic Analysis

Storage and analysis are deeply intertwined. Genomic data often needs to be processed where it is stored, because moving terabytes of data across networks is slow and expensive. Cloud computing has become the default platform for large-scale genomic analysis partly because of this: the data and the computation sit in the same data center. An early demonstration of cloud-based comparative genomics ran more than 300,000 analysis processes across 100 compute nodes in about 70 hours for a total cost of roughly $6,300, showing that massive parallelization in the cloud could make analyses feasible that would have been impractical on local hardware.23PubMed Central. Cloud computing for comparative genomics – Section: RESULTS

Cloud costs have dropped substantially since that early study, and the range of available tools has expanded. But the basic architectural insight remains: for genomic data, storage and compute are not separate budget lines. A lab that chooses its cloud provider based solely on storage pricing may find that analysis costs, egress fees for moving data out, or the lack of pre-installed bioinformatics tools wipe out whatever savings they thought they had. Choosing a genomic storage strategy without thinking about where and how the data will be analyzed is like choosing a warehouse location without considering shipping routes.