How Is DNA Digitized? From Sample to Digital File

DNA digitization converts the physical chemical sequence of a DNA molecule into a computer-readable text file, and the basic pipeline has three broad stages: extract the DNA from a biological sample, run it through a sequencing machine that reads the order of its four chemical bases (A, T, C, G), and then use software to translate the machine’s raw output into letters, quality scores, and standardized file formats. The process sounds straightforward, but each stage involves distinct chemistry, engineering, and computation that have evolved dramatically over the past few decades. What a researcher gets at the end is not a single neat text string but a layered set of files that encode both the sequence itself and how confident the system is about every single letter.

Getting DNA Out of the Sample

Before any machine can read DNA, you need naked strands of it, free from the proteins, membranes, and cellular debris that normally surround it. This is the extraction step, and it typically involves breaking open cells with a chemical or enzymatic solution, then pulling the DNA out of that mess using a filter, magnetic beads, or a silica column that DNA sticks to. Different extraction kits work better for different sample types. A study comparing sequencing outcomes found that bacterial pellets processed with kits like the DNeasy Blood & Tissue Kit or the ChargeSwitch gDNA Mini Bacteria Kit, alongside a separate Easy-DNA Kit protocol, all yielded usable DNA, but the choice of kit and its chemistry influenced downstream sequencing results.

Once the DNA is extracted, it usually needs to be fragmented into manageable pieces. Whole chromosomes are far too long for most sequencing instruments to handle in one go. Fragmentation can be mechanical (physically shearing the DNA with sound waves or pressure) or enzymatic (using proteins that cut the strands at semi-random spots). After fragmentation, short synthetic DNA sequences called adapters are attached to the ends of each fragment. These adapters serve as molecular handles that let the sequencing instrument grab, amplify, and identify each piece. The combination of fragmentation and adapter attachment is called library preparation, and different protocols suit different platforms. One comparison tested Illumina’s Nextera XT kit, which fragments and tags DNA enzymatically in a single step, against the TruSeq Nano kit, which uses mechanical shearing followed by separate adapter ligation, and found that protocol choice affected the final sequencing data.

1PubMed Central. Application of different DNA extraction procedures, library preparation protocols and sequencing platforms: impact on sequencing results

How Sequencing Machines Read the Bases

The heart of digitization is the sequencing instrument itself. There are several fundamentally different approaches, each with its own way of detecting which base sits at each position along a DNA strand. All of them ultimately produce a signal that software interprets as a string of A, T, C, and G letters, but the physics of that signal varies widely.

Sanger Sequencing

The original method for reading DNA, developed in the late 1970s, relies on making copies of a DNA strand that randomly terminate at each of the four bases. A DNA polymerase enzyme builds a new strand one base at a time, but the reaction mixture includes chemically modified bases that halt the copying process wherever they get incorporated. The result is a collection of fragments of every possible length, each ending with a labeled terminator base. By sorting those fragments by size, you can read off the sequence one letter at a time. Early versions used radioactive labels and gel slabs to separate fragments; later versions switched to fluorescent dyes and capillary tubes filled with gel. One study demonstrated that thermally stable polymerases could generate fragments and achieve single-base resolution estimated at over 450 bases from the starting point, with the fragments separated and detected by laser-induced fluorescence on a capillary column.

2PubMed. Sanger DNA-sequencing reactions performed in a solid-phase nanoreactor directly coupled to capillary gel electrophoresis

Sanger sequencing is highly accurate for individual reads but slow and expensive when you need to cover an entire genome. It dominated genomics for two decades and was the method behind the original Human Genome Project. Today it is mostly used for targeted verification, like confirming a single gene variant, rather than for whole-genome work.

Short-Read Sequencing

The technology that made large-scale genome digitization practical is short-read sequencing, most commonly associated with Illumina platforms. The basic idea is called sequencing by synthesis. Millions of DNA fragments from your library are spread across a glass surface called a flow cell, amplified into tiny clusters, and then read simultaneously. In each cycle, a single fluorescently labeled base is added to every growing strand. A camera captures which color lights up at each cluster position, the fluorescent label is chemically removed, and the next base is added. Researchers developed families of reversible terminators, modified nucleotides bearing a cleavable fluorescent dye, that enable this stop-and-read cycle. After each labeled base is incorporated and detected, the blocking group and dye are cleaved, allowing the next base to be added.

3PubMed Central. A new class of cleavable fluorescent nucleotides: synthesis and optimization as reversible terminators for DNA sequencing by synthesis

Each read is relatively short, typically around 150 to 300 bases, but because the instrument processes hundreds of millions of clusters at once, you can cover an entire human genome many times over in a single run lasting a day or two. The sheer parallelism is what makes it affordable enough for routine clinical use.

Long-Read Sequencing

Short reads struggle with repetitive stretches of DNA and complex structural rearrangements because the fragments are too small to span those regions. Long-read platforms solve this by reading much longer stretches, sometimes tens of thousands of bases or more, from a single molecule. Pacific Biosciences (PacBio) uses an approach where a single polymerase enzyme is fixed at the bottom of a tiny well called a zero-mode waveguide. A circular DNA template, called a SMRTbell, drifts into the well, and the polymerase begins copying it. As each fluorescently labeled base is incorporated, a brief light pulse identifies which base it is, and the instrument records that pulse in real time.

4Genomics, Proteomics & Bioinformatics. PacBio Sequencing and its Applications

Oxford Nanopore Technologies takes a completely different approach. Instead of watching a polymerase work, it threads a single DNA strand through a protein pore embedded in a membrane. As each base passes through the pore, it disrupts an electrical current in a characteristic way. The pattern of current changes is recorded and later decoded into a sequence of letters. This platform can produce reads exceeding a million bases in exceptional cases and has the additional advantage of being portable: the smallest device fits in a shirt pocket.

Turning Raw Signals Into Letters

No sequencing machine directly outputs a clean text file of A, T, C, and G. What it produces is raw signal data: fluorescence images in the case of Illumina, light-pulse traces for PacBio, or electrical current squiggles for Nanopore. Converting that raw signal into a sequence of nucleotide letters is called base calling, and it is a computationally demanding step that has become increasingly reliant on machine learning.

For Nanopore data in particular, the raw electrical signal is complex because multiple bases inside the pore influence the current at any given moment. Neural network base callers process this signal through layers that first recognize local patterns in the current (using convolutional layers) and then integrate those patterns over time (using recurrent layers) to assign probabilities to each possible base. A final decoding step converts those probabilities into the most likely DNA sequence.

5GigaScience. Chiron: translating nanopore raw signal directly into nucleotide sequence using deep learning

The quality of base calling has a direct impact on everything downstream. A benchmarking study of neural network base callers for Oxford Nanopore noted that this computational translation step is of critical importance to the platform’s overall performance.

6PubMed Central. Performance of neural network basecalling tools for Oxford Nanopore sequencing

Along with the base calls, the software assigns a quality score to each individual letter. These scores represent how confident the algorithm is that a particular base is correct. A higher score means fewer expected errors. These per-base quality scores travel alongside the sequence data into the first standard digital file the process creates.

The File Formats That Hold a Genome

The most common starting file format for digitized DNA is called FASTQ. Each entry in a FASTQ file contains four lines: a sequence identifier, the string of bases, a separator line, and a matching string of quality-score characters (each character encodes a numerical confidence value for the corresponding base). FASTQ emerged as the de facto standard for sharing sequencing reads, despite lacking a formal specification for years and existing in at least three incompatible variants developed by different platforms. A 2010 paper defined the format explicitly, covering the original Sanger standard and the Solexa/Illumina variants, and established conventions for converting between them.

7Nucleic Acids Research. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants

FASTQ files are large. A single human genome sequenced at standard depth can produce hundreds of gigabytes of raw FASTQ data. Once the reads have been aligned to a reference genome (more on that in a moment), the data is typically stored in SAM format, a tab-delimited text file that records where each read mapped and how well it matched. Because SAM files are enormous and human-readable text, they are almost always compressed into BAM files, the binary equivalent. BAM arranges similar content together to improve compression and is required by most analysis tools.

8Briefings in Bioinformatics. Sequence Alignment/Map format: a comprehensive review of approaches and applications

An even more compact format called CRAM takes the compression further by storing reads as differences from a reference genome rather than as full sequences. Since any two humans share about 99.9% of their DNA, most reads will match the reference almost perfectly, and only the tiny number of differences need to be explicitly recorded. The latest version, CRAM 3.1, is roughly 50 to 70 percent smaller than the equivalent BAM file for Illumina data, with more modest gains for long-read data because its signals carry more inherent randomness.

9PubMed Central. CRAM 3.1: advances in the CRAM file format

Aligning and Assembling the Pieces

A raw FASTQ file is like a box of millions of jigsaw pieces with no picture on the lid. Two main strategies exist for putting those pieces together: alignment and assembly.

Alignment, also called mapping, takes each short read and figures out where it belongs on a known reference genome. This is how most human genome projects work, because we already have a high-quality human reference to compare against. The challenge is doing this efficiently when you have billions of short reads and a reference that is over three billion bases long. One widely used tool, BWA, uses a data structure called the Burrows-Wheeler Transform to enable rapid searching, efficiently aligning short reads against a large reference while allowing for mismatches and small gaps.

10PubMed Central. Fast and accurate short read alignment with Burrows-Wheeler transform

Assembly is needed when there is no reference genome available, as is common when sequencing a newly discovered organism or a heavily rearranged cancer genome. Here, algorithms look for overlaps between reads and try to stitch them into longer contiguous sequences. Long reads have made assembly much more tractable because their greater length spans repetitive regions that would confuse short reads. One assembler, ABruijn, demonstrated how to combine graph-based approaches with overlap-layout strategies to produce accurate genome reconstructions even from long, error-prone reads.

11PubMed Central. Assembly of long error-prone reads using de Bruijn graphs

Finding the Differences That Matter

Once reads are aligned, the next analytical step is variant calling: identifying positions where the sequenced DNA differs from the reference genome. These differences might be single-letter changes, small insertions or deletions, or larger structural rearrangements. This is where digitized DNA becomes medically or scientifically actionable, because variants are what distinguish one person from another or reveal disease-associated mutations.

Multiple variant-calling tools exist, and their performance varies. A comparative analysis found that a deep-learning-based caller, DeepVariant, achieved the highest precision and overall balanced accuracy on a test chromosome, while other callers like Strelka2 excelled in precision on whole-genome analysis and Octopus showed superior recall, meaning it missed fewer true variants.

12PubMed Central. Variant calling in genomics: A comparative performance analysis and decision guide

The output of variant calling is typically a VCF (variant call format) file, which lists each detected difference along with its genomic position, the reference base, the alternative base, and various quality metrics. For clinical use, whole genome sequencing can detect a wider range of variants than older targeted methods like chromosomal microarrays. One prenatal study found that whole genome sequencing detected all the abnormalities that microarray analysis caught, plus additional cases involving small deletions and single-letter changes that the older method missed.

13PubMed. Whole genome sequencing vs chromosomal microarray analysis in prenatal diagnosis

How Accurate Is the Digital Copy?

Not all sequencing platforms produce equally faithful digital copies of the original DNA. Illumina short reads have very low per-base error rates, typically well below one percent. Long-read platforms historically had higher error rates, though both PacBio and Nanopore have improved substantially with newer chemistry and better base callers. Still, when researchers benchmarked assemblies from different platforms head-to-head, Illumina-only assemblies had the lowest error rates. The best Nanopore-only assemblies showed single-letter errors at frequencies about 41 percent higher and insertion/deletion errors about 157 percent higher than Illumina assemblies. PacBio assemblies landed in between, with single-letter errors about 12 percent higher and insertion/deletion errors about 78 percent higher than Illumina.

14PubMed Central. The long and short of it: benchmarking viromics using Illumina, Nanopore and PacBio sequencing technologies

These differences explain why many large-scale projects use a combination of platforms: short reads for high accuracy and long reads for structural completeness. The two complement each other, and hybrid assembly strategies that merge both data types can produce genomes that are both accurate and contiguous.

The Storage Challenge

Genomic data is growing faster than the technology to store it. A single deeply sequenced human genome generates on the order of 100 to 200 gigabytes of raw data, and population-scale projects involving hundreds of thousands of genomes push into the petabyte range. Specialized compression methods beyond standard BAM and CRAM have been developed to cope with this. Reference-based compression, which stores only the differences between a new genome and a reference, achieves excellent compression ratios because the vast majority of the sequence is identical. One such method, HRCM, was designed as a lossless approach for compressing both individual genomes and large collections.

15PubMed Central. HRCM: An Efficient Hybrid Referential Compression Method for Genomic Big Data

International sharing of genomic data adds another layer of complexity. Organizations like the Global Alliance for Genomics and Health (GA4GH) work to establish standards ensuring that data from different institutions and countries can be combined and reused, following principles of findability, accessibility, interoperability, and reusability.

16Cell Press. International federation of genomic medicine databases using GA4GH standards

Privacy Risks of a Digitized Genome

A digitized genome is arguably the most personally identifiable piece of data that exists. Unlike a password, you cannot change it if it is compromised. A 2018 study showed that it is possible to identify individuals by name from genetic data alone, by cross-referencing anonymous sequences against public genealogy databases. The researchers matched an anonymous participant’s genetic data to the GEDmatch database and identified her surname through relatives who had uploaded their own DNA. Estimates suggest that a genetic database covering just two percent of a target population would be enough to provide a third-cousin match to nearly any person in that population. As of 2018, the probability of such a match on GEDmatch was estimated at about 60 percent.

17PubMed Central. Assessing Privacy Vulnerabilities in Genetic Data Sets: Scoping Review

Anonymizing genetic data while keeping it scientifically useful remains an unsolved problem. Many privacy-enhancing approaches try to reduce the information content or restrict access so that only a minimal amount of data is shared, but there is no consensus on whether true anonymization of genetic data is even possible. This tension between data sharing for scientific progress and individual privacy runs through every stage of DNA digitization, from the moment a sample enters the lab to the moment the file lands on a server.

17PubMed Central. Assessing Privacy Vulnerabilities in Genetic Data Sets: Scoping Review

Going Beyond the Base Sequence

Standard sequencing digitizes the order of the four DNA bases, but DNA carries additional layers of information that sit on top of the base sequence. The most studied of these is methylation, a chemical modification where a small molecule is attached to certain bases (usually cytosine) without changing the underlying letter. Methylation patterns control which genes are active in a given cell and are altered in diseases like cancer. Digitizing these patterns requires extra chemistry. Bisulfite treatment chemically converts unmethylated cytosines into a different base while leaving methylated ones intact; sequencing the treated DNA then reveals which positions were methylated. Digital bisulfite sequencing refines this by analyzing individual DNA molecules, enabling researchers to see the methylation state of every position on a single strand rather than averaging across millions of cells.

18Nucleic Acids Research. DNA methylation analysis by digital bisulfite genomic sequencing and digital MethyLight

Another frontier is single-cell sequencing, which digitizes the genome or gene activity of individual cells rather than bulk tissue. A method called inDrops uses droplet microfluidics to encapsulate individual cells into tiny water-in-oil droplets, each containing a barcoded bead. Inside each droplet, the cell is broken open, its messenger RNA is tagged with a unique barcode, and the tagged molecules are later pooled and sequenced together. The barcode lets software assign each read back to its cell of origin. This approach can index over 15,000 cells per hour and has become essential for mapping the diversity of cell types in tissues.

19PubMed Central. Single-cell barcoding and sequencing using droplet microfluidics

Flipping the Script: DNA as a Hard Drive

An intriguing inversion of the digitization concept treats DNA itself as a storage medium for arbitrary digital data. Instead of reading biological DNA to create a computer file, you write a computer file into synthetic DNA. The idea exploits the fact that DNA is extraordinarily information-dense and can last for thousands of years under the right conditions. The basic scheme converts binary data (the 0s and 1s computers use) into sequences of A, T, C, and G, synthesizes those sequences as physical DNA strands, and later reads them back with a sequencer to recover the original file.

20CCF Transactions on High Performance Computing. Mainstream encoding–decoding methods of DNA data storage

The challenge is cost and speed: synthesizing DNA is still orders of magnitude slower and more expensive than writing to a silicon hard drive. Researchers have worked on improving the efficiency of the encoding step. One approach uses composite DNA letters, where each position in a strand contains a deliberate mixture of all four bases in a predetermined ratio rather than a single pure base. This allows more information to be packed into each synthesis cycle, and a proof-of-concept encoded 6.4 megabytes of data using about 20 percent fewer synthesis cycles per unit of data than previous methods.

21Nature Biotechnology. Data storage in DNA with fewer synthesis cycles using composite DNA letters

A newer technique called shortmer combinatorial encoding takes a different path, defining an extended alphabet where each “letter” represents not a single base but a subset of short DNA sequences mixed together at a given position. By analyzing which sequences appear at each position during readout, the system infers which letter was written. This combinatorial approach further increases the logical storage density while keeping error rates manageable.

22Scientific Reports. Efficient DNA-based data storage using shortmer combinatorial encoding

DNA data storage remains a research pursuit rather than a commercial product, but it highlights something worth appreciating: the same molecule that biology uses to store the instructions for life turns out to be an extraordinarily effective medium for storing anything, from text to video to operating systems. The bottleneck is no longer the concept but the economics of reading and writing at scale.