How to Build a DNA Molecule: A Step-by-Step Process

Building a DNA molecule from scratch is a layered process that moves from chemistry to biology, starting with the chemical synthesis of short single-stranded fragments and progressing through increasingly powerful assembly methods that stitch those fragments into genes, chromosomes, and even entire genomes. The foundational chemistry dates to the 1980s, but each layer of the process has seen dramatic improvements in the past few years, including an enzymatic alternative that may eventually replace the chemical approach altogether. What follows is the actual sequence of steps researchers and commercial providers use to construct DNA of any length, from a few dozen bases to over a million.

Synthesizing Short DNA Strands

Every synthetic DNA project begins with short single-stranded pieces called oligonucleotides, or “oligos.” The workhorse method for making these is phosphoramidite chemistry, a cyclic four-step process developed in the early 1980s that remains the industry standard. A starting nucleoside is anchored to a solid support bead inside a column, and then additional nucleotides are added one at a time in a repeating cycle. Each cycle involves removing a protective chemical group from the growing strand, coupling the next nucleotide, capping any strands that failed to couple (so they don’t cause errors later), and oxidizing the new linkage to stabilize it. The strand grows one base per cycle, from the 3′ end toward the 5′ end. Modern synthesizers can complete each cycle in minutes.

The catch is that no chemical reaction works perfectly every time. Even with coupling efficiencies above 99%, each cycle has a small failure rate, and those failures compound. After a hundred or so cycles, the cumulative effect means a shrinking fraction of the strands on the support are the correct, full-length product. For most commercial and research purposes, the practical ceiling has long been considered about 200 bases per oligo. Recent work has pushed dramatically beyond that boundary: one group reported directly synthesizing an 800-base green fluorescent protein gene and even a 1,728-base DNA polymerase gene on an automated synthesizer, demonstrating that careful optimization of the chemistry can stretch chemical synthesis far past its traditional limits.1Chemical Science. Long oligos: direct chemical synthesis of genes with up to 1728 nucleotides

A newer twist on the classic cycle uses dinucleotide building blocks instead of single nucleotides. By coupling two bases at a time, this approach cuts the number of cycles in half for a given oligo length, which reduces the accumulated error rate and speeds up the process.2Tetrahedron Letters. A practical dinucleotide phosphoramidite chemistry for de novo DNA synthesis via block coupling Modified building blocks also allow researchers to incorporate non-natural bases, unusual sugar backbones, or other chemical tweaks directly during synthesis.3PubMed. Novel phosphoramidite building blocks in synthesis and applications toward modified oligonucleotides

Enzymatic Synthesis as an Emerging Alternative

Phosphoramidite chemistry relies on harsh organic solvents and generates chemical waste, and it has that fundamental length ceiling. A different strategy sidesteps the chemistry altogether by using an engineered enzyme, terminal deoxynucleotidyl transferase (TdT), to add bases one at a time in water. In nature, TdT randomly adds nucleotides to DNA ends during immune-cell development. Researchers have re-engineered it to add only one specific base per step by attaching a reversible chemical “cap” to each nucleotide that blocks further additions until it is removed.

Getting this to work reliably took extensive protein engineering. One commercial effort put the enzyme through 32 rounds of laboratory evolution, introducing about 80 amino acid changes (roughly a fifth of the protein). The result was a version whose incorporation efficiency jumped around 200-fold to above 99% per step, with individual extension times dropping by more than 600-fold to about 90 seconds. The engineered enzyme is also far more heat-stable, gaining about 20°C in thermostability.4Nucleic Acids Research. Evolving a terminal deoxynucleotidyl transferase for commercial enzymatic DNA synthesis Enzymatic synthesis is not yet as mature as phosphoramidite chemistry, but it promises greener chemistry, aqueous conditions, and potentially longer reads as the technology matures.

Assembling Fragments into Genes

Even with the longest oligos available, most genes are thousands of bases long, and whole genomes stretch into millions. Bridging that gap requires assembly: taking dozens or hundreds of short oligos and stitching them together in the correct order. Several methods exist, each with trade-offs in complexity, fidelity, and scale.

Polymerase Cycling Assembly

The oldest and most straightforward approach is polymerase cycling assembly (PCA), which works like a modified PCR reaction. Overlapping oligos are mixed together, heated to separate any partial duplexes, cooled to let the overlapping ends anneal, and then extended by a DNA polymerase. Over repeated thermal cycles, the overlapping ends guide the fragments to hybridize in the correct order, and the polymerase fills in the gaps, progressively building up the full-length product.5PubMed Central. Polymerase cycling assembly In one large-scale application, unpurified 40-base oligos were built into fragments of 500 to 800 bases using automated PCR-based gene synthesis, which then served as building blocks for even larger assemblies.6PubMed Central. Total synthesis of long DNA sequences: synthesis of a contiguous 32-kb polyketide synthase gene cluster

Gibson Assembly

Gibson Assembly, introduced in 2009, takes a different approach. Instead of thermal cycling, it uses three enzymes working together in a single tube at a constant temperature: a 5′ exonuclease chews back the ends of each fragment to create single-stranded overhangs, a DNA polymerase fills in any remaining gaps, and a DNA ligase seals the nicks.7PubMed. Enzymatic assembly of DNA molecules up to several hundred kilobases The overlapping single-stranded regions anneal to their partners, and the result is a seamless join with no leftover sequences at the junctions. Gibson Assembly can join multiple fragments in a single reaction and has been used to assemble constructs up to several hundred kilobases in length.8PubMed. Assembling Multiple Fragments: The Gibson Assembly

Golden Gate Assembly

Golden Gate Assembly uses a clever trick with restriction enzymes. Special “Type IIS” enzymes cut DNA at a defined distance from their recognition site rather than within it. By designing fragments so that the enzyme’s recognition site sits just outside the desired junction, the enzyme generates custom sticky ends that snap together in a predetermined order. Because the cut removes the enzyme’s own recognition site from the final product, the assembly can be considered functionally “scarless,” meaning no unwanted sequences are left behind at the junctions.9PubMed Central. A User’s Guide to Golden Gate Cloning Methods and Standards The restriction enzyme and a DNA ligase work together in the same tube, cutting and joining in a cyclical reaction that drives the assembly toward the full-length product.

The efficiency of Golden Gate Assembly depends on the strength of the sticky-end overhangs. Using a 10-fragment assembly as a benchmark, researchers have shown that choosing strong overhang sequences increases the yield of correct assemblies, while weak overhangs produce more misassembled products.10Nucleic Acids Research. Enhanced Golden Gate Assembly: evaluating overhang strength for improved ligation efficiency Expanded versions of the method have also been developed to make it compatible with a wider range of vectors and restriction sites, broadening its flexibility for different projects.11PubMed Central. An efficient cloning method to expand vector and restriction site compatibility of Golden Gate Assembly

Catching and Fixing Errors

Chemical synthesis introduces errors at a low but consistent rate, typically one incorrect base per few hundred synthesized. When dozens of error-containing oligos are assembled into a gene, those mistakes stack up. A gene assembled without error correction might contain so many mutations that only a small fraction of the resulting clones encode a functional product.

One effective error-removal strategy uses enzymes that recognize mismatched base pairs in double-stranded DNA. After the gene is assembled, the product is incubated with a mismatch-cleaving endonuclease, which cuts the DNA backbone at or near mismatches. Exonucleases then chew away the exposed single-stranded overhangs, and the remaining correct fragments are re-amplified by PCR. In controlled experiments, treatment with T4 endonuclease VII followed by exonuclease digestion raised the fraction of correct clones more than 10-fold, from about 4% to about 47%. A similar treatment with E. coli endonuclease V achieved roughly an 8-fold improvement, from about 4% to about 31%.12Nucleic Acids Research. Removal of mismatched bases from synthetic genes by enzymatic mismatch cleavage The exonuclease step proved essential; without it, the mismatch-cleaving enzyme alone showed no improvement at all.

Commercial gene synthesis providers typically layer multiple rounds of error correction with sequence verification by Sanger sequencing or next-generation sequencing. Even so, customers often order several clones and sequence them all, selecting the one with the fewest remaining errors for downstream work.

Scaling Up to Chromosomes and Whole Genomes

Gibson Assembly and Golden Gate can handle constructs in the tens of kilobases, but building a chromosome or an entire genome requires assembling hundreds of kilobases to megabases of DNA. At that scale, researchers rely heavily on the natural recombination machinery of baker’s yeast, Saccharomyces cerevisiae. Yeast cells are remarkably good at joining pieces of DNA that share short overlapping sequences at their ends, a process called homologous recombination. Overlapping regions as short as 24 base pairs are enough for yeast to reliably stitch together up to a dozen fragments in a single transformation.13PubMed. High-Throughput DNA Assembly Using Yeast Homologous Recombination

This approach achieved one of the landmark results in synthetic biology: the one-step assembly of 25 overlapping DNA fragments into the complete 592-kilobase synthetic genome of Mycoplasma genitalium inside yeast cells.14PubMed Central. One-step assembly in yeast of 25 overlapping DNA fragments to form a complete synthetic Mycoplasma genitalium genome To go even larger, an iterative method called YLC-assembly exploits the natural yeast life cycle of mating and sporulation to repeatedly merge assembled pieces across generations, enabling the construction of megabase-scale DNA.15Nucleic Acids Research. YLC-assembly: large DNA assembly via yeast life cycle

The Sc2.0 project, an international effort to build a completely synthetic yeast genome, illustrates how these steps chain together at the grandest scale. The design involved reducing the native yeast genome by about 8%, systematically removing destabilizing elements and adding engineered flexibility. Individual synthetic chromosomes built by teams around the world will eventually be consolidated into a single strain, producing the first fully synthetic eukaryotic genome.16PubMed. Design of a synthetic yeast genome Early milestones included the construction of partially synthetic chromosome arms that functioned in living yeast with near-normal fitness.17PubMed Central. Synthetic chromosome arms function in yeast and generate phenotypic diversity by design

From Synthetic DNA to a Living Cell

A synthetic genome sitting inside a yeast cell is an impressive piece of molecular engineering, but making it function as the operating system of a living organism is a separate challenge. The field calls this step “booting up,” and it involves transplanting the synthetic genome into a recipient cell whose own genome has been removed or displaced.18PubMed Central. Synthetic chromosomes, genomes, viruses, and cells

The proof of concept came with genome transplantation in bacteria: the entire genome of one Mycoplasma species was extracted as naked DNA and introduced into cells of a different Mycoplasma species using a chemical transformation method. The recipient cells adopted the donor genome and essentially became the donor species, expressing the donor’s proteins and abandoning their original identity.19PubMed. Genome transplantation in bacteria: changing one species to another Building on these techniques, a later project created JCVI-syn3.0, a minimal bacterial cell running on a synthetic genome of about 531,000 base pairs encoding just 473 genes, the smallest set of genes known to support independent life.20PubMed Central. JCVI-syn3.0 – A synthetic genome stripped bare!

High-Throughput and Miniaturized Synthesis

Making a single gene is one thing. Making thousands at once is the challenge that drove the development of microfluidic and microchip-based synthesis platforms. Early work showed that gene synthesis could be performed inside tiny reaction chambers holding just 500 nanoliters, assembling genes up to 1 kilobase in length from oligo concentrations as low as 10 to 25 nanomolar, in parallel across multiple reactors on a single chip.21PubMed Central. Parallel gene synthesis in a microfluidic device Another approach used a “PicoArray” platform with massively parallel picoliter-scale reaction chambers that could synthesize and purify oligos simultaneously, designed from the start for multiplex gene assembly.22Nucleic Acids Research. Microfluidic PicoArray synthesis of oligodeoxynucleotides and simultaneous assembling of multiple DNA sequences

More recent microchip-based systems have pushed the throughput further while solving a persistent limitation: earlier chip-based methods produced oligos at very low concentrations, making downstream assembly of longer DNA difficult. A newer massive-in-parallel system uses an iterative cycle of identifying, sorting, synthesizing, and recycling on a microchip, boosting DNA product concentrations by four to six orders of magnitude compared to earlier chip methods and simplifying the path to large-scale gene assembly.23PubMed Central. Scaling DNA synthesis with a microchip-based massively parallel synthesis system

Quality Control for Long Synthetic DNA

Verifying that a synthetic DNA molecule has the intended sequence is straightforward for short oligos but becomes harder as the length increases. Standard sequencing methods like Sanger sequencing work well for fragments up to about 1,000 bases, and next-generation sequencing handles longer constructs. But for therapeutic applications, where regulatory agencies demand high confidence in every base, additional analytical methods are becoming important.

Mass spectrometry has emerged as a powerful quality-control tool, especially for shorter synthetic DNA and RNA used in drug development. For DNA molecules longer than about 60 bases, direct mass spectrometry has historically struggled. A newer approach addresses this by using restriction enzymes and short guide oligos to cut the long synthetic DNA into smaller pieces, then analyzing those pieces with high-resolution tandem mass spectrometry. This lets researchers confirm the sequence of synthetic DNA well beyond the 60-nucleotide threshold that used to be the practical limit for mass-spectrometry-based verification.24PubMed. Sequence confirmation of synthetic DNA exceeding 100 nucleotides using restriction enzyme mediated digestion combined with high-resolution tandem mass spectrometry

Biosecurity Screening

The ability to build DNA to order raises obvious safety concerns. A customer could, in principle, order the sequence of a dangerous pathogen. To address this, most commercial DNA synthesis providers run screening at two levels. They verify the identity of the purchaser and their affiliation with a legitimate institution. Separately, the requested sequence is computationally compared against databases of pathogenic sequences of concern. Any sequence that is a close match to something dangerous gets flagged for human review. The provider then follows up with the customer to verify that they have a legitimate use, proper biosafety protocols, and any required export licenses. If the follow-up does not resolve the concern, the order is terminated.25PubMed Central. Screening State of Play: The Biosecurity Practices of Synthetic DNA Providers Customer screening and sequence screening typically run in parallel, so neither step delays the other for routine orders.

Storing Digital Information in Synthetic DNA

One application that stretches the idea of “building a DNA molecule” in an unexpected direction is DNA data storage. Because DNA encodes information in a four-letter alphabet that is extraordinarily stable and dense, researchers have explored encoding digital files (text, images, software) into synthetic DNA sequences, storing the physical molecules, and then reading the data back by sequencing.

In one demonstration, a 1,854-base-pair DNA fragment encoding digital information was synthesized and then sequenced using long-read nanopore technology. Out of over 100,000 reads, a consensus sequence was assembled that showed 100% identity to the designed sequence, meaning the digital information could be recovered perfectly without any additional error correction.26PubMed Central. Storing Digital Information in Long-Read DNA Detailed protocols for translating digital files into DNA sequences, physically handling and storing the molecules, and retrieving the information by sequencing have been published as step-by-step guides.27Nature Protocols. Reading and writing digital data in DNA The economics are not yet competitive with conventional storage media for everyday use, but DNA’s theoretical storage density and longevity (intact DNA has been recovered from samples thousands of years old) keep this a lively area of research. The synthesis and assembly steps are the same ones used to build genes; the difference is simply what the sequence encodes.