Every time one of your cells divides, it first copies all of its DNA in three broad stages: initiation, where the cell identifies where to start and pries the double helix apart; elongation, where new DNA strands are assembled by reading each old strand as a template; and termination, where converging replication machinery meets, finishes the job, and is dismantled. These stages involve dozens of specialized proteins working in tight coordination, and the process is so accurate that it typically introduces fewer than one mistake per complete copy of the human genome.
Initiation Starts at Specific Locations
Replication does not start randomly. Cells mark specific stretches of DNA called replication origins where the process is allowed to begin. In bacteria like E. coli, there is a single origin on the circular chromosome, known as oriC. A protein called DnaA, fueled by ATP, binds to specific sequence motifs within this origin, assembles into a multi-protein complex, and forces open a nearby AT-rich region of the double helix. AT-rich regions are easier to separate because the bonds holding A-T base pairs together are weaker than those holding G-C pairs. Once the strands are pried apart, DnaA helps load the replicative helicase (DnaB), which will go on to unwind DNA ahead of the copying machinery.
The details of how DnaA opens the origin are more intricate than a simple “bind and pull.” In E. coli, the origin has two functional halves. DnaA molecules assemble on both halves but adopt different shapes depending on their position. Specific amino acid residues in DnaA are required for the left-half sub-complex to become competent for unwinding, and a DNA-bending protein called IHF helps recruit the separated single strands toward the DnaA complex on the opposite side, promoting helicase loading.
Human cells face a bigger challenge: their genomes are roughly a thousand times larger than a bacterial genome, so relying on a single origin would make replication impossibly slow. Instead, human chromosomes contain tens of thousands of replication origins. The cell licenses these origins during a specific window before DNA synthesis begins, using a complex called ORC (Origin Recognition Complex) along with a ring-shaped helicase called MCM2-7. ORC loads two copies of MCM2-7 onto each origin in a head-to-head arrangement, forming what is called a double hexamer. Once loading is complete, ORC is displaced from efficient origins in a loading-dependent manner, which helps prevent the same origin from being re-licensed prematurely.
Priming the Pump
DNA polymerases, the enzymes that actually build new DNA strands, have a frustrating limitation: they cannot start a new strand from scratch. They can only add nucleotides to an existing strand. This means every new stretch of DNA synthesis needs a short starter piece, called a primer, laid down first. That job falls to an enzyme called primase, which synthesizes short RNA molecules that serve as starting points for DNA polymerases.
On the leading strand (described below), a single primer is enough to get synthesis rolling in one continuous stretch. On the lagging strand, primase has to lay down a new RNA primer repeatedly, every few hundred nucleotides or so, because that strand is synthesized in short, discontinuous pieces. Without primase, no replication fork could advance.
Elongation and the Two-Strand Problem
Once the helicase unwinds the double helix and primers are in place, elongation begins. This is the main copying phase. DNA polymerase reads the exposed template strand and adds matching nucleotides one at a time, building the new complementary strand. But there is a fundamental asymmetry at every replication fork that makes this process surprisingly complicated.
DNA polymerase can only build a new strand in one direction (conventionally called 5′ to 3′). Since the two strands of the double helix run in opposite directions, only one strand, the leading strand, can be copied continuously in the same direction the helicase is traveling. The other strand, the lagging strand, has to be copied in the opposite direction, away from the fork, in short fragments. These fragments, called Okazaki fragments, are each about 100 to 200 nucleotides long in eukaryotes and somewhat longer in bacteria.
The leading and lagging strand polymerases do not work independently. Studies using the bacteriophage T7 replication system showed that leading and lagging strand synthesis are coupled at the fork, meaning the activities of both polymerases are coordinated. When the T7 helicase-primase pauses to synthesize a new lagging-strand primer, leading strand synthesis slows down as well. The replication proteins appear to be recycled rather than falling off and being replaced each time a new Okazaki fragment begins.
How Okazaki Fragments Become a Continuous Strand
Each Okazaki fragment starts with an RNA primer that eventually must be removed and replaced with DNA. This process, called Okazaki fragment maturation, requires several enzymes working in sequence. In eukaryotic cells, there appear to be two parallel pathways for getting the job done. One involves nucleases called Dna2 and Fen1 that clip away flaps of displaced RNA-DNA primer material. The other involves exonucleases RNase H2 and Exo1 that digest the primer from its end.
Once the RNA is removed and replaced with DNA, the remaining nick between adjacent fragments is sealed by DNA ligase. In eukaryotes, this enzyme (DNA ligase 1) is remarkably picky: it strongly discriminates against sealing nicks where the bases are incorrectly paired, acting as a final quality-control checkpoint before the lagging strand becomes a continuous, finished molecule.
Keeping the Polymerase on Track
DNA polymerases by themselves are not especially clingy. Left on their own, they tend to fall off the template after copying a short stretch. To keep them attached for the thousands of nucleotides they need to copy without detaching, cells use ring-shaped proteins called sliding clamps. In bacteria, the clamp is called the beta clamp; in human cells, it is called PCNA. These rings encircle the double-stranded DNA and tether the polymerase to it, functioning as a nearly frictionless bearing that allows fast, processive synthesis. They essentially “water skate” along the DNA surface, held close but free to slide.
Getting the clamp onto the DNA in the first place requires a separate machine called a clamp loader, which uses ATP to open the ring, thread it around DNA at a primer-template junction, and snap it shut. Once in place, the clamp does more than just hold the polymerase on. It also serves as an attachment point for other enzymes involved in repair, maturation, and checkpoint signaling, making it a central organizing hub at the replication fork.
Managing Tension Ahead of and Behind the Fork
As the helicase unwinds DNA at the fork, it creates a problem: the DNA ahead of the fork becomes overwound, building up positive supercoiling, while the DNA behind the fork can become tangled in structures called precatenanes. If this tension is not relieved, the fork will stall. Topoisomerases solve this problem by cutting one or both strands of DNA, allowing the helix to relax, and then resealing the break.
In human cells, topoisomerase IIα rapidly relaxes positively supercoiled DNA, suggesting it plays a key role in relieving torsional stress ahead of the advancing fork. Bacteria rely on a specialized topoisomerase called DNA gyrase, which actively introduces negative supercoils to counteract the positive ones generated by helicase activity. Without topoisomerases, replication forks grind to a halt within seconds.
Meanwhile, the single-stranded DNA exposed at the fork is vulnerable. It can form hairpins or other secondary structures, and it is susceptible to damage by nucleases. A protein called RPA (Replication Protein A) in eukaryotes coats this exposed single-stranded DNA, protecting it from degradation and keeping it in a straightened-out conformation so that polymerases can read it properly. RPA clusters into tetramers that efficiently coat stretches of single-stranded DNA created by the advancing fork, and its role goes beyond simple protection: it also participates in signaling and coordinating repair pathways.
Error Correction During Copying
Given that a human cell copies roughly six billion base pairs every time it divides, accuracy is critical. The replication machinery achieves a final error rate of less than one mistake per genome duplication through three layered defenses.
The first layer is base selectivity: DNA polymerases preferentially insert the correct nucleotide because the geometry of a proper base pair fits the enzyme’s active site far better than a mismatch. The second layer is proofreading. Most replicative polymerases have a built-in exonuclease domain that acts like a backspace key. When a wrong nucleotide is incorporated, the polymerase stalls, and the newly added strand is shuttled to the exonuclease site, which clips off the mismatch. The strand then returns to the polymerase site for another attempt. This back-and-forth between the polymerase and proofreading sites improves fidelity by more than a thousand-fold on its own.
The third layer, mismatch repair, catches errors that slip past proofreading. After the fork has moved on, dedicated repair proteins scan the freshly made DNA for mismatches, excise the incorrect section, and resynthesize it. Together, these three systems bring the error rate down to approximately one mutation per genome per cell division.
Termination in Bacteria
Bacteria with circular chromosomes have a neat termination system. In E. coli, replication starts at a single origin and two forks travel in opposite directions around the circle. To prevent one fork from overshooting and entering territory already copied by the other, the chromosome contains a series of ter sequences arranged in a region called the replication fork trap. A protein called Tus binds to these ter sites, creating one-way barriers: a fork approaching from one direction passes through, but a fork coming from the opposite direction is blocked. Termination occurs when the two converging forks meet, one at the permissive face of a Tus-ter complex and the other at the nonpermissive face that halted it.
After the forks meet, the two daughter chromosomes are often still interlinked, like chain links that have not been separated. Topoisomerase IV resolves these catenanes by cutting through one circle and passing the other through, fully separating the two finished chromosomes so they can be partitioned into the two daughter cells.
Termination in Eukaryotes
Eukaryotic cells do not use a fork-trap system. Because they fire thousands of origins, termination simply happens wherever two converging forks collide. The more interesting question in eukaryotes is how the replication machinery is dismantled afterward.
The key helicase complex in eukaryotes, called CMG (for Cdc45-MCM-GINS), must be removed from the DNA once its job is done. During elongation, the Y-shaped structure of the replication fork actively represses the ubiquitin-tagging system that would mark CMG for destruction. When two forks converge and the fork structure disappears, this repression is lifted. The ubiquitin ligase SCF-Dia2 then attaches long chains of ubiquitin to the Mcm7 subunit of CMG. Once the ubiquitin chain exceeds a threshold of about five ubiquitin molecules, an ATP-powered unfoldase called Cdc48 grabs the tagged Mcm7, unfolds it, and pulls CMG apart. This elegant system ensures the helicase is destroyed only after replication is complete, not while it is still working.
The End Replication Problem and Telomeres
Linear chromosomes, like those in human cells, have a problem that circular bacterial chromosomes avoid. Because the lagging strand polymerase needs an RNA primer upstream of wherever it is copying, there is always a small stretch at the very tip of each chromosome that cannot be fully replicated. When the final RNA primer is removed, there is nothing upstream to fill the resulting gap. This means that with each round of cell division, the ends of chromosomes get a little shorter.
Cells solve this with telomeres: long stretches of repetitive DNA sequence at chromosome ends that act as a disposable buffer. Telomeres can afford to lose a few dozen nucleotides per division because they do not contain essential genes. But this erosion has a limit. Once telomeres become critically short, the cell enters a state of growth arrest or dies, which is one reason most human cells can only divide a finite number of times.
Cells that need to keep dividing indefinitely, like stem cells and immune cells, express an enzyme called telomerase, a specialized reverse transcriptase that extends telomeres by adding repetitive DNA sequences back to the ends. Most cancer cells also reactivate telomerase, which is part of what makes them effectively immortal. Telomerase has become a target for cancer research, though developing drugs that block it without harming normal stem cells has proven difficult.
Ensuring DNA Is Copied Exactly Once Per Cycle
Copying the genome once per cell cycle is essential. Copying a region twice would create extra gene copies that could disrupt cell function, while skipping a region would lose vital information. Cells enforce this “once and only once” rule through an elegant system tied to cyclin-dependent kinases, or CDKs.
The basic principle, worked out in yeast and conserved in human cells, is that the same enzymes that trigger origin firing simultaneously prevent new origins from being licensed. At low CDK activity (during the G1 phase), origins can be loaded with the MCM helicase in preparation for replication, but firing is not yet triggered. At intermediate CDK activity, origins fire, but any attempt to reload an already-fired origin is blocked. At high CDK activity, the cell moves into mitosis. This ratchet-like mechanism means that by the time an origin has fired, the conditions that allowed it to be prepared in the first place no longer exist.
One specific example of this dual control involves the licensing factor Cdc6. Cyclin E-CDK2 phosphorylates Cdc6, which stabilizes it during a narrow window before S phase, allowing origins to be licensed. But once S phase is underway and cyclin A levels rise, Cdc6 is degraded and licensing is shut off. This creates a one-way gate: origins get loaded, they fire, and then the cell makes it chemically impossible to reload them until the next cell cycle.
When Forks Stall and How Cells Cope
Real genomes are not pristine templates. DNA lesions, tightly bound proteins, unusual secondary structures, and collisions with the transcription machinery can all cause a replication fork to stall. This is called replication stress, and cells have evolved multiple recovery strategies to deal with it.
One response is translesion synthesis, in which specialized, error-prone polymerases temporarily replace the high-fidelity replicative polymerase and copy past the damaged site. This gets the fork moving again but at the cost of a higher chance of introducing a mutation. Another response is fork reversal: the stalled fork literally backs up, reannealing the parental strands and extruding the newly synthesized strands into a four-way structure that can be processed or protected until the obstacle is resolved. Proteins that catalyze fork reversal and those that protect the reversed fork from being chewed up by nucleases are both critical. If a stalled fork collapses into a double-strand break, the cell can use homologous recombination to rebuild the fork and resume replication.
PCNA, the sliding clamp, plays a central role in coordinating which recovery pathway is used. Post-translational modifications to PCNA, including ubiquitin and SUMO tags, act as molecular switches that recruit different sets of repair and bypass proteins depending on the type of damage encountered.
Drugs That Target the Replication Machinery
Because DNA replication is essential for cell growth, it is a natural target for both antibiotics and cancer drugs. Fluoroquinolone antibiotics, a widely used class that includes ciprofloxacin and levofloxacin, work by targeting the bacterial topoisomerases DNA gyrase and topoisomerase IV. These drugs lock the topoisomerase onto the DNA after it has cut the strands but before it can reseal them, converting the enzyme from a helpful tension reliever into an agent that generates lethal double-strand breaks. Because human topoisomerases are structurally different from the bacterial versions, fluoroquinolones can kill bacteria without directly harming human cells.
Cancer chemotherapy takes a similar approach with human topoisomerase inhibitors. Drugs like camptothecin (and its derivatives topotecan and irinotecan) trap topoisomerase I on DNA, while etoposide and mitoxantrone target topoisomerase II. In both cases, the stalled topoisomerase-DNA complex collides with the replication fork, generating double-strand breaks that trigger cell death. Because cancer cells replicate more frequently than most normal cells, they are disproportionately vulnerable, though the overlap with rapidly dividing healthy tissues like bone marrow and gut lining explains many chemotherapy side effects.
Replication Outside the Nucleus
Mitochondria, the energy-producing compartments inside human cells, contain their own small circular genomes and replicate them independently of the nuclear DNA. Mitochondrial DNA replication follows a strand-displacement model: one strand is copied first, peeling away from the other, and the second strand is copied only after the first has progressed far enough to expose a separate origin. This is fundamentally different from the simultaneous leading-and-lagging-strand synthesis used in the nucleus.
Some viruses use yet another strategy called rolling-circle replication. In this mode, one strand of a circular DNA molecule is nicked, and a polymerase extends the cut end while continuously displacing the old strand, producing long, linear copies that are later cut into individual genome-length pieces. This mechanism is found in certain bacteriophages, plant viruses, and animal viruses such as porcine circovirus, where a conserved stem-loop structure at the origin is essential for proper termination of the rolling-circle process.