Whole genome bisulfite sequencing, usually abbreviated WGBS, maps chemical tags on DNA called methyl groups at every single position across an entire genome. These methyl tags sit on cytosine bases and influence which genes get turned on or off in a given cell, without changing the underlying genetic code. WGBS remains the most comprehensive way to survey this landscape, but its chemistry comes with trade-offs that have shaped both the technology itself and the biological questions researchers can answer with it.
How Bisulfite Conversion Works
The method hinges on a chemical reaction discovered independently by two research groups in 1970: sodium bisulfite selectively strips an amino group from unmethylated cytosine, converting it to uracil, while leaving methylated cytosine largely untouched.1PubMed Central. Discovery of bisulfite-mediated cytosine conversion to uracil, the key reaction for DNA methylation analysis–a personal account When the treated DNA is then amplified and sequenced, every uracil reads as thymine. A cytosine that still shows up as cytosine in the sequence was methylated; one that became thymine was not. By comparing the sequencing readout to the known reference genome, researchers can determine the methylation status of each cytosine position.
The selectivity of this reaction is what makes the whole technique possible. Bisulfite acts as a catalyst for deamination, lowering the energy barrier for the chemical step that removes the amino group from cytosine. Direct hydrolysis of cytosine without bisulfite is too slow to be practical, but bisulfite makes the reaction proceed readily under mild conditions.2PubMed. The conversion of protonated cytosine-SO3(-) to uracil-SO3(-): insights into the novel induced hydrolytic deamination through bisulfite catalysis Methylated cytosine resists this conversion because the methyl group physically blocks the bisulfite from attacking the same position on the ring. The end result is a binary readout for each cytosine in the genome.
The Price of Conversion
Bisulfite treatment is harsh. DNA is a fragile molecule, and soaking it in bisulfite at elevated temperatures breaks it into small fragments. Systematic testing of the reaction conditions found that getting maximum conversion of unmethylated cytosines required incubation at 55°C for four to eighteen hours or at 95°C for one hour, and under those conditions, roughly 84 to 96 percent of the input DNA was degraded.3PubMed Central. Bisulfite genomic sequencing: systematic investigation of critical experimental parameters That means researchers have to start with enough DNA to tolerate losing most of it, which has traditionally been a major barrier when working with rare cell types or tiny tissue samples.
Beyond sheer destruction of material, the fragmentation introduces bias. Short fragments from GC-rich regions of the genome behave differently during amplification and sequencing than fragments from other regions, so certain parts of the genome end up over-represented while others are under-represented. The downstream data can look patchy as a result, forcing researchers to sequence more deeply to compensate and driving up costs.
A Computational Headache
Even after the chemistry is done and the sequencing data rolls in, WGBS creates a distinctive computing problem. Because bisulfite conversion turns most unmethylated cytosines into thymines, the resulting sequences have far fewer distinct letters than normal DNA. A stretch that might have been ACGTCGA becomes ATGTUGA and then reads as ATGTTGA. With so many cytosines converted to thymines, sequences lose their uniqueness, making it harder for alignment software to figure out where each fragment belongs in the reference genome.4PubMed Central. Analysis and performance assessment of the whole genome bisulfite sequencing data workflow: currently available tools and a practical guide to advance DNA methylation studies – Section: Alignment
There is also an asymmetry problem. A thymine in the sequencing read could map to a cytosine in the reference genome (because it was unmethylated and got converted), but a cytosine in the reference can never map to a thymine in the read. Standard alignment tools treat mismatches symmetrically, so specialized bisulfite-aware aligners had to be built.5Briefings in Bioinformatics. Methodological aspects of whole-genome bisulfite sequencing analysis – Section: MAPPING BISULFITE-CONVERTED READS TO A GENOME REFERENCE Several dedicated pipelines have been developed to handle the full chain of steps from raw data to methylation calls, managing trimming, alignment, and downstream analysis in an automated way.6PubMed Central. MethylStar: A fast and robust pre-processing pipeline for bulk or single-cell whole-genome bisulfite sequencing data Some of these can generate publication-ready visualizations of methylation profiles, differential methylation between samples, and annotation of results against known genomic features.7PubMed Central. msPIPE: a pipeline for the analysis and visualization of whole-genome bisulfite sequencing data
The Blind Spot for Hydroxymethylcytosine
Standard bisulfite sequencing has a well-known limitation that took years to appreciate fully. In addition to ordinary methylcytosine (5mC), mammalian DNA contains another modified base called 5-hydroxymethylcytosine (5hmC). Both of these modified forms resist bisulfite conversion and read as cytosine in the final sequence, which means bisulfite sequencing cannot tell them apart.8PubMed Central. The behaviour of 5-hydroxymethylcytosine in bisulfite sequencing This matters because the two modifications play different biological roles. Methylcytosine typically silences genes, while hydroxymethylcytosine is often an intermediate step in active demethylation and tends to be enriched in active regulatory regions, particularly in the brain.
To solve this, researchers developed oxidative bisulfite sequencing. The approach chemically oxidizes 5hmC into 5-formylcytosine before the bisulfite step, and formylcytosine does get converted to uracil by bisulfite. The oxidized sample therefore gives a readout of 5mC alone, while a standard bisulfite sample gives 5mC plus 5hmC combined. Subtracting one from the other reveals the 5hmC signal at each position.9PubMed. Quantitative sequencing of 5-methylcytosine and 5-hydroxymethylcytosine at single-base resolution This was one of the first methods to quantify 5hmC genome-wide at single-base resolution, and it requires running two parallel sequencing experiments for every sample, effectively doubling the sequencing cost.10PubMed. Oxidative Bisulfite Sequencing: An Experimental and Computational Protocol
Enzymatic Alternatives to Bisulfite
The damage that bisulfite inflicts on DNA has motivated the development of gentler approaches. Enzymatic methyl-seq (EM-seq) replaces the chemical conversion with a two-step enzymatic process. First, an enzyme protects methylated and hydroxymethylated cytosines by adding a chemical group that shields them. Then a second enzyme deaminates the unprotected (unmethylated) cytosines to uracil, achieving the same C-to-T conversion that bisulfite does but without the heat, acid, and fragmentation.11PubMed Central. Enzymatic methyl sequencing detects DNA methylation at single-base resolution from picograms of DNA
In practice, EM-seq libraries end up with longer DNA fragments, fewer duplicate reads, a higher fraction of reads that map cleanly to the reference, and less GC bias compared to bisulfite-converted libraries.12PubMed Central. Enzymatic Methyl-seq: Next Generation Methylomes The reduced damage also means EM-seq can work with much smaller amounts of starting material, down to picograms of DNA rather than the nanograms typically needed for bisulfite sequencing. For researchers working with precious clinical samples or small cell populations, this is a significant practical advantage. Despite these benefits, bisulfite sequencing retains its status as the reference standard, partly because decades of data have been generated with it and partly because most published computational tools and benchmarks are built around bisulfite-converted data.
Long-Read Sequencing Skips Conversion Entirely
A more radical departure from bisulfite chemistry has come from long-read sequencing platforms that can detect base modifications directly, without any chemical or enzymatic conversion at all. Both nanopore sequencing and single-molecule real-time sequencing can sense methylated cytosines as the DNA passes through the instrument, because methylation slightly alters the electrical signal (in nanopore) or the polymerase kinetics (in single-molecule real-time sequencing) compared to unmodified bases.13PubMed Central. A comparison of methods for detecting DNA methylation from long-read sequencing of human genomes – Section: BACKGROUND
The appeal is obvious: no conversion step means no DNA damage, no reduced sequence complexity, and reads that span tens of thousands of bases instead of a few hundred. Long reads also preserve the physical linkage between distant methylation sites on the same molecule, which short-read bisulfite sequencing destroys when it fragments the DNA. The trade-off has been accuracy; calling methylation from the subtle signal changes in long reads is noisier than the binary readout of bisulfite conversion, though computational methods for this are improving rapidly.
Scaling Down to Single Cells
One of the most active frontiers in methylation profiling is doing it cell by cell. Bulk WGBS averages the methylation signal across millions of cells, masking the differences between cell types in a tissue. Single-cell WGBS (scWGBS) overcomes this by profiling individual cells, revealing epigenetic heterogeneity that would otherwise be invisible.
The challenge is that a single cell contains just a few picograms of DNA, and bisulfite treatment destroys most of it. Early protocols could cover only a modest fraction of the genome per cell. Newer methods have pushed coverage higher; one approach, scSPLAT, employs a pooling strategy to prepare libraries at higher throughput and achieved coverage of up to about 40 percent of all CpG sites in the human genome from individual cells.14Scientific Reports. scSPLAT, a scalable plate-based protocol for single cell WGBS library preparation Another improved protocol, scDEEP-mC, generates high-coverage libraries that allow cell type identification, genome-wide profiling of hemi-methylation (where only one DNA strand is methylated), and allele-specific analysis.15Nature Communications. High-coverage allele-resolved single-cell DNA methylation profiling reveals cell lineage, X-inactivation state, and replication dynamics
Even with higher coverage per cell, the data from any single cell is still sparse and stochastic, with many CpG sites missing entirely. Machine-learning approaches are being developed to fill in the gaps and extract biological signal from these patchy datasets.16bioRxiv. scWGBS-GPT: A Foundation Model for Capturing Long-Range CpG Dependencies in Single-Cell Whole-Genome Bisulfite Sequencing to Enhance Epigenetic Analysis
Reduced Representation as a Cost Compromise
Sequencing an entire genome after bisulfite conversion is expensive, and not every study needs single-base resolution across billions of bases. Reduced representation bisulfite sequencing (RRBS) addresses this by using a restriction enzyme to cut the genome at predictable sites, enriching for short fragments that contain a high density of CpG sites. Only those fragments are bisulfite-treated and sequenced, covering the most informationally dense regions at a fraction of the cost. Multiplexed versions of this protocol can process 96 or more samples per week, making it practical for large cohorts of clinical or population samples.17PubMed Central. Gel-free multiplexed reduced representation bisulfite sequencing for large-scale DNA methylation profiling The trade-off is that RRBS misses methylation changes in regions far from its enzyme cut sites, so researchers have to decide whether breadth or depth matters more for their question.
What WGBS Has Revealed About Embryonic Development
Some of the most striking biological insights from WGBS have come from mapping methylation during the earliest stages of life. In human embryos, WGBS data show a dramatic genome-wide drop in methylation during early development, declining from about 0.61 to 0.30 as embryogenesis progresses.18Cell Discovery. DNA methylation reprogramming of functional elements during mammalian embryonic development This massive erasure and subsequent re-establishment of methylation patterns is thought to reset the epigenetic slate, allowing the specialized methylation patterns of different cell types to be laid down fresh.
WGBS studies in cattle have added a comparative dimension. In bovine embryos, the major wave of genome-wide demethylation is complete by the eight-cell stage, after which new methylation patterns begin to be deposited.19PubMed Central. Methylome Dynamics of Bovine Gametes and in vivo Early Embryos The broad strokes of this reprogramming are conserved across mammals, but the timing and fine details differ between species, something that only genome-wide approaches like WGBS can capture.
Cancer Methylation Landscapes
WGBS has reshaped how researchers think about the methylation changes in cancer. Tumors do not simply gain or lose methylation uniformly. Instead, they typically show a global loss of methylation across large domains of the genome while simultaneously gaining methylation at specific small regions, particularly at CpG islands near gene promoters.20PubMed Central. Whole-genome bisulfite sequencing of cell-free DNA identifies signature associated with metastatic breast cancer
WGBS of breast cancer genomes has shown that large “partially methylated domains” become hypervariable, and the normal distinction between heavily methylated gene bodies and unmethylated promoter islands breaks down inside these domains. DNA methylation inside these regions appears to be deposited without regard to the function of the underlying genomic elements, which can drive CpG island hypermethylation, a hallmark of many cancers.21Nature Communications. Partially methylated domains are hypervariable in breast cancer and fuel widespread CpG island hypermethylation – Section: CpG island methylation in breast cancer PMDs
These methylation signatures are detectable not just in tumor tissue but in fragments of tumor DNA floating in the bloodstream. Because abnormal methylation arises early during cancer development, methylation-based analysis of cell-free DNA in blood is being developed as a noninvasive tool for early cancer detection, monitoring for residual disease after treatment, and predicting treatment outcomes.22PubMed. Liquid Biopsy of Methylation Biomarkers in Cell-Free DNA23PubMed Central. DNA Methylation-Based Testing in Liquid Biopsies as Detection and Prognostic Biomarkers for the Four Major Cancer Types
Non-CpG Methylation in the Brain
Most discussions of DNA methylation focus on CpG sites, where a cytosine is followed by a guanine. For decades, methylation at other sequence contexts (CA, CT, CC, collectively called non-CpG or CpH methylation) was thought to be negligible in mammals. WGBS proved that wrong, at least in one organ. In adult mouse neurons, about a quarter of all methylation occurs at non-CpG sites.24PubMed Central. Distribution, recognition and regulation of non-CpG methylation in the adult mammalian brain This neuronal non-CpG methylation is conserved in human brains and tends to accumulate in regions with low CpG density. It correlates negatively with gene expression, meaning higher non-CpG methylation around a gene is associated with that gene being turned down.
Comparative WGBS across species revealed that this brain non-CpG methylation system is present across vertebrates but absent in invertebrates.25PubMed Central. The emergence of the brain non-CpG methylation system in vertebrates Developmental WGBS studies of the human cortex have further shown that CpG and non-CpG methylation follow distinct trajectories during brain maturation, with neighboring sites in the same regions showing correlated changes over development.26PubMed Central. Divergent neuronal DNA methylation patterns across human cortical development reveal critical periods and a unique role of CpH methylation None of these discoveries would have been possible with targeted methods that only look at CpG sites; only whole-genome approaches capture the full picture.
Genomic Imprinting and Parent-of-Origin Effects
Certain genes are expressed from only one of your two copies, depending on whether the copy came from your mother or your father. This phenomenon, genomic imprinting, is controlled by regions called imprint control regions where methylation is locked in on one parental copy and absent on the other. WGBS across DNA from different tissue types and from gametes has been used to map these regions genome-wide, identifying 1,488 candidate imprint control regions in the human genome, including 332 where methylation in sperm and eggs approached either zero or full methylation, consistent with strict parent-of-origin marking.27PubMed Central. Genomic map of candidate human imprint control regions: the imprintome The number is far larger than the roughly 25 imprint control regions that had been previously characterized, suggesting there may be many more imprinted loci than textbooks traditionally list.
WGBS in other species has uncovered imprinted loci that are lineage-specific. In pigs, for instance, a maternally methylated region was discovered at the ZNF791 locus, along with an unannotated antisense transcript that had not been described before.28PLOS Genetics. Evolutionary lineage-specific genomic imprinting at the ZNF791 locus Findings like this underscore how much of the epigenetic landscape is still being mapped for the first time, even in well-studied mammalian genomes.
Aging, Epigenetic Clocks, and Cell Type Identity
One of the more unexpected applications of methylation profiling has been building “epigenetic clocks” that predict biological age from methylation patterns. These clocks work because certain CpG sites gain or lose methylation at a remarkably consistent rate as organisms age. WGBS studies have extended this concept beyond mammals; in haddock, for example, age-related methylation changes were found throughout the genome and were enriched in genes related to immune function, metabolism, and development.29bioRxiv. Building an ‘epigenetic clock’: Utilizing whole genome DNA methylation patterns to predict age in haddock, Melanogrammus aeglefinus
A deeper question, though, is what aging actually does to the methylation landscape at cell-type resolution. A study of over 20 million CpG sites in human neurons and non-neuronal brain cells found that aging is the single strongest predictor of methylation variation, outweighing sex and psychiatric diagnosis. The CpG sites that differ most between neurons and non-neuronal cells are particularly vulnerable to age-related drift, leading to a gradual blurring of epigenetic cell type identities over time. Interestingly, the sites used in most published epigenetic clocks tend to be ones that are not highly differentiated between cell types, meaning the clock signal and the cell-identity erosion signal are capturing distinct facets of aging.30PubMed Central. Human brain aging is associated with dysregulation of cell type epigenetic identity
Environmental Exposures Written Into the Methylome
WGBS has also been applied to study how the environment leaves marks on DNA methylation. A study comparing 120 people living in a heavily polluted region of China with those in a less polluted area (where concentrations of five air pollutants were 1.6 to 6.6 times higher) found 371 differentially methylated regions between the two groups. These regions clustered in gene regulatory elements like promoters and enhancers, and the affected genes were enriched for functions related to mitochondrial assembly, immune signaling, pulmonary disorders, and cancer.31Elsevier / ScienceDirect (Environ Pollut). Genome-wide DNA methylation analysis reveals significant impact of long-term ambient air pollution exposure on biological functions related to mitochondria and immune response
Studies like this illustrate both the power and the limits of WGBS for environmental health research. The technique can detect subtle methylation differences across the entire genome in an unbiased way, without requiring researchers to guess in advance which genes to look at. But linking those differences to specific health outcomes requires careful study design and replication. A cross-sectional comparison between people in two regions cannot easily separate the effects of air pollution from the many other factors that differ between any two populations. WGBS provides the map; establishing causation requires additional layers of evidence.
Genetic Variation Shaping the Methylome
An underappreciated complication in interpreting WGBS data is that your DNA sequence itself influences your methylation patterns. Certain genetic variants, called methylation quantitative trait loci (mQTLs), directly affect the methylation level at nearby CpG sites. A large-scale study of human fetal brain tissue tested hundreds of thousands of genetic variants against hundreds of thousands of methylation sites and identified over 16,000 mQTLs, many of which overlapped with genomic regions already linked to schizophrenia risk.32PubMed Central. Methylation quantitative trait loci in the developing brain and their enrichment in schizophrenia-associated genomic regions
This finding complicates the interpretation of any WGBS study that compares methylation between groups of people. If group A differs from group B at certain methylation sites, some of that difference may reflect genuine environmental or disease-related epigenetic change, but some may simply reflect different allele frequencies between the groups at mQTLs. Separating genetic from epigenetic effects is one of the trickiest analytical challenges in population-level WGBS research, and it is why many studies now integrate genotyping alongside methylation profiling.