Single-cell RNA sequencing, or scRNA-seq, is a technology that measures gene activity in individual cells rather than averaging across an entire tissue sample. Where older methods blended thousands or millions of cells into a single readout, scRNA-seq can profile up to 20,000 individual cells simultaneously, revealing which genes are switched on or off in each one.1International Journal of Oral Science. From bulk, single-cell to spatial RNA sequencing That resolution has made it one of the most transformative tools in modern biology, used to discover rare cell types, trace how cells change during development, and understand how tumors evolve resistance to drugs.
Why Measuring One Cell at a Time Matters
To appreciate what scRNA-seq does, it helps to understand the limitation it was designed to overcome. Traditional bulk RNA sequencing grinds up a tissue sample and reads out the combined gene expression of every cell in it. The result is a single averaged profile, useful for comparing broad differences between, say, a healthy liver and a diseased one. But a liver contains hepatocytes, immune cells, endothelial cells lining blood vessels, and many other types. Bulk sequencing blurs all of them together, so you cannot tell whether a gene that appears active is turned on in every cell or cranked up in just a small subset.2PubMed Central. Bioinformatics perspectives on transcriptomics: A comprehensive review of bulk and single-cell RNA sequencing analyses
ScRNA-seq solves that by assigning every transcript it reads back to the individual cell it came from. This lets researchers uncover cellular heterogeneity, identify rare cell populations that might make up less than one percent of a tissue, and distinguish between cell populations that look identical under a microscope but behave very differently at the molecular level.2PubMed Central. Bioinformatics perspectives on transcriptomics: A comprehensive review of bulk and single-cell RNA sequencing analyses A tumor biopsy that looks uniform might contain dozens of molecularly distinct subpopulations when profiled cell by cell. That kind of detail is invisible to bulk methods.
How Cells Get Isolated
The first physical challenge is separating a tissue into individual cells and then keeping track of which RNA molecules belong to which cell. Two main strategies dominate the field: droplet-based methods and microwell-based methods.3PubMed Central. Microfluidic design in single-cell sequencing and application to cancer precision medicine
In the droplet-based approach, a suspension of single cells flows through a microfluidic chip alongside a stream of tiny gel beads. Oil pinches the flow into nanoliter-scale droplets, each ideally containing one cell and one bead. The bead carries short DNA sequences, called barcodes, that will tag every RNA molecule in that droplet so it can later be traced back to a single cell. Technologies like Drop-seq and the widely used 10x Genomics Chromium system work this way. They are popular because they are relatively cheap per cell, easy to operate, and can process thousands of cells in a single run.3PubMed Central. Microfluidic design in single-cell sequencing and application to cancer precision medicine
Microwell-based approaches use plates with thousands of tiny wells, each designed to trap a single cell. Barcoded beads are loaded into those wells alongside the cells. While these systems can be more sensitive and accurate than droplet methods, they tend to be more expensive to commercialize, which is one reason droplet-based platforms have become the industry default.3PubMed Central. Microfluidic design in single-cell sequencing and application to cancer precision medicine
From RNA to Readable Data
Once a cell is isolated inside a droplet or microwell, the cell membrane is broken open (lysed), releasing its messenger RNA. Mature mRNAs in animal cells have a characteristic tail made of repeated adenine bases, called a poly-A tail. The barcoded beads carry short stretches of complementary thymine (oligo-dT) that grab onto these tails, physically capturing the mRNA from each cell.4Communications Biology. A practical guide to targeted single-cell RNA sequencing technologies The beads then serve as a surface for reverse transcription, a reaction that converts each captured mRNA into a more stable DNA copy (cDNA) while simultaneously attaching the bead’s barcode to it.5Scientific Reports. An Automated Microwell Platform for Large-Scale Single Cell RNA-Seq
At this stage, each cDNA molecule carries two important tags: a cell barcode (identifying which cell it came from) and a unique molecular identifier, or UMI. The UMI is a short random sequence that distinguishes genuinely different original RNA molecules from copies produced during a later amplification step. Without UMIs, you could not tell whether a gene appeared highly active because the cell truly produced a lot of its RNA, or because the same molecule was duplicated many times during sample preparation.6PubMed Central. Elimination of PCR duplicates in RNA-seq and small RNA-seq using unique molecular identifiers
After barcoding and reverse transcription, all the cDNA from thousands of cells gets pooled into a single tube. This is possible because every molecule is already tagged. The pooled cDNA is amplified and then prepared into a sequencing library, which is fed into a high-throughput sequencer.7PubMed Central. Preparation of Single-Cell RNA-Seq Libraries for Next Generation Sequencing The sequencer reads millions of short fragments, and software sorts each one back to its cell of origin using the barcode. The end product is a large matrix: rows are genes, columns are individual cells, and each number in the matrix represents how many unique RNA molecules of a given gene were detected in that cell.
Cleaning Up the Raw Data
Raw scRNA-seq data is noisy. Not every droplet captures exactly one healthy cell. Some droplets end up empty, some capture two cells stuck together (called doublets), and some capture a dying cell that has already lost most of its RNA and is leaking mitochondrial material. Before any biological analysis, researchers filter out these low-quality events. A common approach is to remove cells whose reads are dominated by mitochondrial genes (a sign of a damaged cell) or whose total UMI count is suspiciously high (suggesting a doublet) or suspiciously low (suggesting an empty droplet or a cell that was mostly dead).8PubMed Central. Multimodal single-cell analysis of nonrandom heteroplasmy distribution in human retinal mitochondrial disease
Even after quality control, the data contains a huge number of zeros. A typical cell expresses thousands of genes, but any individual gene might not be detected in many cells simply because the sequencing did not go deep enough to catch every low-abundance transcript. This “dropout” phenomenon, sometimes called zero inflation, means that a zero in the matrix does not always mean a gene is truly silent in that cell. Analytical software has gotten quite good at handling this. For example, the popular tool Seurat maintains accurate cell-clustering results even when up to about 60% of real counts are artificially masked as zeros.9PubMed Central. Statistics or biology: the zero-inflation controversy about scRNA-seq data
Visualizing Thousands of Cells at Once
A single scRNA-seq experiment can measure expression of 20,000 genes across thousands of cells. No human can look at a spreadsheet that large and spot patterns. Researchers use dimensionality reduction algorithms that compress all that information into two-dimensional plots where each dot represents a cell and cells with similar gene expression profiles cluster together. If the algorithm works well, cells of the same type form tight groups on the plot, while distinct cell types separate into their own neighborhoods.
The two most widely used algorithms for this are t-SNE and UMAP.10PubMed Central. Visualizing Single-Cell RNA-seq Data with Semisupervised Principal Component Analysis A systematic comparison of ten such methods across dozens of datasets found that t-SNE delivered the best accuracy overall, while UMAP offered the highest stability and did a strong job preserving the natural separation between cell populations.11PubMed Central. A Comparison for Dimensionality Reduction Methods of Single-Cell RNA-seq Data UMAP also tends to run faster and produce more reproducible results, which has made it the go-to choice in many labs.12Nature Biotechnology. Dimensionality reduction for visualizing single-cell data using UMAP The colorful “island plots” you see in genomics papers, where clusters of dots represent different cell types, are almost always UMAP or t-SNE projections of scRNA-seq data.
What Researchers Do with the Clusters
Once cells are grouped into clusters, the real biology begins. The most common first step is differential expression analysis, which identifies genes that are turned on more (or less) in one cluster compared to others. These marker genes tell you what each cluster is. If a cluster lights up for well-known T-cell markers, you can label it as T cells. If another cluster expresses liver-specific enzymes, it is likely hepatocytes. Differential expression analysis is the primary downstream step for detecting cell types and feeds into almost every other kind of follow-up analysis.13PubMed Central. Differential Expression Analysis of Single-Cell RNA-Seq Data: Current Statistical Approaches and Outstanding Challenges
Another powerful use is trajectory inference, sometimes called pseudotime analysis. Cells captured from a dynamic biological process, like stem cells differentiating into mature blood cells, do not all arrive at the same stage simultaneously. By computationally ordering the cells along a trajectory based on their progressively changing gene expression, researchers can reconstruct the sequence of molecular events even though all the cells were captured at a single time point.14Nature Communications. A statistical framework for differential pseudotime analysis with multiple single-cell RNA-seq samples This approach has become especially important for studying embryonic development and how immune cells mature in response to infection.15PubMed Central. Gene trajectory inference for single-cell data by optimal transport metrics
When You Cannot Use Fresh Tissue
Standard scRNA-seq requires fresh tissue that can be dissociated into a clean single-cell suspension. That is a problem for archived samples sitting in biobank freezers, or for tissues like brain and skeletal muscle that are notoriously difficult to break apart without destroying the cells. The workaround is single-nucleus RNA sequencing (snRNA-seq), which profiles the RNA inside isolated cell nuclei instead of whole cells.16Nature Medicine. A single-cell and single-nucleus RNA-Seq toolbox for fresh and frozen human tumors
Nuclei are sturdier than whole cells and survive the freezing and thawing process much better. Researchers have developed reliable protocols to isolate nuclei from frozen human skeletal muscle, including tissue that has been stored for years and tissue with significant disease-related damage.17PubMed Central. A protocol for single nucleus RNA-seq from frozen skeletal muscle Similar methods have been optimized for long-term-frozen brain tumor tissue, dramatically expanding the potential for studying rare cancers where fresh samples are hard to come by.18Scientific Reports. A simplified preparation method for single-nucleus RNA-sequencing using long-term frozen brain tumor tissues The trade-off is that nuclear RNA does not perfectly mirror whole-cell RNA; some transcripts that are mainly found in the cytoplasm will be underrepresented. But for many research questions, snRNA-seq captures the essential biology and is the only practical option.
Applications in Cancer and Disease
Cancer research has been one of the biggest beneficiaries of scRNA-seq. Tumors are not uniform masses of identical cancer cells. They contain a patchwork of molecularly distinct cancer subpopulations, immune cells that have infiltrated the tumor, blood vessel cells, and supporting tissue. ScRNA-seq can pull these apart and characterize each one individually. In glioblastoma, the deadliest brain cancer, scRNA-seq has mapped the transcriptome of both primary and recurrent tumors at single-cell resolution, revealing how the tumor microenvironment shifts and how drug-resistance mechanisms develop after standard treatment.19PubMed Central. Single-cell RNA sequencing reveals tumor heterogeneity, microenvironment, and drug-resistance mechanisms of recurrent glioblastoma Similar work in lymphoma has uncovered heterogeneity, microenvironment interactions, and prognostic biomarkers that would be invisible to bulk sequencing.20PubMed Central. Single-cell RNA sequencing in diffuse large B-cell lymphoma: tumor heterogeneity, microenvironment, resistance, and prognostic markers
Beyond cancer, scRNA-seq is reshaping immunology, developmental biology, and neuroscience. It has been used to track how embryonic cells differentiate into the hundreds of specialized cell types in the adult body, to identify which immune cells respond to specific infections, and to catalog the extraordinary diversity of neurons in the brain. Large collaborative projects have begun building comprehensive single-cell atlases across more than 125 human fetal and adult tissues, combining scRNA-seq with other single-cell measurements into searchable reference maps.21PubMed Central. Single Cell Atlas: a single-cell multi-omics human cell encyclopedia These atlases serve as baselines against which disease states can be compared, cell type by cell type.
Validating What scRNA-seq Finds
A common misconception is that scRNA-seq results are self-contained. In practice, findings almost always need orthogonal validation, meaning confirmation by an independent method. RNA levels do not perfectly predict protein levels, and computational clusters sometimes group cells in ways that reflect technical noise rather than true biology. Researchers routinely follow up scRNA-seq with techniques like flow cytometry, mass cytometry, or in-situ RNA staining to confirm that the cell types and marker genes identified computationally really exist in the tissue.22PubMed Central. Human bone marrow assessment by single-cell RNA sequencing, mass cytometry, and flow cytometry In one bone marrow study, direct comparison between scRNA-seq and flow cytometry revealed discrepancies in how certain immune cell populations were quantified, underscoring why relying on a single technology can be misleading.22PubMed Central. Human bone marrow assessment by single-cell RNA sequencing, mass cytometry, and flow cytometry
Adding Location with Spatial Transcriptomics
One thing scRNA-seq inherently loses is spatial information. By dissociating tissue into a cell suspension, you lose track of where each cell was physically sitting. A T cell that was right next to a tumor cell behaves differently from one floating in the bloodstream, but once both are in a suspension, they look the same. Spatial transcriptomics methods address this by measuring gene expression while preserving the tissue’s architecture, either by sequencing RNA directly on a tissue slice or by staining tissue sections with fluorescent probes that light up specific transcripts.
Integrating scRNA-seq with spatial transcriptomics is now a rapidly growing research frontier. The idea is to use scRNA-seq’s deep per-cell profiling to define cell types, then map those types back onto tissue sections using spatial data, reconstructing the tissue’s cellular neighborhoods.23PubMed Central. Integrating single-cell and spatial transcriptomics to elucidate intercellular tissue dynamics New computational tools are emerging to handle this integration, including methods that can align not just RNA data but also chromatin accessibility and DNA methylation measurements onto spatial coordinates.24Nature Communications. Spatial integration of multi-omics single-cell data with SIMO The combination promises to answer questions that neither technology can address alone: which cells are talking to their neighbors, and what those conversations look like at the molecular level.
AI Models Trained on Single-Cell Data
The sheer volume of scRNA-seq data now available has attracted attention from the artificial intelligence community. Researchers have begun building what are called foundation models for single-cell biology: large AI systems, typically based on transformer architectures similar to those behind modern language models, trained on tens of millions of single-cell profiles to learn general patterns of gene regulation.25Experimental & Molecular Medicine. Single-cell foundation models: bringing artificial intelligence into cell biology
One such model, scGPT, was trained on over 33 million cells and has shown it can annotate cell types, integrate data from different experiments, and predict how cells will respond to genetic perturbations.26Nature Methods. scGPT: toward building a foundation model for single-cell multi-omics using generative AI Another, scFoundation, used 100 million parameters trained on over 50 million human single-cell profiles and achieved strong performance across tasks ranging from predicting drug responses to inferring gene regulatory networks.27Nature Methods. Large-scale foundation model on single-cell transcriptomics These models are still in early stages, but they represent a shift from analyzing one experiment at a time to learning generalizable rules across the entire corpus of available single-cell data. If they mature as hoped, a researcher could feed in a new dataset and get high-quality cell-type labels, pathway predictions, or perturbation forecasts without building a bespoke analysis pipeline from scratch.
Privacy Risks in Published Datasets
A less obvious issue with scRNA-seq is privacy. Because the technology reads RNA from individual cells, and RNA expression patterns carry genotypic signatures, publicly shared scRNA-seq datasets can potentially be traced back to the person they came from. A 2024 study demonstrated this concretely: using only the publicly available count matrices (the standard output researchers share), attackers were able to re-identify over 93% of individuals in one large dataset and over 98% in another, using mainly the most common cell types in each sample.28Cell. scRNA-seq datasets are vulnerable to linking attacks Even though single-cell measurements are noisier than bulk sequencing, the genotypic information embedded in the expression data is enough to link a sample to an individual across separate datasets.29Patterns. Privacy of single-cell gene expression data
This matters because open data sharing has been a cornerstone of the scRNA-seq field’s rapid progress. Cell atlases, cancer datasets, and developmental biology resources are typically deposited in public repositories specifically so other labs can reuse them. The discovery that these datasets are vulnerable to re-identification attacks has prompted discussion about whether new anonymization or access-control strategies will be needed as the volume of human single-cell data continues to grow. For participants in biomedical studies, the implication is that “de-identified gene expression data” may not be as anonymous as it sounds.