Robust Cell Type Decomposition, or RCTD, is a statistical method that takes the blurry, mixed signals from spatial transcriptomics experiments and separates them into the individual cell types that produced them. Spatial transcriptomics lets researchers see where genes are active across a tissue slice, but many of these technologies capture gene activity from spots that contain several cells at once, making it hard to tell which cell type is doing what. RCTD solves this by using single-cell RNA sequencing data as a reference guide, estimating the proportion of each cell type at every location. It has become one of the most widely used tools in the field, especially in cancer and neuroscience research, and its performance has held up well across multiple independent benchmarks.
The Mixed-Signal Problem in Spatial Transcriptomics
Spatial transcriptomics technologies measure gene expression while preserving the physical layout of a tissue. A thin slice of tissue gets placed on a specially prepared surface, and the RNA from each spot is captured and sequenced. The result is a map showing which genes are turned on and where. The catch is resolution. Many of these platforms, including the widely used Visium system, measure gene expression not from individual cells but from spots that cover an area containing anywhere from a handful to dozens of cells. Those cells can be quite different from one another. A spot near a tumor boundary might contain cancer cells, immune cells, and connective tissue cells all mixed together.1Nature Biotechnology. Spatially informed cell-type deconvolution for spatial transcriptomics
What the sequencer reads from that spot is a blend of all those cells’ gene activity. Without a way to pull apart those signals, you cannot say which genes belong to the tumor cells versus the immune cells, and you lose the biological story that spatial data is supposed to tell. This is the deconvolution problem: given a mixed signal, figure out what cell types contributed to it and in what proportions.
How RCTD Separates Cell Types
RCTD works by comparing the gene expression measured at each spot to known cell type profiles, then figuring out what mixture of those profiles best explains the data. The known profiles come from single-cell RNA sequencing, a separate technology that sequences individual cells and tells you which genes are active in each cell type. Think of it like identifying ingredients in a smoothie by comparing its flavor profile to the individual fruits you already know.
For each spot, RCTD estimates what fraction belongs to each cell type. If a spot’s gene expression looks like it came from roughly 60 percent epithelial cells and 40 percent immune cells, RCTD will report those proportions. The method accounts for the total number of RNA molecules captured at each spot and uses a statistical framework that handles the noisiness inherent in sequencing data.2PubMed Central. Robust decomposition of cell type mixtures in spatial transcriptomics
One feature that sets RCTD apart is how it handles platform differences. Single-cell RNA sequencing and spatial transcriptomics use different technologies, and these technologies do not capture the same genes with equal efficiency. A gene that appears highly active in single-cell data might look quieter in spatial data, or vice versa, purely because of technical differences. RCTD includes a correction step that accounts for these technology-specific biases, estimating and adjusting for gene-by-gene differences between platforms.3PubMed Central. Robust decomposition of cell type mixtures in spatial transcriptomics This built-in correction means you can pair spatial data from one technology with a reference dataset generated on a completely different sequencing platform and still get reliable results.
RCTD also handles batch effects through this same normalization approach, quantifying a random effect for each gene to absorb systematic differences that arise when data is generated at different times, in different labs, or on different machines.4Briefings in Bioinformatics. A comprehensive comparison on cell-type composition inference for spatial transcriptomics data
Choosing a Reference Dataset
Because RCTD depends on a single-cell reference to define what each cell type looks like, the quality and composition of that reference matters. A natural question is how large the reference needs to be, and whether it needs to come from the same patient or tissue type as the spatial data.
The answer is more forgiving than you might expect. In a study of breast cancer spatial transcriptomics, RCTD maintained stable accuracy with as few as 500 reference cells, showing no meaningful improvement when the reference was scaled up to 20,000 cells. Matched references, where the single-cell data came from the same type of tissue as the spatial data, offered modest accuracy gains, but the effect was small. Broad cancer atlases also produced reliable results.5PubMed. Impact of Single-Cell RNA Reference Selection for the Deconvolution of Breast Cancer Spatial Transcriptomics Datasets RCTD was also less sensitive to reference choice than some competing methods, meaning you can get trustworthy results even without a perfectly matched dataset.6bioRxiv. Impact of single-cell RNA reference selection for the deconvolution of breast cancer spatial transcriptomics datasets
Where reference selection does become critical is when cell types present in the tissue are missing from the reference altogether. A systematic evaluation found that the performance of deconvolution methods, including RCTD, drops in proportion to the number of cell types absent from the reference.7bioRxiv. Systematic evaluation of robustness of deconvolution methods for spatial transcriptomics data in case of cell type mismatch If the reference is missing a cell type that is genuinely present in the tissue, the method has no choice but to assign those cells’ expression to the nearest available category, which distorts the proportions. The practical advice is straightforward: make sure your reference covers the cell types you expect to find. You do not need a massive reference, but you need a complete one.
How RCTD Compares to Other Methods
Several deconvolution methods exist for spatial transcriptomics, and RCTD has been evaluated alongside them in multiple independent benchmarking studies. The consensus is that RCTD is consistently among the top performers, though it rarely claims the absolute top spot across every scenario.
A large benchmark published in Nature Communications evaluated many methods across different tissue types and platforms. CARD, Cell2location, Tangram, and RCTD were identified as the best-performing methods overall. RCTD performed especially well on Slide-seqV2 and Stereo-seq datasets, where it ranked among the top three.8PubMed Central. A comprehensive benchmarking with practical guidelines for cellular deconvolution of spatial transcriptomics A separate benchmark published in Bioinformatics found Cell2location in first place overall, followed by RCTD and spatialDWLS. On a mouse brain dataset, RCTD achieved very low error rates and high correlation with ground truth, closely trailing Cell2location.9Bioinformatics. Benchmarking and integration of methods for deconvoluting spatial transcriptomic data
The picture that emerges from these comparisons is that RCTD is not always the single best method, but it is reliably near the top. Cell2location tends to edge it out slightly in controlled comparisons when neither method faces missing cell types or platform mismatches. RCTD’s strengths show up more clearly in its robustness: it is less sensitive to reference size, reference matching, and platform differences. For researchers who want strong performance without needing a perfectly curated setup, that resilience matters.
Mapping the Tumor Microenvironment
Cancer research has been one of the most active areas for RCTD. Tumors are not just masses of cancer cells. They are complex ecosystems containing immune cells, blood vessel cells, fibroblasts, and other supporting tissue, all arranged in spatial patterns that affect how the tumor grows, spreads, and responds to treatment. Understanding where immune cells are located relative to cancer cells is directly relevant to predicting whether immunotherapy will work.
In prostate cancer research, RCTD was used with Slide-seqV2 data to map cell type distributions across tumor tissue. The deconvolution revealed that the spatial data showed much larger proportions of epithelial cells and fibroblasts, and far fewer immune cells, than single-cell sequencing alone had suggested. This kind of discrepancy highlights why spatial context matters: bulk or single-cell data alone can overrepresent immune cells simply because they are easier to isolate.10Nature Communications. Dissecting the immune suppressive human prostate tumor microenvironment via integrated single-cell and spatial transcriptomic analyses
RCTD has also been applied in combination with more specialized spatial techniques. One study paired it with Slide-TCR-seq, a method that simultaneously captures gene expression and T cell receptor sequences at high spatial resolution, to map immune infiltration patterns in mouse spleen and tumor tissue.11Immunity. Slide-TCR-seq couples 10-μm-resolution spatial transcriptomics with T cell receptor sequencing In metastatic renal cell carcinoma, RCTD was applied to both Visium and COSMx datasets to assign cell types to spatial spots and individual cells, forming part of an analysis aimed at predicting which patients would respond to immunotherapy.12PubMed Central. Spatial transcriptomic profiling of metastatic renal cell carcinoma identifies chemokine-driven macrophage and CD8+ T-cell interactions predictive of immunotherapy response
Across cancer types including breast, prostate, and kidney cancer, RCTD has been used alongside other top-performing tools. In a comparison across seven cancer types, RCTD and spatialDWLS both achieved high accuracy for deconvolving cell types within tumors.13Nature Communications. Estimation of cell lineages in tumors from spatial transcriptomics data The method has become something of a standard first step when researchers want to assign cell type identities in tumor spatial transcriptomics data.
Resolving Neuron Subtypes in the Brain
The brain presents a particularly demanding test for deconvolution methods. Brain tissue contains dozens of closely related neuron subtypes whose gene expression profiles overlap substantially. Telling apart two types of excitatory neurons in the hippocampus requires much finer discrimination than distinguishing, say, an epithelial cell from a macrophage in a tumor.
In benchmarking studies on mouse brain tissue, RCTD demonstrated an ability to classify excitatory neurons into specific hippocampal subtypes, including CA1, CA2, CA3, and dentate gyrus neurons. Several competing methods could not resolve CA2 neurons at all, folding them into neighboring subtypes.14Briefings in Bioinformatics. Benchmarking mapping algorithms for cell-type annotating in mouse brain by integrating single-nucleus RNA-seq and Stereo-seq data CA2 is a small hippocampal region with a relatively small number of cells, so the ability to pick it out from its neighbors is a good stress test of a method’s sensitivity. This kind of fine-grained resolution is not just an academic exercise; the hippocampus is central to memory, and understanding which neuron subtypes are affected in diseases like Alzheimer’s depends on being able to map them accurately in tissue.
Going Beyond Proportions With C-SIDE
Knowing that a spot contains 60 percent of one cell type and 40 percent of another is useful, but it only answers the “who is here” question. Researchers often want to know whether the same cell type behaves differently depending on where it is in the tissue. An immune cell near a tumor might express different genes than an immune cell far from it, even though both are classified as the same cell type.
C-SIDE, or Cell type-Specific Inference of Differential Expression, was developed by the same group behind RCTD to address exactly this. C-SIDE takes the cell type proportions that RCTD estimates and uses them to identify genes whose expression changes across tissue locations within a single cell type. The key challenge is that changes in gene expression at a spot could be caused by a change in cell type proportions rather than a real change in how a cell type behaves. If a spot near the tumor has more macrophages, overall expression of macrophage genes will go up even if each individual macrophage is behaving the same way. C-SIDE accounts for this confounding by jointly modeling cell type composition and gene expression changes.15PubMed Central. Cell type-specific inference of differential expression in spatial transcriptomics
C-SIDE is distributed within the same R package as RCTD, called spacexr, so the two methods are designed to work together in a single analysis pipeline. You run RCTD first to get cell type proportions, then feed those results into C-SIDE to ask cell type-specific questions about spatial gene expression. This combination moves spatial transcriptomics analysis from a descriptive exercise (what cells are where) to a mechanistic one (how do cells change their behavior based on their spatial context).
Practical Details for Running RCTD
RCTD is implemented in R and distributed through the spacexr package, available on GitHub. The input requirements are relatively straightforward: you need a spatial transcriptomics dataset with gene expression counts per spot and a single-cell RNA-seq reference dataset with cell type labels. The method first uses the reference to estimate average gene expression profiles for each cell type, selects genes that differ meaningfully across cell types, estimates platform effects, and then fits the mixture model to each spot.
RCTD offers multiple running modes. The default “full” mode estimates a continuous proportion for every cell type at every spot. A “doublet” mode is designed for platforms with near-single-cell resolution, like Slide-seq, where each spot likely contains at most two cell types. There is also a “multi” mode for spots with potentially many cell types. The choice depends on the platform’s resolution and the biological context.2PubMed Central. Robust decomposition of cell type mixtures in spatial transcriptomics
The method was originally developed for Slide-seq data, which captures gene expression at roughly 10-micrometer resolution, but it has since been applied across many platforms including Visium, Stereo-seq, and COSMx.4Briefings in Bioinformatics. A comprehensive comparison on cell-type composition inference for spatial transcriptomics data This cross-platform flexibility, supported by the built-in platform-effect correction described earlier, is one reason RCTD appears so frequently in published spatial transcriptomics analyses.
When RCTD Is Not the Best Fit
No tool works perfectly in every situation, and RCTD has specific limitations worth knowing about. The most significant is its dependence on an external single-cell reference. If you are studying a tissue or organism for which no good single-cell atlas exists, you cannot run RCTD. Reference-free methods like STdeconvolve, which infer cell types directly from spatial data without external input, fill that gap at the cost of less precise cell type annotations.
In the metastatic breast cancer study, researchers compared RCTD against TACCO-OT for label transfer to spatial data and ultimately selected TACCO-OT for downstream analyses because it handled both count-based and non-count-based data more flexibly.16Nature Medicine. A multi-modal single-cell and spatial expression map of metastatic breast cancer biopsies across clinicopathological features This reflects a broader pattern: RCTD is excellent for the core deconvolution task but may not be the ideal choice when the analysis workflow involves data types beyond standard count matrices.
Cell2location also tends to outperform RCTD in fully controlled benchmarks where the reference is well matched and complete.7bioRxiv. Systematic evaluation of robustness of deconvolution methods for spatial transcriptomics data in case of cell type mismatch The trade-off is that Cell2location is a Bayesian deep-learning method that requires more computational resources and longer run times. For large-scale studies or exploratory analyses where speed and simplicity matter, RCTD remains a practical choice.
The Rise of Subcellular-Resolution Platforms
Newer spatial transcriptomics platforms like MERFISH, Xenium, and CosMx are approaching or achieving true single-cell resolution, which raises a reasonable question: if each spot contains only one cell, do you still need deconvolution at all? The short answer is that even at high resolution, the data is not always clean. Some spots may still capture transcripts from neighboring cells, especially in densely packed tissues. And deconvolution tools like RCTD have found a second life in these contexts as a way to assign cell type labels to individual cells by matching their expression profiles against a reference, effectively using the same framework for cell type annotation rather than mixture decomposition.12PubMed Central. Spatial transcriptomic profiling of metastatic renal cell carcinoma identifies chemokine-driven macrophage and CD8+ T-cell interactions predictive of immunotherapy response
Even as the resolution of spatial platforms improves, the computational logic behind RCTD remains relevant. Tissues are inherently mixed environments. Cells sit next to each other, signals bleed across boundaries, and biological questions often require knowing not just what type a cell is but what fraction of a neighborhood it represents. Whether you are decomposing a 55-micrometer Visium spot into seven cell types or confirming the identity of a single cell on a Xenium slide, the reference-matching approach RCTD pioneered continues to anchor how the field interprets spatial gene expression data.