What Is Squidpy? A Tool for Spatial Transcriptomics

Squidpy is an open-source Python library designed to analyze spatial transcriptomics data, the kind of data that tells you not just which genes are active in a cell, but where that cell sits in a tissue. Published in Nature Methods, it provides a collection of tools for building spatial graphs, running neighborhood enrichment analyses, computing spatial statistics, and processing tissue images, all within the familiar Python data-science ecosystem.1PubMed Central. Squidpy: a scalable framework for spatial omics analysis If you work with spatial omics data or are just trying to understand what researchers mean when they reference it, here is what Squidpy does and why it matters.

Why Location Matters in Gene Expression

For years, the standard way to study gene activity across thousands of cells was single-cell RNA sequencing. That technology revolutionized biology by letting researchers identify distinct cell types within a tissue sample. The catch is that the tissue has to be broken apart first. Once you dissociate a piece of lung or brain into individual cells, you lose all information about which cells were sitting next to each other and how they were communicating.2Nature Reviews Genetics. Integrating single-cell and spatial transcriptomics to elucidate intercellular tissue dynamics That spatial context turns out to be critical. A tumor cell behaves differently depending on whether it is surrounded by immune cells or by other tumor cells. A neuron in the hippocampus expresses different genes depending on whether it sits in the CA1 or CA3 region.

A wave of newer technologies, from multiplexed in situ hybridization to spatial barcoding platforms, now lets researchers measure gene expression while keeping cells in place within intact tissue sections.2Nature Reviews Genetics. Integrating single-cell and spatial transcriptomics to elucidate intercellular tissue dynamics These platforms produce rich, complex datasets that combine gene expression matrices with spatial coordinates and, often, high-resolution tissue images. Analyzing that data requires tools built specifically for the spatial dimension, and that is the gap Squidpy fills.

How Squidpy Fits Into the Python Ecosystem

Squidpy was designed to slot into an existing, widely used set of Python tools rather than reinvent them. It sits on top of two foundational libraries: Scanpy, the go-to package for single-cell analysis, and AnnData, a data structure that stores gene expression matrices alongside cell-level metadata.1PubMed Central. Squidpy: a scalable framework for spatial omics analysis If you have already used Scanpy to cluster cells and identify marker genes, Squidpy extends that analysis into space without forcing you to switch platforms or reformat your data.

Under the hood, Squidpy draws on several established scientific computing libraries. It uses Scikit-image for image processing, Napari for interactive image visualization, and Dask for handling datasets too large to fit in memory all at once.1PubMed Central. Squidpy: a scalable framework for spatial omics analysis The practical benefit of this modular design is that a researcher can move from quality control to clustering to spatial analysis to image exploration in a single notebook, using compatible data objects throughout. A practical tips guide for the field describes Squidpy as broadly used in the Python spatial transcriptomics ecosystem for exactly this reason.3PLOS Computational Biology. Ten quick tips for spatial transcriptomics analysis

Building Spatial Graphs

The first thing Squidpy typically does with a spatial dataset is construct a spatial graph. Think of it as drawing invisible lines between cells (or spots, depending on the technology) that are near each other in the tissue. This graph is the backbone of most downstream spatial analyses, because once you know which cells are neighbors, you can start asking whether certain cell types tend to cluster together, avoid each other, or interact in specific patterns.

Squidpy offers multiple ways to define what “neighbor” means. You can set a fixed distance threshold, so any two cells within a certain number of micrometers are connected. Or you can use methods that connect each cell to its nearest handful of neighbors regardless of absolute distance. The choice depends on the technology and the biology. In a dense imaging-based assay with subcellular resolution, a tight distance cutoff makes sense. In a spot-based technology where each measurement captures a small patch of tissue rather than a single cell, a grid-based neighborhood is more appropriate.

This spatial graph construction step is also the entry point for more complex quantitative analyses. Once the neighbor relationships are defined, Squidpy can calculate things like neighborhood enrichment scores, which measure whether pairs of cell types appear next to each other more or less often than you would expect by chance.3PLOS Computational Biology. Ten quick tips for spatial transcriptomics analysis The result is a matrix showing which cell types are spatially associated and which tend to occupy different tissue regions, a basic but powerful readout of tissue organization.

Measuring Spatial Patterns

Beyond simple neighborhood counts, Squidpy includes several statistical tools for characterizing how cells or features are distributed across a tissue. These fall into a few broad categories.

Co-occurrence scores quantify how often two categories of cell, or two subcellular features, appear together at different distance scales. In the original Squidpy paper, the authors demonstrated this by examining subcellular compartment annotations and showing that nucleus-associated measurements co-occurred with nuclear envelope features at short distances, exactly as you would expect anatomically. Applied to brain tissue data from a SlideseqV2 experiment, co-occurrence scores provided a quantitative measure of an observation researchers could see qualitatively: endothelial tip cells and ependymal cells showed strong spatial co-occurrence.1PubMed Central. Squidpy: a scalable framework for spatial omics analysis

Ripley’s L statistic is another spatial tool available through Squidpy. It evaluates whether a particular cell type is randomly scattered across the tissue, clustered into groups, or dispersed at regular intervals. Researchers have used the Ripley function in Squidpy to compare how cell types distribute themselves across different spatial transcriptomics chip designs. In one study of mouse brain tissue, oligodendrocytes, a widely distributed cell type, showed increasing aggregation at longer distances on one chip design but were more constrained on another, likely because the second chip covered a smaller area.4Nature Genetics. Custom microfluidic chip design enables cost-effective three-dimensional spatiotemporal transcriptomics with a wide field of view That kind of comparison matters when researchers are evaluating whether a new technology faithfully reproduces the tissue’s real spatial architecture.

Moran’s I, a classical measure of spatial autocorrelation, is also accessible through Squidpy. It tells you whether a gene’s expression is spatially patterned, meaning that nearby cells tend to have similar expression levels, rather than randomly sprinkled. In a study of mouse embryo development, researchers used Moran’s I to confirm that the signaling molecule Shh was highly spatially variable across the embryo at a specific developmental stage.5Nature Methods. Inferring pattern-driving intercellular flows from single-cell and spatial transcriptomics Identifying spatially variable genes is often a first step toward understanding which molecular signals are driving tissue patterning.

Ligand-Receptor Interaction Inference

One of the most compelling reasons to keep spatial information is the ability to study cell-cell communication in situ. Cells talk to each other by sending out signaling molecules (ligands) that bind to receptors on neighboring cells. When you know which cells express which ligands and receptors, and you also know which cells are sitting next to each other, you can infer which communication channels are likely active in a given tissue region.

Squidpy implements ligand-receptor interaction inference as part of its spatial analysis toolkit.3PLOS Computational Biology. Ten quick tips for spatial transcriptomics analysis The basic logic is straightforward: if cell type A expresses a ligand, cell type B expresses the corresponding receptor, and the two cell types are spatial neighbors, that ligand-receptor pair is a candidate for active signaling. By scoring these interactions across the tissue, researchers can build maps of intercellular communication that go far beyond what single-cell sequencing alone could reveal.

This capability has proven valuable in cancer research, where the communication between tumor cells and surrounding immune cells determines whether the immune system mounts an effective attack or gets suppressed. In a study of KRAS-mutant colorectal cancer, researchers used Squidpy’s spatial neighbor functions to calculate heterotypic proximity scores, essentially measuring how close different cell types sit to each other compared to what random mixing would produce.6PubMed Central. Single-cell and spatial transcriptome profiling identifies the immunosuppressive spatial niche in KRAS-mutant colorectal cancer Those scores helped identify immunosuppressive niches within the tumor microenvironment, zones where tumor cells had surrounded themselves with cell types that dampen immune activity.

Image Analysis Capabilities

Many spatial transcriptomics technologies generate high-resolution tissue images alongside gene expression data. A histology stain might accompany a Visium slide, or a fluorescence image might come from a MERFISH experiment. Squidpy’s image analysis module lets researchers extract quantitative features from these images and connect them to the gene expression layer.

In practice, this means you can compute things like texture, color intensity, or morphological features for the tissue region surrounding each spatial measurement point, then use those image-derived features in clustering, classification, or correlation analyses alongside the molecular data. Because Squidpy integrates with Napari, researchers can also interactively explore tissue images overlaid with gene expression patterns or cell-type annotations. For large images, Dask handles lazy loading so that only the portion of the image currently being viewed or analyzed needs to be in memory.

The image side of Squidpy is less commonly discussed than its spatial statistics, but it fills a real gap. Without it, researchers would need to switch to a separate image-processing pipeline, manually match coordinates between image features and gene expression spots, and lose the convenience of having everything in one data object.

Compatibility with Multiple Platforms

Spatial transcriptomics is not one technology but a growing family of platforms, each with its own resolution, throughput, and data format. Squidpy was designed to work with data from a range of these. The published demonstrations include analyses on data from Visium, SlideseqV2, MERFISH, and seqFISH, among others.1PubMed Central. Squidpy: a scalable framework for spatial omics analysis

In the years since Squidpy’s release, newer platforms have continued to appear, and the research community has adopted Squidpy for many of them. A study of mouse lung tissue used Squidpy alongside Seurat and Scanpy to analyze data generated by 10x Genomics’ Xenium platform, which provides single-cell resolution spatial data.7PubMed Central. Enhanced Spatial Transcriptomics Analysis of Mouse Lung Tissues Reveals Cell-Specific Gene Expression Changes Associated with Pulmonary Hypertension Another group applied Squidpy’s Ripley’s L function to data from Stereo-seq, a platform that uses patterned arrays to achieve large field-of-view spatial capture.4Nature Genetics. Custom microfluidic chip design enables cost-effective three-dimensional spatiotemporal transcriptomics with a wide field of view This cross-platform flexibility is a major reason Squidpy has become a common reference tool. Researchers do not need to learn a new analysis framework every time a new spatial technology enters their lab.

The trade-off, as with any general-purpose tool, is that Squidpy may not implement the most cutting-edge, technology-specific analysis methods. Platform developers sometimes release their own software with optimized algorithms for their data. Squidpy’s role is more foundational: it provides the spatial graph construction, standard statistics, and visualization hooks that nearly every spatial analysis project needs regardless of the upstream platform.

Applications in Disease Research

The biological questions spatial transcriptomics can address are broad, and Squidpy has appeared across a range of disease-focused studies. Cancer research has been an especially active area, because tumors are heterogeneous tissues where spatial organization directly affects disease progression and treatment response. The colorectal cancer study mentioned earlier used Squidpy’s spatial neighbor analysis to dissect how immune cells were spatially organized around KRAS-mutant tumor cells, revealing immunosuppressive neighborhoods that could inform immunotherapy strategies.6PubMed Central. Single-cell and spatial transcriptome profiling identifies the immunosuppressive spatial niche in KRAS-mutant colorectal cancer

Lung disease is another domain where spatial context proves essential. In a study of pulmonary hypertension in mice, researchers used Squidpy as part of their analysis pipeline for Xenium spatial data to identify cell-type-specific gene expression changes, revealing how disease processes played out differently in different regions of the lung tissue.7PubMed Central. Enhanced Spatial Transcriptomics Analysis of Mouse Lung Tissues Reveals Cell-Specific Gene Expression Changes Associated with Pulmonary Hypertension

Developmental biology is a natural fit too, since embryonic development is fundamentally about cells organizing themselves in space. The FlowSig study that used Squidpy’s Moran’s I statistic examined Stereo-seq data from mouse embryos to trace how signaling molecules like Shh, Bmp4, and Wnt5a drive pattern formation during early development.5Nature Methods. Inferring pattern-driving intercellular flows from single-cell and spatial transcriptomics Identifying the upstream drivers and downstream targets of spatially variable signals is the kind of systems-level question that neither single-cell sequencing nor bulk tissue analysis could answer alone.

Where Squidpy Fits Among Other Tools

Squidpy is not the only software for spatial transcriptomics analysis. In the R ecosystem, packages like Seurat (which added spatial capabilities in later versions) and Giotto offer overlapping functionality. Specialized tools exist for particular tasks: cell segmentation, deconvolution of spot-based data, or deep-learning-based spatial domain identification. The field is moving fast, and new methods papers appear regularly.

What distinguishes Squidpy is its position as a general-purpose spatial extension of the Scanpy/AnnData stack, which is already the dominant framework for single-cell analysis in Python. For researchers whose workflows are Python-based, Squidpy offers the path of least resistance into spatial analysis. A practical guide to the field describes it in exactly those terms, noting that Scanpy together with Squidpy is broadly used and that Squidpy implements spatial statistics including spatial graph construction, neighborhood-enrichment analyses, and ligand-receptor interaction inference.3PLOS Computational Biology. Ten quick tips for spatial transcriptomics analysis

One thing worth noting is that Squidpy is a framework for working with spatial data, not a turnkey pipeline that runs itself. It gives you the building blocks, and you need to make decisions at every step: how to define neighborhoods, which statistics to compute, how to threshold significance, how to interpret the results in context. That flexibility is its strength for experienced analysts, but it means there is a learning curve for newcomers who may be more comfortable with point-and-click interfaces. The growing library of tutorials and the integration with Scanpy’s documentation ecosystem help, but spatial analysis still requires genuine statistical thinking about what “neighbor” means in your tissue and what spatial patterns are biologically meaningful versus artifacts of tissue processing.

Scalability and Large Datasets

As spatial transcriptomics technologies improve, datasets are getting dramatically larger. Newer platforms can capture tens of thousands of genes across millions of cells in a single tissue section, and the accompanying images can be gigabytes in size. Squidpy was built with scalability in mind, and its reliance on Dask for out-of-core computation is part of that design. Rather than loading an entire dataset into memory at once, Dask breaks it into manageable chunks that are processed sequentially or in parallel.

In practice, scalability also depends on the specific analysis. Building a spatial graph for a million cells is computationally heavier than doing so for ten thousand. Neighborhood enrichment analyses that involve permutation testing, where the observed neighbor patterns are compared against many rounds of random shuffling, can become time-consuming at scale. Researchers working with very large datasets sometimes pre-filter to regions of interest or subsample before running computationally expensive steps, then validate findings on the full dataset.

The Squidpy team has continued to update the library since its initial release. Newer versions have added support for additional data formats and refined existing functions. Because it is open source and hosted on GitHub, the community can contribute bug fixes, new features, or integrations with other tools. That open development model matters in a field where new technologies and data formats appear regularly, and no single development team can anticipate every use case.

Practical Considerations for Getting Started

If you are considering using Squidpy, a few practical points are worth knowing. Installation is through pip or conda, the standard Python package managers, and it pulls in its dependencies automatically. The documentation includes tutorials built around publicly available spatial transcriptomics datasets, so you can work through complete analyses before applying the tools to your own data.

The most common entry point is loading a spatial dataset into an AnnData object (often directly from the output of a spatial transcriptomics platform’s processing software), then using Squidpy’s graph-building functions to define spatial neighborhoods. From there, you branch into whichever analyses your biological question demands: neighborhood enrichment if you want to know which cell types co-localize, spatial autocorrelation if you want to find spatially variable genes, co-occurrence or Ripley’s statistics if you want to characterize distribution patterns at different length scales, or image feature extraction if you want to integrate histological information.

One thing to keep in mind is that spatial transcriptomics analysis is still a young field, and best practices are evolving. Choices that seem standard today, like which permutation test to use for neighborhood enrichment or how to handle multiple testing correction across thousands of genes, may be refined as the community gains more experience. Squidpy gives you the tools, but staying current with the methods literature matters as much as knowing the software.

Leave a Reply

Your email address will not be published. Required fields are marked *