SQANTI3: Quality Control of Long-Read Transcriptomes

SQANTI3 is a bioinformatics tool built specifically to evaluate and curate the transcript models that come out of long-read RNA sequencing experiments. It works by comparing each transcript to a reference genome and annotation, classifying it into structural categories, and then flagging potential artifacts based on features like splice junctions, transcription start sites, and polyadenylation signals.1PubMed Central. SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms For anyone working with PacBio or Oxford Nanopore data to study gene isoforms, SQANTI3 has become a central piece of the quality control workflow, and its classification system is now used as a shared vocabulary across the field.

Why Long-Read Transcriptomes Need Dedicated Quality Control

Short-read sequencing, the workhorse of genomics for over a decade, chops RNA into small fragments and then computationally stitches them back together to infer full-length transcripts. That reconstruction step introduces uncertainty, especially for genes that produce many different isoforms through alternative splicing. Long-read platforms from PacBio and Oxford Nanopore changed the game by reading entire RNA molecules in a single pass, eliminating the assembly problem. But they introduced new ones.

Long reads have higher per-base error rates than short reads, and those errors can create phantom splice sites or shift the apparent boundaries of a transcript. Reverse transcription artifacts, incomplete captures of the 5′ end, and internal priming at adenine-rich stretches can all generate transcript models that look real but are not. Without a systematic way to separate genuine isoforms from noise, a long-read experiment might report tens of thousands of “novel” transcripts, many of which are technical artifacts. One benchmarking study found that some transcript-calling tools produced annotations where more than a third of transcript models could not be confirmed by orthogonal evidence.2Nature Biotechnology. Accurate isoform discovery with IsoQuant using long reads That kind of false-discovery rate makes quality control not optional but essential.

How SQANTI3 Classifies Transcript Models

At its core, SQANTI3 takes a set of transcript models (usually in GTF or GFF format) and compares each one against a reference annotation to assign it a structural category. These categories form a naming system that has become standard shorthand in the field. A study of human blood cell transcriptomes illustrates the typical breakdown: each transcript gets labeled based on how its splice junctions and exon structure match up to known reference transcripts.3G3 Genes|Genomes|Genetics. Reference long-read isoform-aware transcriptomes of 4 human peripheral blood lymphocyte subsets

The main categories include:

  • Full Splice Match (FSM): Every splice junction in the transcript matches a known reference transcript exactly. This is the highest-confidence category for known isoforms.
  • Incomplete Splice Match (ISM): The transcript matches a known reference but is missing some junctions, typically because it was truncated at the 5′ or 3′ end.
  • Novel in Catalog (NIC): The transcript uses known donor and acceptor splice sites from the reference but combines them in a new way not previously annotated.
  • Novel Not in Catalog (NNC): The transcript contains at least one splice junction that uses a donor or acceptor site not found anywhere in the reference annotation.
  • Other categories: Transcripts landing between annotated genes (intergenic), on the opposite strand from a known gene (antisense), spanning two genes (fusion), or falling within a gene’s intron or genomic region without matching known splicing patterns.

In practice, most transcripts from a well-prepared library fall into the FSM category, meaning they match known biology. But substantial numbers of NIC and NNC transcripts also appear, and the question is always which of those represent genuine biological novelty and which are errors. That is where the rest of SQANTI3’s toolkit comes in.

Quality Descriptors and Artifact Detection

Beyond structural classification, SQANTI3 calculates a battery of quality descriptors for each transcript model, its junctions, and its transcript ends. These descriptors provide the evidence a researcher needs to decide whether a given transcript should be kept or discarded. The tool examines whether splice junctions use canonical dinucleotide motifs (the GT-AG pairs that mark nearly all real splice sites in mammals), whether the 5′ end overlaps with known transcription start sites, and whether the 3′ end sits near a recognized polyadenylation signal or motif.1PubMed Central. SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms

One particularly useful check involves short-read RNA-seq data. If you have matching Illumina data for the same sample, SQANTI3 can verify whether each splice junction in a long-read transcript is also supported by short reads. A junction seen in both platforms is far more likely to be real than one found only in long reads, where a single base error could create a spurious splice site. CAGE (cap analysis of gene expression) data can independently confirm transcription start sites, and poly(A) site databases can confirm 3′ ends. The tool integrates all of these orthogonal data types to build a multi-layered confidence picture for every transcript.

SQANTI3 also flags transcripts that show signs of specific known artifacts. Internal priming, where the reverse transcriptase latches onto an internal stretch of adenines rather than the true poly(A) tail, is a common problem in long-read RNA-seq. The tool detects these by looking for adenine-rich sequences immediately downstream of the apparent 3′ end. Similarly, transcripts with non-canonical splice sites, unusually short exons, or other red flags get flagged for closer inspection or automatic removal during filtering.

Bringing Orthogonal Evidence Together

The real power of SQANTI3 lies in how it synthesizes independent lines of evidence to evaluate each transcript. A lung cancer study using long-read sequencing recently demonstrated this approach in a clinical context: the researchers identified over 38,000 previously unannotated isoforms in non-small-cell lung cancer samples, then used SQANTI3 alongside CAGE-defined transcription start sites from The Cancer Genome Atlas and poly(A) site databases to confirm the boundaries of those novel transcripts.4American Journal of Respiratory Cell and Molecular Biology. Long-Read Sequencing Reveals Tumor-Specific Splicing Isoforms as Therapeutic Targets in Non–Small-Cell Lung Cancer Without that orthogonal validation step, reporting 38,000 novel isoforms would be hard to take seriously. With it, the researchers could credibly narrow down to 269 tumor-specific splicing events and identify specific exon-skipping and alternative first-exon events absent from existing gene annotations.

This integration philosophy extends to short-read data as well. A mouse transcriptome study fed SQANTI3 not only the long-read Iso-Seq data but also STAR-aligned short-read splice junction files, transcription start site databases, and curated poly(A) site catalogs to evaluate their merged transcriptome.5bioRxiv. Plasticity and evolutionary dynamics of alternative RNA splicing The combination of evidence types lets SQANTI3 build a much more reliable picture than any single data source could provide on its own.

The Rescue Module

Artifact filtering is a balancing act. Set the thresholds too loose and your final transcript set is polluted with false positives. Set them too strict and you lose real biology, particularly lowly expressed isoforms or transcripts from genes with unusual structural features. SQANTI3 addresses this tension with a Rescue module, designed to recover transcripts that initially fail quality filters but still show genuine evidence of expression.6BioRxiv. SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms

The idea is straightforward: if a transcript gets flagged as potentially artifactual but its gene is known to be expressed (supported by read counts or other evidence), discarding it could mean losing the only representative of that gene in the final annotation. The Rescue module checks for this scenario and reinstates transcripts when removing them would leave an expressed gene with no representation at all. It is a safety net that prevents overzealous filtering from creating gaps in the transcriptome.

Detecting Functional Consequences of Novel Transcripts

SQANTI3 does not stop at structural classification and quality filtering. It also evaluates the coding potential of transcript models by predicting open reading frames and checking for features that would trigger nonsense-mediated decay (NMD), a cellular surveillance mechanism that degrades transcripts containing premature stop codons. In the mouse transcriptome study mentioned earlier, novel transcripts with coding potential had a substantially higher probability of being NMD targets compared to known transcripts, with an average NMD probability of 0.17 for novel versus 0.02 for known.5bioRxiv. Plasticity and evolutionary dynamics of alternative RNA splicing

That finding is biologically meaningful. It suggests that many of the novel isoforms captured by long-read sequencing are not just rare variants of known genes but are likely subject to rapid degradation in the cell. Some may represent regulatory “noise” that the cell produces and immediately destroys. Others could be functional in specific contexts, such as particular cell types or stress conditions, where NMD is less active. Either way, knowing which novel transcripts carry premature stop codons helps researchers prioritize which discoveries to follow up on experimentally.

SQANTI3’s Role in Community Benchmarking

The Long-Read RNA-seq Genome Annotation Assessment Project (LRGASP) was a community effort to systematically compare methods for transcript identification and quantification from long-read data. SQANTI3 served as the evaluation framework for the project, providing the structural categories and performance metrics used to assess submissions from participating teams.7PubMed Central. Systematic assessment of long-read RNA-seq methods for transcript identification and quantification The evaluation drew on spike-in controls, simulated data, and manually curated transcript models defined by the GENCODE annotation project to measure how well each pipeline detected real transcripts and avoided false ones.

Using SQANTI3 as a shared evaluation standard accomplished something important: it gave the community a common language and set of metrics to compare results across very different computational pipelines. When one tool produces a transcript set and another produces a different set, you need an independent yardstick to say which is closer to the truth. SQANTI3’s classification system, combined with orthogonal data checks, provided that yardstick.8Nature Methods. Systematic assessment of long-read RNA-seq methods for transcript identification and quantification The LRGASP results showed wide variation in performance across tools, which reinforced the importance of post-processing quality control regardless of which transcript caller you choose.

The Broader SQANTI Ecosystem

SQANTI3 is the main tool, but it has spawned related tools that address specific needs in the long-read RNA-seq workflow. SQANTI-SIM is a simulator that generates controlled benchmark datasets, including simulated short-read Illumina data and CAGE-peak data, so researchers can test their pipelines against a known ground truth.9PubMed Central. SQANTI-SIM: a simulator of controlled transcript novelty for lrRNA-seq benchmark Having simulated data where you know exactly which transcripts are real and which are artifacts is invaluable for tuning filtering parameters and comparing tool performance.

SQANTI-reads takes a different angle. Rather than evaluating transcript models after they have been collapsed and called, it works at the level of individual sequencing reads across replicated experiments. It uses SQANTI3’s structural categories to assess the distribution of reads and unique junction chains across samples, making it easier to spot outlier samples or systematic quality problems in multi-sample designs.10PubMed Central. SQANTI-reads: a tool for the quality assessment of long read data in multi-sample lrRNA-seq experiments Think of SQANTI3 as quality control for your final transcript catalog and SQANTI-reads as quality control for your raw sequencing runs before you even get to transcript calling.

What Alternative Splicing Looks Like Through the SQANTI3 Lens

One of the things SQANTI3 helps researchers quantify is the relative contribution of different splicing mechanisms to isoform diversity. A study of neuronal differentiation used long-read sequencing with SQANTI3-based quality control to map how isoform usage changed as stem cells became neurons. Roughly 55% of the isoform changes came from classical alternative splicing events like exon skipping and intron retention, while the remaining 45% arose from the use of alternative transcription start sites and polyadenylation sites.11PubMed Central. Uncovering the dynamics and consequences of RNA isoform changes during neuronal differentiation

That near-even split is worth noting because alternative start sites and polyadenylation changes are often underappreciated compared to splicing. They do not alter the internal exon structure of a transcript but can dramatically change which protein is produced (different start sites can mean different N-terminal domains) or how stable the mRNA is (different 3′ ends carry different regulatory elements). SQANTI3’s ability to characterize both transcript ends and internal junctions means it captures this full spectrum of isoform variation rather than focusing only on splicing in the narrow sense.

Applications in Disease Research

The lung cancer study mentioned earlier illustrates how SQANTI3-curated transcriptomes feed directly into biomedical discovery. After identifying 269 tumor-specific splicing events, the researchers found 17 with significant associations to specific lung cancer subtypes and 13 enriched across all cases. Specific exon-skipping events in genes like IFI27, PUF60, and ANAPC11, plus an alternative first exon in YBEY, were absent from GENCODE annotations entirely.4American Journal of Respiratory Cell and Molecular Biology. Long-Read Sequencing Reveals Tumor-Specific Splicing Isoforms as Therapeutic Targets in Non–Small-Cell Lung Cancer These are not academic curiosities. Tumor-specific isoforms that are absent from normal tissue are exactly the kind of molecular targets that drug developers look for, whether for antibody-based therapies, splice-switching oligonucleotides, or neoantigen-based immunotherapy.

Without rigorous quality control, these discoveries would be buried in noise. A researcher might report hundreds of novel isoforms in tumor tissue, but without confirming that the transcript boundaries are real and that the splice junctions are supported by independent evidence, the findings would be unreliable. SQANTI3 provides the confidence layer that turns raw long-read output into actionable biology.

Common Misconceptions About Long-Read Transcript Quality

A persistent misunderstanding is that long-read sequencing, because it captures full-length transcripts, produces inherently clean and reliable isoform catalogs. The reality is more complicated. While long reads eliminate the transcript assembly problem, they introduce their own error modes. A single long read that spans an entire transcript can still contain sequencing errors near splice sites, carry artifacts from library preparation, or represent a degraded molecule captured mid-decay. The structural category system in SQANTI3 makes this visible: a transcript classified as NNC (novel not in catalog) might represent a genuinely new isoform, but it might also reflect a sequencing error that shifted one splice site by a few bases into uncharted territory.

Another misconception is that filtering out artifacts is a one-size-fits-all process. The appropriate stringency depends heavily on the biological question. A project cataloging the reference transcriptome of a species wants very high confidence and should filter aggressively, accepting that some real but poorly supported transcripts will be lost. A project looking for rare, condition-specific isoforms in a disease context might use more lenient thresholds and rely on the Rescue module to preserve transcripts from expressed genes. SQANTI3 provides the information to make these decisions but deliberately leaves the final judgment to the researcher rather than imposing a single cutoff.

Practical Inputs and What You Need to Run SQANTI3

Running SQANTI3 requires a few key inputs beyond the transcript models themselves. You need a reference genome and a reference gene annotation (typically from GENCODE or Ensembl). Optionally but strongly recommended, you can supply short-read RNA-seq alignments and splice junction files, CAGE-peak data for transcription start site validation, poly(A) site catalogs, and expression quantification data. The mouse study described earlier provides a good example of a comprehensive input setup: Ensembl gene annotation, STAR-aligned short-read data with junction files, the refTSS database for transcription start sites, and the PolyASite portal for polyadenylation data.5bioRxiv. Plasticity and evolutionary dynamics of alternative RNA splicing

The more orthogonal data you feed in, the more informative the quality assessment becomes. A bare-bones run with just long reads and a reference will still give you structural categories and basic junction checks, but you lose the ability to independently validate transcript boundaries. For studies where the goal is discovering new isoforms rather than simply confirming known ones, that extra validation data makes a meaningful difference in how much you can trust the results.

SQANTI3 produces detailed output tables listing every transcript with its structural category, quality descriptors, junction-level metrics, and coding-potential predictions. These tables are designed to be filtered downstream according to the researcher’s needs. The SQANTI-reads extension adds multi-sample visualizations organized by experimental design factors, which helps catch batch effects or technical outliers before they contaminate the analysis.10PubMed Central. SQANTI-reads: a tool for the quality assessment of long read data in multi-sample lrRNA-seq experiments For large-scale projects with dozens of samples, catching a single problematic library early can save weeks of wasted analysis time.