What Is a FASTQ File in Bioinformatics and Genomics?

A FASTQ file is a plain-text file that stores the raw output of a DNA or RNA sequencing run: both the sequence of nucleotide bases (the A, T, G, and C letters) and a per-base quality score indicating how confident the sequencing machine was about each letter it called. The format emerged as a practical standard for sharing sequencing reads without any governing body formally defining it, and it remains the starting point for nearly every genomics analysis pipeline today. Its simplicity is part of its appeal, but that simplicity also creates quirks around storage, variant formats, and preprocessing that anyone working with sequencing data will bump into quickly.

Anatomy of a FASTQ File

Every individual sequencing read in a FASTQ file occupies exactly four lines. The first line starts with an “@” symbol followed by a sequence identifier, which often includes information about the instrument, the run, and the position on the sequencing chip where that read was generated. The second line is the actual nucleotide sequence itself, a string of characters like AGTCCTGA. The third line begins with a “+” symbol and sometimes repeats the identifier from line one, though many files leave it blank after the plus sign to save space. The fourth line is the quality string, a line of ASCII characters the same length as the sequence on line two, where each character encodes a score for the corresponding base.

That four-line structure repeats for every read in the file. A single sequencing run on a modern instrument can produce hundreds of millions of reads, so a raw FASTQ file can easily reach tens or hundreds of gigabytes. Despite its size, the format is human-readable: you can open it in a text editor and see the sequences directly, which makes it easy to inspect but expensive to store.

What the Quality Scores Actually Mean

The fourth line in each read block is where FASTQ files distinguish themselves from the older FASTA format, which stores sequences without quality information. Each character in the quality string maps to a number on the Phred scale, which encodes an estimate of the probability that a given base was called incorrectly.1Oxford Academic. GeneCodeq: quality score compression and improved genotyping using a Bayesian framework A Phred score of 10 means there is roughly a 1-in-10 chance the base is wrong. A score of 20 means about 1-in-100. A score of 30 means 1-in-1,000. The scale is logarithmic, so each jump of 10 represents a tenfold improvement in confidence.

The scores are encoded as ASCII characters so that the quality string stays the same length as the sequence string and the whole file remains plain text. The letter “I” might represent a quality of 40 (very high confidence), while a “#” might represent a quality of 2 (essentially untrustworthy). This encoding is compact, but it introduced one of the format’s most persistent headaches: different sequencing platforms chose different ASCII offset values, meaning the same character could represent a different quality score depending on which variant of FASTQ you were looking at.

The Three Incompatible Variants

The FASTQ format was never formally standardized by any official body, and it drifted into at least three incompatible variants. The original Sanger variant uses an ASCII offset of 33, meaning a quality score of zero is encoded as the character “!” (ASCII 33). Early Solexa (later Illumina) instruments used a different offset of 64 and a different quality formula. A later Illumina variant kept the offset of 64 but switched to the standard Phred formula.2PubMed Central. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants

In practice, this meant that mixing files from different instruments, or feeding a file into software expecting the wrong variant, could silently corrupt every quality score in the dataset. The bioinformatics community eventually converged on the Sanger/Phred+33 standard, and modern Illumina instruments now output in that format. Oxford Nanopore’s long-read sequencers also follow the Phred+33 standard, with quality values that can theoretically range from 0 to 93.3Nature Communications. Restoring flowcell type and basecaller configuration from FASTQ files of nanopore sequencing data But older datasets still floating around public repositories may use the legacy Illumina encoding, so anyone downloading archived data should check which variant they are dealing with before running an analysis.

Paired-End Reads and File Conventions

Many modern sequencing experiments produce paired-end reads, where the instrument sequences both ends of a DNA fragment. Instead of one FASTQ file, you get two: commonly named something like sample_R1.fastq and sample_R2.fastq. The first read from the left end of a fragment lives in R1, and its partner from the right end lives in R2, at the same line position in both files. Keeping these two files synchronized matters because downstream alignment tools use the known distance between paired reads to place them more accurately on a reference genome.

If one file gets corrupted or a read is accidentally dropped from one file but not the other, the pairing breaks and the analysis can silently go wrong. This is one of the first things quality-control tools check for, and it is a surprisingly common source of errors when files are transferred between servers or downloaded from public archives.

Quality Control Before Anything Else

Quality control is an essential first step in sequencing data analysis, and the tools built for this purpose have become deeply embedded in standard pipelines at most sequencing centers.4PubMed Central. Falco: high-speed FastQC emulation for quality control of sequencing data The most widely known tool is FastQC, which generates a visual report summarizing the quality profile of a FASTQ file. It checks things like how quality scores are distributed across positions in a read, whether certain bases are overrepresented, what the GC content looks like, and whether adapter sequences have contaminated the reads.5Bioinformatics. FQC Dashboard: integrates FastQC results into a web-based, interactive, and extensible FASTQ quality control tool

A typical red flag in a quality report is a drop in base quality toward the ends of reads. Sequencing chemistry tends to degrade over the length of a read, so the last 20 or 30 bases might have noticeably lower confidence scores than the first 20 or 30. Another common issue is the presence of adapter sequences, short artificial DNA sequences ligated onto fragments during library preparation. If a fragment is shorter than the read length, the sequencer reads past the insert and into the adapter, contaminating the sequence data with non-biological sequence.

Preprocessing and Trimming

After quality control flags the problems, preprocessing tools fix them. This step typically involves removing adapter contamination, filtering out reads whose overall quality is too low to be useful, and sometimes correcting individual bases that appear to have been miscalled.6PubMed Central. Ultrafast one-pass FASTQ data preprocessing, quality control, and deduplication using fastp Tools like fastp can handle all of these tasks in a single pass through the file, which matters when the file is hundreds of gigabytes and reading through it multiple times would take hours.

Adapter trimming deserves special mention because it is not always straightforward. Most trimming tools work by searching for known adapter sequences and clipping them off the ends of reads. But some tools take a different approach: rather than looking for the adapter itself, they identify the portion of the read that matches actual biological sequence and trim everything else.7Bioinformatics. EARRINGS: an efficient and accurate adapter trimmer entails no a priori adapter sequences This can be useful when the adapter sequence is unknown or when multiple adapter types were used in the same experiment.

Demultiplexing and Sorting Reads by Sample

Modern sequencing instruments often run multiple samples at the same time on a single flow cell, distinguished by short barcode sequences (sometimes called index sequences) attached to each sample’s DNA during library preparation. The raw output is one enormous FASTQ file containing reads from all samples mixed together. Demultiplexing is the process of sorting those reads back into separate files, one per sample, by matching the barcode sequence on each read against the known list of sample barcodes.

This matching needs to be tolerant of small errors, since the barcodes themselves can contain sequencing mistakes. Demultiplexing tools typically allow a certain number of mismatches (measured as hamming distance) when comparing a read’s barcode to the expected barcode for each sample. If a read matches one sample closely but no others, it gets assigned to that sample. If it matches two samples equally well, it gets flagged as ambiguous and is usually discarded.8PubMed Central. mgikit: demultiplexing toolkit for MGI fastq files Single-cell sequencing experiments add another layer, where each individual cell has its own barcode, and specialized tools use edit-distance algorithms to match cell barcodes against reference lists numbering in the tens of thousands.9Bioinformatics. Flexiplex: a versatile demultiplexer and search tool for omics data

Where FASTQ Files Go After Preprocessing

Once you have clean, trimmed, demultiplexed FASTQ files, the reads are ready to enter whatever analysis pipeline your experiment calls for. For whole-genome or exome sequencing, the standard workflow aligns reads to a reference genome using tools like BWA, then calls genetic variants using the Genome Analysis Toolkit (GATK).10PubMed Central. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit best practices pipeline For metagenomics or de novo assembly of organisms without a reference genome, the reads feed directly into assemblers that piece overlapping reads together into longer contiguous sequences. Some assembly pipelines can operate directly on raw FASTQ files without any prior processing.11PLOS ONE. An Integrated Pipeline for de Novo Assembly of Microbial Genomes

For RNA sequencing experiments, cleaned FASTQ reads are typically aligned to a reference transcriptome or genome to quantify gene expression levels. Long-read technologies from Oxford Nanopore and PacBio produce FASTQ files with much longer individual reads (thousands to tens of thousands of bases rather than the 150 to 300 bases common with Illumina), and specialized tools have been developed to handle the different error profiles of those longer reads in applications like transcript discovery for single-cell and spatial transcriptomics.12Oxford Academic. A systematic benchmark of bioinformatics methods for single-cell and spatial RNA-seq nanopore long reads data Regardless of the experiment type, the FASTQ file is always the common handoff point between the sequencer and the analysis software.

The Storage Problem

Because FASTQ files are plain text with highly repetitive content, they compress well compared to most file types. But even compressed, the sheer volume of data produced by modern sequencing centers is staggering. A single human genome sequenced at standard depth can produce FASTQ files totaling 50 to 100 gigabytes before compression. Multiply that by the thousands of genomes sequenced in large cohort studies, and you are looking at petabytes of raw FASTQ data.

This has driven the development of specialized compression tools that go beyond general-purpose methods like gzip. These specialized compressors exploit the specific structure of FASTQ data: the fact that sequence data is drawn from a four-letter alphabet, that quality scores follow predictable patterns, and that the identifier lines share a lot of repeated structure across reads.13PubMed Central. FastqCA: an effective FASTQ compressor through 2D spatial redundancy reduction In practice, many labs simply store gzip-compressed FASTQ files (recognizable by their .fastq.gz extension) because most bioinformatics tools can read them directly. The more aggressive compression schemes save more space but add complexity.

Some researchers and institutions have moved toward storing data in the more compact BAM or CRAM formats, which represent aligned reads rather than raw sequences. The trade-off is that these formats encode the alignment to a specific reference genome, so going back to the raw data requires either keeping the original FASTQ files or accepting some irreversible data loss. For archival purposes, FASTQ remains the most universal raw format.

Public Repositories and Data Sharing

The Sequence Read Archive (SRA), maintained by NCBI, is the largest public repository of sequencing data and stores raw reads from platforms including Illumina, Roche 454, PacBio, Oxford Nanopore, and others.14PubMed Central. SRAdb: query and use public next-generation sequencing data from within R When a study is published, journals increasingly require that the underlying sequencing data be deposited in SRA or an equivalent repository (like the European Nucleotide Archive), making it possible for other researchers to reanalyze the raw FASTQ files and reproduce or challenge the original findings.

Downloading data from SRA is notoriously cumbersome. The archive stores data in its own SRA format and converts it back to FASTQ on the fly using a tool called fastq-dump (or its faster successor, fasterq-dump). Several community tools have been built to simplify the process, including utilities that let you query metadata and download FASTQ files programmatically.15PubMed Central. pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive Still, downloading a large dataset can take hours or days depending on your bandwidth, and corrupted downloads that break paired-end file synchronization are a frequent frustration.

Privacy Risks in Raw Sequencing Data

Something that often surprises people outside the field is that FASTQ files can carry more personal information than the experiment intended to reveal. A gene expression study, for example, generates RNA sequencing data to measure which genes are turned on or off. But the raw FASTQ reads from that experiment also contain the donor’s genomic variants, including single-nucleotide differences that could theoretically be used to re-identify an anonymous participant. Raw and minimally processed data formats like FASTQ contain information that is not of primary research interest but can be exploited for reidentification attacks.16PubMed Central. Assessing Privacy Vulnerabilities in Genetic Data Sets: Scoping Review

The saving grace, to some extent, is that extracting variant information from raw FASTQ files requires meaningful bioinformatics expertise and computational resources, creating a higher barrier than working with pre-processed variant call files. But the risk is real enough that data access committees and institutional review boards increasingly scrutinize how raw sequencing data is shared, and controlled-access repositories (where researchers must apply for permission to download data) have become the norm for human genomic studies. If you are depositing FASTQ files from a human study, checking with your ethics board about what level of access control is appropriate is not optional overhead: it is a necessary step.

Common Misconceptions About FASTQ Files

One common misunderstanding is that a FASTQ file contains a genome. It does not. It contains millions of short (or long) overlapping fragments of a genome, a transcriptome, a metagenome, or whatever was sequenced. Assembling those fragments into a coherent genome, or aligning them to a known reference, is a separate computational step that happens downstream. The FASTQ file is the raw ingredient, not the finished product.

Another misconception is that quality scores are objective measurements of accuracy. They are actually estimates produced by the base-calling software, and different versions of the same base caller can assign different quality scores to identical raw signal data. Oxford Nanopore data illustrates this clearly: their Guppy base caller assigns quality values ranging from 1 to 90, while the newer Dorado base caller restricts the range from 1 to 50 for the same underlying signal.3Nature Communications. Restoring flowcell type and basecaller configuration from FASTQ files of nanopore sequencing data The quality scores are useful, but treating them as ground truth rather than as calibrated guesses can lead to overconfident variant calls or overly aggressive quality filtering.

A third misconception, mostly among people who are new to sequencing, is that FASTQ files are interchangeable regardless of the sequencing platform. While the four-line format is the same, the error profiles and read lengths differ dramatically. Illumina reads are short with occasional single-base substitution errors. Nanopore reads are long with more frequent insertion and deletion errors. Feeding one platform’s data into a tool designed for the other without adjusting parameters can produce misleading results, even though the input files look structurally identical.

When FASTQ Files Are Not the Right Format

Despite their ubiquity, FASTQ files are not always the ideal choice. Once reads have been aligned to a reference genome, the standard practice is to convert the data into BAM format (or its more compressed cousin, CRAM), which stores the alignment coordinates alongside the sequence and quality data. Keeping data in FASTQ format after alignment would throw away the computationally expensive mapping work and leave you with raw reads you would have to re-align later.

For long-term archival of very large datasets, some institutions have moved to compressed reference-based formats that store only the differences between each read and the reference genome, dramatically reducing file sizes. And for data sharing at the variant level, researchers typically share VCF (variant call format) files rather than FASTQ files, since VCF files are orders of magnitude smaller and contain the biologically interesting findings without the overhead of billions of individual reads.

FASTQ’s role is specifically as the lingua franca of raw sequencing output. It occupies the narrow but critical window between the sequencing instrument and the first computational step. Everything upstream (the chemistry, the optics, the base-calling algorithm) produces it, and everything downstream (alignment, assembly, variant calling, expression quantification) consumes it. That positioning, combined with the fact that the format is simple enough to inspect by eye, is why it has persisted for well over a decade despite the compression headaches and the variant-encoding confusion that dogged its early years.