Flye is a long-read genome assembler that builds assemblies by constructing repeat graphs from raw, error-prone sequencing reads, and multiple independent benchmarks have ranked it among the top-performing tools for both Oxford Nanopore (ONT) and PacBio CLR data. Its distinctive approach to handling repetitive DNA, which is the single biggest obstacle in genome assembly, has made it a go-to choice across fields from microbial genomics to large eukaryotic projects. But Flye’s strengths are not universal, and understanding where it excels and where other tools pull ahead matters if you are choosing an assembler for your own data.
How Flye Approaches Assembly Differently
Most long-read assemblers work by finding overlaps between reads and building a layout from those overlaps. Flye takes a different path. It first concatenates reads into long, error-filled sequences called disjointigs, which trace arbitrary paths through the genome’s repeat structure. From these messy sequences, Flye constructs what is known as a repeat graph, a structure that explicitly represents repeated regions of the genome rather than trying to collapse or ignore them.1PubMed Central. Assembly of long, error-prone reads using repeat graphs Once the repeat graph is built, Flye resolves its tangles to produce final contiguous sequences. This graph-first strategy is what distinguishes Flye from overlap-layout-consensus assemblers and gives it an edge in genomes packed with repetitive elements.
Flye also has built-in options for recovering small plasmids and handling uneven coverage depth, the latter being critical for metagenomic samples where different organisms are present at wildly different abundances.2PubMed Central. Benchmarking of long-read assemblers for prokaryote whole genome sequencing – Section: Assemblers tested These are not afterthoughts bolted onto the software; they are design choices baked into its pipeline.
Benchmarking Against Other Assemblers
When Flye was first published, its developers reported that it nearly doubled the contiguity of a human genome assembly compared with existing assemblers, as measured by the NGA50 metric, while running an order of magnitude faster than competitors.3PubMed. Assembly of long, error-prone reads using repeat graphs Since then, independent groups have put Flye through extensive head-to-head tests. A broad evaluation of long-read assemblers for eukaryotic genomes concluded that Flye was the best-performing assembler overall for both PacBio CLR and ONT reads, on both real and simulated datasets.4PubMed Central. Evaluating long-read de novo assembly tools for eukaryotic genomes: insights and considerations – Section: CONCLUSIONS A yeast genome benchmarking study found the same pattern for ONT data specifically, with Flye achieving the highest composite quality scores among the tools tested.5Briefings in Bioinformatics. Benchmarking of long-read sequencing, assemblers and polishers for yeast genome – Section: Results and discussion
For plant and crop genomes, an evaluation described Flye and Canu as the heavyweight tools that should be the first choice when the goal is an accurate and complete assembly.6PubMed. Comparative Evaluation of Genome Assemblers from Long-Read Sequencing for Plants and Crops A molluscan genome study offered a more nuanced recommendation: Flye worked best for compact, less heterozygous genomes, while NextDenovo outperformed it on more repetitive and heterozygous ones, with roughly 40 to 50× ONT coverage being sufficient for high-quality results in either case.7PubMed Central. Benchmarking Oxford Nanopore read assemblers for high-quality molluscan genomes – Section: Abstract
Where Flye Falls Short
Flye’s dominance with error-prone long reads does not carry over cleanly to PacBio HiFi data, which are much more accurate to begin with. In a benchmark using HiFi reads for maize and a human cell line, Flye produced the least contiguous assemblies among the tools tested. For the human dataset, Flye’s N50 was about 29 megabases compared with roughly 87 megabases for Hifiasm. However, there was a tradeoff: Flye assemblies consistently had the fewest total errors, including as few as three insertion or deletion errors per 100,000 bases.8PubMed Central. New algorithms for accurate and efficient de novo genome assembly from long DNA sequencing reads – Section: Benchmark with PacBio HiFi data So if your priority is accuracy over contiguity and you are working with HiFi reads, Flye still has something to offer, but for the most contiguous HiFi assemblies, tools like Hifiasm are the better bet.
A recent agricultural species benchmarking study reinforced this picture. Most ONT and CLR assemblies, including those from Flye on ONT data, had k-mer completeness below 95%, reflecting the small errors that creep in from noisy input reads. CLR Flye was a notable exception, suggesting that Flye’s internal error-correction pipeline works better with CLR chemistry. HiFi assemblies from dedicated HiFi tools consistently achieved higher quality scores and fewer errors overall.9bioRxiv. Benchmarking long-read genome assemblers for three sequencing protocols and three agricultural species – Section: Results
Highly Heterozygous Genomes
Diploid organisms, and especially those with high heterozygosity, present a particular challenge. When the two copies of a chromosome differ substantially, an assembler can inflate the assembly size by treating the two haplotypes as separate sequences rather than collapsing them. In a study of genomes with varying heterozygosity levels, Flye’s assembly of the Pacific oyster genome (heterozygosity around 3%) exceeded an assembly ploidy of 2, meaning it was assembling both haplotypes rather than producing a single reference with the alternate haplotype represented separately.10Briefings in Bioinformatics. A practical assembly guideline for genomes with various levels of heterozygosity – Section: RESULTS For highly heterozygous organisms, you may need a haplotype-aware assembler like PECAT, which explicitly identifies and separates haplotype-specific overlaps to produce phased assemblies.11PubMed Central. De novo diploid genome assembly using long noisy reads – Section: Result
This is not a fatal flaw so much as a design limitation. Flye was built to handle repeats, not to phase haplotypes. If your organism has low to moderate heterozygosity, Flye handles it well. If your genome looks more like an oyster’s, you will want a different primary assembler or a downstream phasing step.
Metagenomic Assembly With metaFlye
One of Flye’s most significant extensions is metaFlye, its metagenome assembly mode. Environmental and clinical samples routinely contain dozens to hundreds of microbial species at vastly different abundances, plus closely related strains that share large stretches of sequence. The repeat graph framework is well suited to this: shared conserved sequences between related strains produce bubble structures in the graph, and metaFlye detects and simplifies these strain-induced subgraphs to produce separate, contiguous assemblies for each genome in the community.12PubMed Central. metaFlye: scalable long-read metagenome assembly using repeat graphs – Section: Assembling multiple closely-related bacterial genomes
The practical upshot is that metaFlye can take a complex microbial community sample sequenced on a long-read platform and pull out individual genomes, including those of closely related species that would be hopelessly tangled in a short-read assembly. This has become especially important for clinical microbiology and environmental monitoring, where knowing which specific strains are present and what genes they carry is the whole point.
Tracking Drug Resistance in Pathogens
Flye has carved out a particularly strong niche in antimicrobial resistance (AMR) surveillance. A study comparing assemblers for clinical isolates found that Flye was the best tool for detecting and correctly locating both virulence genes and AMR genes, distinguishing whether resistance genes sat on the chromosome or on mobile plasmids.13PubMed Central. The impact of applying various de novo assembly and correction tools on the identification of genome characterization, drug resistance, and virulence factors of clinical isolates using ONT sequencing – Section: Abstract That chromosomal-versus-plasmid distinction matters a great deal in epidemiology, because plasmid-borne resistance genes can transfer between species and are a much higher public health concern than chromosomal mutations.
A comparison of DNA extraction kits and assemblers for detecting AMR genes in two clinically important species found that Flye outperformed Unicycler across all workflows, boosting detection by 2 to 14 percentage points. The best-performing combination achieved about 95% detection of AMR determinants, compared with roughly 68% for the lowest-performing pipeline, with the largest gap seen in efflux pump genes.14PubMed. Comparison of three commercial DNA extraction kits and assemblers for AMR determinant detection in Pseudomonas aeruginosa and Enterobacter cloacae using long-read sequencing – Section: RESULTS These are not abstract benchmarking numbers. Missing efflux pump genes means underestimating a pathogen’s ability to pump antibiotics out of its cells, which directly affects treatment decisions.
Flye has also been incorporated into genomic surveillance pipelines for multidrug-resistant organisms alongside other assemblers, reflecting its status as a standard tool in the field.15PubMed Central. Genomic surveillance of multidrug-resistant organisms based on long-read sequencing – Section: METHODS
The Plasmid Recovery Problem
Small circular DNA molecules like bacterial plasmids are chronically underassembled by long-read tools, and Flye is no exception. A study that tested multiple assemblers found that Flye-raw recovered about 91% of plasmid sequences, Flye-meta about 90%, and Flye-hq about 88%. For comparison, Unicycler recovered 100% of plasmids, and Canu managed 96%. Recovery rates depended heavily on plasmid size and bacterial species, and all long-read-only assemblers except Unicycler, which incorporates short reads, struggled with the smallest plasmids.16PubMed Central. Long read genome assemblers struggle with small plasmids – Section: Results and discussion
This gap matters because small plasmids often carry critical resistance genes. If your assembly misses a 2-kilobase plasmid harboring a carbapenem resistance gene, your surveillance pipeline has a blind spot. For projects where complete plasmid recovery is essential, a hybrid approach combining long and short reads, or using Unicycler, remains safer. When working with Flye alone, be aware that some small circular elements may drop out.
Computational Resources and Speed
Flye is fast, often dramatically so compared with tools like Canu, which can take hours to complete assemblies that Flye finishes in a fraction of the time.17PubMed Central. Benchmarking of long-read assemblers for prokaryote whole genome sequencing – Section: Results and discussion That speed advantage was part of its original design goal and remains one of its most practical benefits, especially for labs processing many samples.
The tradeoff is memory. The same prokaryote benchmarking study found that Flye had the highest RAM usage among the assemblers tested, with memory consumption scaling with both read length and dataset size. For bacterial genomes this is manageable, but for large eukaryotic assemblies it can push hardware requirements into the hundreds of gigabytes. A separate project demonstrated that re-engineering Flye’s data structures could reduce memory consumption by 22% to 47% without affecting assembly results, and could also cut processing time by up to 25%.18PubMed. Memory-Efficient Assembly Using Flye This kind of optimization is relevant for labs running assemblies on shared computing clusters or institutional servers with fixed memory limits.
Polishing the Output
A Flye assembly from noisy long reads is not the final product. Post-assembly polishing, which aligns reads back to the draft assembly and corrects remaining errors, is a standard step. A typical workflow uses multiple rounds of Racon polishing with the original long reads, followed by Medaka for ONT-specific error correction. In one case, four rounds of Racon followed by Medaka brought a Flye assembly of the black carpenter ant genome up to a BUSCO completeness score of 97%, across a 309-megabase assembly in 1,625 contigs.19Nucleic Acids Research. De novo sequencing, diploid assembly, and annotation of the black carpenter ant, Camponotus pennsylvanicus, and its symbionts by one person for $1000, using nanopore sequencing – Section: MATERIALS AND METHODS
With the newest high-accuracy ONT chemistry and basecalling models, the polishing burden has shrunk considerably. An optimized workflow using Nanopore R10.4.1 flow cells with super-accurate basecalling, followed by Flye assembly and Medaka polishing, achieved an error rate of just 0.0007% for bacterial genomes, with 100% concordance in downstream typing and resistance gene detection compared with gold-standard hybrid assemblies.20PubMed Central. Optimized De novo assembly of Haemophilus influenzae using oxford nanopore sequencing without short-read data on a graphical user interface-based galaxy platform That level of accuracy from long reads alone would have been unthinkable a few years ago, and it means that for many bacterial projects, a Flye-plus-Medaka pipeline can now replace hybrid assembly entirely.
Using Multiple Assemblers in Practice
The genomics community has settled into a practical consensus that no single assembler is best for every genome. A study assembling fungal genomes with Nanopore data illustrates why. Flye avoided certain misassemblies that a different tool introduced, but it created its own artifacts in another genome, including an erroneous fusion of contigs that appeared to result from incorrect handling of telomeric repeats.21GigaScience. Long-read only assembly of Drechmeria coniospora genomes reveals widespread chromosome plasticity and illustrates the limitations of current nanopore methods – Section: Results The researchers’ conclusion, which echoes the broader field’s practice, was that using more than one assembler and comparing results is the safest approach.
This is especially true for organisms where no closely related reference genome exists. Without a reference to validate against, you cannot be sure which assembler got the right answer in a disputed region. Running Flye alongside a tool with a different algorithmic approach, then looking for disagreements, is one of the most reliable ways to catch assembly errors before they propagate into downstream analyses.
Organelle Genomes and Specialized Applications
Plant mitochondrial genomes are notoriously difficult to assemble. They are riddled with large repeats, exist in multiple structural configurations within a single cell, and share sequences with the nuclear and chloroplast genomes. A recent evaluation of 13 assembly tools for plant mitochondrial genomes found that while some tools excelled at contiguity and completeness, Flye and GetOrganelle stood out specifically for correctness, producing assemblies with fewer errors in their final sequences.22PubMed Central. Advance in the assembly of the plant mitochondrial genomes using high-throughput DNA sequencing data of total cellular DNAs For researchers who need a trustworthy organelle assembly and plan to curate it manually, starting with a correct but possibly fragmented Flye assembly is a defensible strategy.
Flye has also found use in settings far from its original design. Galaxy-based graphical pipelines have wrapped Flye into point-and-click workflows, making it accessible to researchers without command-line experience. These platforms handle read filtering, assembly, and polishing in a single automated pipeline, lowering the barrier to entry for clinical and teaching labs that want to generate genome assemblies from Nanopore data without scripting every step themselves.