After sequencing a genome, the next step is to assemble the millions or even billions of DNA fragments into a complete and continuous sequence. Given the size and complexity, like the ~3 billion bases in the human genome, constructing this sequence presents substantial challenges. This article reviews the current state of genome assembly, highlighting key technical obstacles, recent advancements in sequencing technologies and assembly tools, and emerging trends influencing the future of the field.
Current Challenges in Genome Assembly
Arranging raw sequencing reads into a complete genome has notoriously been a difficult process, and despite recent developments, several challenges still make genome assembly far from straightforward. Rayan Chikhi, Ph.D., Bioinformatics and Group Leader at Institut Pasteur, explained that genome assembly challenges differ depending on the type of genome. His group focuses on metagenome assembly and identifies two major areas of difficulty: sequencing data quality and computational algorithms.
For the first challenge, Chikhi noted that while sequencing data has improved significantly, it is still not consistently sufficient to fully resolve all genomes. On the computational side, challenges include producing high-quality assemblies that are accurate, contiguous, and haplotype resolved, as well as ensuring that tools are computationally efficient and accessible to labs with limited resources. Although significant progress has been made, Chikhi shared that many tools prioritize quality over efficiency. His lab has taken time to address this tradeoff by developing more resource-efficient tools that still produce reliable results.
Rasmus Kirkegaard, staff scientist at Aalborg University, added that genomic repeats have presented a big challenge for genome assembly for many years. These repeats are sequences that occur multiple times in the genome, either in tandem or scattered across different regions1. Because they are often nearly identical, it becomes difficult for assembly algorithms to determine their correct placement, especially when using short-read data.
In addition to these challenges, Ibrahim Jivanjee, Senior Director of Product Management and Marketing at Arima Genomics, outlined several other hurdles. These include difficulty achieving chromosome-scale continuity, even with long-read data, as contigs often remain unlinked or poorly ordered. Accurately phasing haplotypes in diploid and polyploid organisms is another key issue, especially for studying allele-specific features. The computational demands of assembling large genomes add further complications, requiring significant processing power and memory. Lastly, Jivanjee noted that ensuring base-pair accuracy and complete genome coverage, particularly in hard-to-sequence regions, continues to be a major challenge.
Advances Driving Progress
Fortunately for researchers, genome assembly has benefited from numerous developments in data quality, techniques, and algorithmic design. One key recent advancement has been the significant improvement in read accuracy. Chikhi emphasized how PacBio and Oxford Nanopore Technologies (ONT) now produce long reads with much lower error rates. This shift toward highly accurate, longer read lengths has improved assembly practices and enabled more efficient methods. “These technology developments are really transforming the way assemblies are being done,” stated Chikhi. He noted that this increase in accuracy has been critical for human genome assembly as well as advancements in metagenome assembly.
Similarly, Kirkegaard noted how the improved quality of Nanopore reads has enabled the research community to create Nanopore-only-based genomes. He also pointed out that affordable and accessible Nanopore sequencing has allowed many labs to produce closed, high-quality genomes. Jivanjee also discussed the improvements from long-read technologies, noting how they have made it possible to resolve complex genomic regions and achieve near-complete assemblies, as well as supporting large-scale efforts like the T2T consortium, the Vertebrate Genomes Project (VGP), and the Darwin Tree of Life (DToL) project.
Along with long-read sequencing, Jivanjee highlighted the important role of Hi-C technology in improving genome assembly. Hi-C helps order and orient contigs into larger scaffolds, enabling chromosome-scale assemblies. It identifies misassemblies by revealing structural inconsistencies through 3D contact maps. The technology aids in anchoring contigs to specific chromosomes by locating centromeric and telomeric regions and supports haplotype phasing by grouping contigs based on interaction patterns. Additionally, Hi-C maps serve as a visual quality check for detecting structural issues, and the technology pairs effectively with long-read sequencing. “When combined with long-read sequencing, Hi-C provides a powerful approach for achieving high-quality, chromosome-scale assemblies, bridging gaps left by previous genome projects,” he stated.
Tools for High-Quality Assembly
Successful genome assembly depends heavily on selecting the right tools for the task. Over the years, a wide range of software has been developed to support various assembly goals and sample types. Chikhi highlighted several key tools for producing high-quality assemblies. For human genomes, he recommended Verkko and HiFiasm, both of which support telomere-to-telomere assemblies and perform efficiently across genome types. He also suggested LJA, a lesser-known tool that introduced valuable ideas to the field. For metagenomics, Chikhi recommended using MetaMDBG, developed by his group, which generates high-quality assemblies from long-read data. Additionally, he acknowledged ongoing work on NanoMDBG for Nanopore data, and mentioned Metaflye as another option for Nanopore-based metagenomics.
“We are particularly fond of Flye, as it is fast and produces genomes of very high quality from both Nanopore and PacBio reads,” stated Kirkegaard. This versatile genome assembly tool, developed by the same team behind Metaflye, serves as a general-purpose assembler. It supports a wide range of datasets, from small bacterial genomes to large mammalian assemblies, and includes a dedicated mode for metagenome assembly. Building on this, Jivanjee highlighted the value of learning from established pipelines developed by the VGP and Sanger Institute, as well as taking advantage of training resources available through the Galaxy platform.
Tips for Complex Genome Assembly
Even when a reference genome is available, assembling complex genomes remains a formidable task. Here, our experts share key strategies to navigate the process. Jivanjee urges researchers just learning how to perform genome assemblies to step back and thoughtfully plan their strategy. “When budgeting for a de novo assembly project, we have found that many researchers believe that long-read sequencing is enough,” he stated. “This may be true for some smaller genomes (like bacteria), but is often not the case for larger, more complex genomes.” In most cases, additional data types are required, including Hi-C for scaffolding, RNA-seq for gene annotation, and reliable computational tools like Galaxy.
For non-model organisms, Chikhi observed that assembly is more complex, owing to characteristics such as increased heterozygosity and polyploidy. He pointed to projects like VGP and DToL as leading examples for addressing these challenges with reliable sequencing pipelines. These consortia have published guidelines to help researchers produce high-quality assemblies. However, Chikhi noted that this process is far from automated and requires substantial expertise. He anticipates that continual improvements in read length and accuracy will lead to more efficient and user-friendly tools.
Emerging Trends and Future Directions
Ongoing innovations in lab-based methods and bioinformatics are expected to open new possibilities for genome assembly. Kirkegaard believes that developing reliable methods to extract and process DNA fragments longer than 100 kilobases for Nanopore sequencing will significantly impact the future of genome assembly, and notes that many labs and companies are actively working on this advancement. Additionally, Kirkegaard emphasized the importance of the recently developed hifiasm (ONT) from the Heng Li Lab. This method enables efficient near telomere-to-telomere genome assembly using standard ONT Simplex reads, while eliminating costly ultra-long reads and reducing computational demands. As for other promising areas within the field, Jivanjee expressed enthusiasm for the various large-scale initiatives, citing their potential to significantly benefit both human and environmental health.
Lastly, Chikhi explained that future improvements in genome assembly will likely come from a combination of factors, including longer and more accurate reads, reduced sequencing costs, and more efficient computational tools. He noted that while high-quality assemblies are often prioritized, balancing quality with efficiency is essential to make assembly tools more accessible. “A dual focus on assembly quality and efficiency will help the community greatly,” noted Chikhi. “But I must admit this is very hard to achieve.” His group is also exploring environmental metagenomics, where sequencing genomes from soil and ocean samples presents unique challenges due to their high complexity and diversity. Current sequencing efforts often lack the depth needed to fully reconstruct these communities, and most tools cannot handle the scale of the data. Chikhi sees this area as a promising direction for future research.
References
- Liao, X., Zhu, W., Zhou, J. et al. Repetitive DNA sequence detection and its role in the human genome. Commun Biol 6, 954 (2023). https://doi.org/10.1038/s42003-023-05322-y