I’m not a sequencing expert. I’m a purification scientist who uses NGS to evaluate workflows my group develops. With this perspective, we think about the sample first and the NGS workflow second. The sequencer is an exceptionally honest reporter, but it can only report on what you give it, so whether you get clean, interpretable data from an NGS workflow is largely determined before you begin.
Here are nine questions we think about, in roughly the order they matter, before sequencing our samples.
What’s actually in the tube?
1. What measurement methods are you using to quantitate your sample, and are they appropriate?
Absorbance-based methods like NanoDrop are the most used in nucleic acid workflows. Absorbance is quick and cheap, but it will measure single nucleotides, degraded nucleic acid fragments, and other molecules left over from the purification chemistry. Nucleic-acid selective dyes give more precise data for sequencing, and amplification-based quantification is often an even better choice. Quantitation data from fluorometry and PCR can differ, and PCR is often the better predictor of library prep success. Best practice is to align QC with the actual workflow. For example, if your sample is going to be template for a polymerase, QC it with a polymerase. There are PCR-based quantification kits on the market built specifically for this purpose.
2. What other analytes may be present that interfere with measurement of the target?
Common contaminants include genomic DNA contamination in an RNA prep, host DNA in a microbiome workflow, and ribosomal RNA in a transcriptomics experiment. Ideally, your quantification method should distinguish between the target analyte and off-target sequences present in the sample.
3. What is the absolute yield of the analyte that’s the ultimate target for library prep and sequencing?
This is often expressed in mass (e.g., nanograms) but can also be refined to number of molecules (e.g., copy numbers, genomic equivalents, etc.). Think about your number in context of what else might be present in your sample and the library input recommendations of the NGS workflow.
4. Do you need depletion or enrichment to improve the signal-to-noise of the target?
RNA without ribodepletion is mostly rRNA reads. Whole blood RNA without globin depletion is mostly globin transcripts. Lysed-cell genomic DNA contamination in plasma cfDNA reduces variant frequency. Low microbial inputs are masked in the presence of high abundance and ubiquitous microbial sequences. Knowing whether you need to enrich or deplete is often the difference between a productive run and a wasted one.
What happened to the sample on its way to you?
5. Does the eluate contain co-purified inhibitors that may interfere with library prep enzymes or reactions?
Inhibitors come from two sources: the sample matrix (humic acids from stool, heme from blood, urea from urine, polysaccharides from plants) and the purification chemistry itself (residual chaotropes, detergents, organics). This is one of the underappreciated advantages of commercial purification technologies: Providers evaluate their formulations against downstream NGS applications to confirm compatibility with library prep enzymology. This is possible with in-house methods, but the burden rests entirely on the lab.
6. Did any step during purification introduce damage or quality issues?
FFPE decrosslinking is a classic example. The high temperature conditions that reverse nucleic acid crosslinks in formalin-fixed tissue also drive cytosine deamination. These deamination events present as C-to-T artifacts and can be mistaken as low-frequency variants. Temperature and denaturing agents can also change the strandedness or conformation of the nucleic acid, and depending on your library prep, can bias reads.
7. Did sample storage introduce damage?
Use of proper storage buffers is critical. Typically, these include a buffer to control pH plus a low concentration of chelator like EDTA as insurance against nucleases. Avoid repeated freeze-thaws. Damage from freeze-thaws can include fragmentation and degradation at the 5' and 3' ends of fragments, which negatively impact library prep methods that depend on intact molecules for adapter ligation. Finally, an often-underappreciated failure point is an unsealed lid, which over time can introduce CO2 exposure to the sample, acidifying the eluate and accelerating hydrolysis of the phosphodiester backbone.
Other factors to consider
8. Did the purification introduce bias in the analyte?
Some purification mechanisms have differential binding affinity for GC-rich versus AT-rich nucleic acid polymers, or for different molecular weight ranges, or for different nucleic acid conformations. The result is a sample that no longer faithfully represents what was in the sample. This kind of bias often goes undetected with most quantification methods. Quantification may look fine, the library may look fine, but coverage uniformity or variant calls may skew in ways that are hard to trace back upstream.
9. Was your sample cross-contaminated or mixed up?
Cross-contamination or sample mix-up is a real concern, especially in high-throughput manual workflows. Automated nucleic acid purification systems help significantly, both by reducing handling and by producing audit trails. As a process check, we will often run an every-other or checkerboard pattern of XY and XX chromosome samples to monitor for cross-contamination.
Looking forward
A lot of pre-library enrichment and depletion strategies have existed for years. For example, poly-A selection, ribodepletion, CRISPR-based depletion of unwanted targets. What’s changing is that the chemistries are getting more specific and the workflows are getting more compatible with low-input samples.
From where I sit, the most interesting question in sample prep over the next few years isn’t “How do we extract more efficiently?” Instead, we’re asking “How do we deliver a sample that's already enriched for what the sequencer should be reading?”
The nine factors outlined above are not glamorous. But the variability that exists in NGS workflows often lives disproportionately upstream of the instrument, and if you’re troubleshooting unexpected results and the run metrics look fine, the answer almost always lies within your sample.
Kevin Mayer, Ph.D., is a Senior Research Scientist in the Protein and Nucleic Acid Analysis group at Promega Corporation. He develops and optimizes nucleic acid purification chemistries and workflows, with a particular focus on difficult oncology samples, for use on Promega's Maxwell automated extraction platforms. Before joining Promega, Kevin received his Ph.D. in Genetics from the University of Wisconsin-Madison, where he studied genetic pathways underlying plant flowering time.