The Hidden DNA Quality Bottleneck in Long-Read Sequencing

Long-read sequencing performance begins with DNA fragment-size distribution. Learn how short fragments arise and why post-isolation DNA preparation matters.


Long-read performance starts before library preparation

Long-read sequencing has changed what researchers can see in a genome. Longer reads can span repetitive regions, improve de novo assembly, and support the detection of structural variants that are difficult to resolve using short-read approaches. Yet the ability to generate long reads depends on more than the sequencing platform. It depends heavily on the physical condition of the DNA entering library preparation.

In practice, DNA samples rarely consist of uniformly long molecules. Even when a laboratory uses a high-molecular-weight (HMW) DNA isolation method, the resulting material may contain a broad distribution of fragment sizes. That distribution can become even more heterogeneous as the sample moves through routine laboratory operations.


How short fragments enter an HMW DNA sample

DNA is vulnerable to mechanical and environmental stress. Pipetting, vortexing, repeated transfers, freeze-thaw cycles, prolonged storage, and transportation can all contribute to fragmentation. The magnitude of that effect varies with sample type, DNA concentration, handling technique, storage conditions, and the number of processing steps.

The result is often a mixed population: some molecules remain long enough to support long-read applications, while others have been reduced to much shorter fragments. A sample can therefore meet conventional concentration and purity specifications while still having a fragment-size profile that is not well matched to the sequencing objective.


Why the distribution matters

Library preparation does not automatically distinguish between every desirable long molecule and every undesirable short one. Shorter fragments may enter the library and compete with longer molecules during downstream processing and sequencing. As their relative abundance increases, the library may generate a larger proportion of shorter reads than the research question requires.

This is particularly important for applications that benefit directly from long molecules, including structural variant analysis, genome assembly, phasing, repetitive-region analysis, and ultra-long-read sequencing. Concentration and absorbance ratios remain useful quality indicators, but they do not provide a complete picture. Fragment-size distribution is another critical part of sample readiness. 


Isolation and post-isolation optimization are different steps

HMW DNA isolation and DNA size redistribution solve related but distinct problems. An HMW isolation kit is designed to recover long DNA from a biological sample. A post-isolation purification and size-redistribution reagent is used afterward to adjust the fragment population by reducing shorter DNA molecules and enriching longer ones.

SFD-HT (Short Fragment Depletor–High Throughput) is being developed by MagBio Genomics for this second role. It is a patent-pending, magnetic bead-based DNA purification and size-redistribution reagent applied after HMW DNA or genomic DNA isolation and before long-read library preparation.


A better question for sample QC

Instead of asking only, ‘How much DNA do we have?’ laboratories may also need to ask, ‘What fragment-size distribution are we placing into library preparation?’ That shift in perspective can help identify a hidden source of variability before sequencing resources are committed.

As long-read sequencing becomes more widely adopted, upstream workflows will increasingly need to protect long molecules and manage unwanted short-fragment carryover. Better prepared DNA provides a stronger foundation for generating the read-length profile a study is designed to achieve.

A recent MagBio study illustrates why conventional QC is only part of the story. Human-blood gDNA isolated with the HighPrep Blood & Tissue DNA Kit entered the experiment with a DQN of 9.2, and N50 of 12.6 kb. After a 0.40x SFD-HT treatment configured to deplete fragments below 10 kb, it improved to a DQN of 9.8, and N50 of 27.8 kb. In conclusion, even with strong input metrics (DQN 9.2; 91.8% in the 10–16.6 kb window), the read-length profile still skewed short. After SFD-HT’s <10 kb cleanup and size selection, the size distribution shifted toward longer fragments, increasing the long-fragment fraction by 6% (91.8% to 97.8%) and boosting N50 by ~121% (12.6 kb to 27.8 kb). Therefore, supporting high recovery and optimizing libraries for long-read sequencing.


Interested in testing SFD-HT?

Learn more about the technology and evaluation opportunities through MagBio Genomics’ SFD-HT Early Adopter & Collaboration Program.



Quick Links: