Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • Brian Bushnell
    Super Moderator
    • Jan 2014
    • 2709

    #16
    Oh, sorry, the command I gave you was for interleaved reads, so those are all false positive merges. For pairs in separate files, the command would be:

    bbmerge.sh in1=R1.fq in2=R2.fastq ihist=ihist.txt reads=100000

    As for whether you merge the reads before assembly, that depends on the assembler. AllPathsLG does its own merging; Ray seems to do well with merged reads; SoapDenovo creates worse assemblies with merged reads; there's no single rule. I've never used Geneious, so I don't know if it would help. As with many options, such as trimming and subsampling, sometimes the only way to get the best assembly is to try both ways.

    If you do merge the reads, I suggest using the default settings:

    bbmerge.sh in1=r1.fq in2=r2.fq out=merged.fq outu1=unmerged1.fq outu2=unmerged2.fq

    ...then feed the assembler both the merged and unmerged reads. Many or most assemblers will accept both paired and unpaired reads; merging should not be done for assemblers that don't allow you to feed them both paired and unpaired reads simultaneously, as low-complexity genomic regions will not merge as well.

    Comment

    • Marisa_Miller
      Member
      • Aug 2010
      • 34

      #17
      Originally posted by Brian Bushnell View Post
      The fraction joined and the position of the peak in the graph will make it clear what the real distribution is like. If the graph is still rising then abruptly drops to zero just before 2x(read length) then the insert sizes are generally too long for merging.
      Hi Brian,
      I re-ran bbmerge with the correct command, and it looks like the distribution shows the insert sizes are too big for merging. Although, the fraction joined is around 60% for most libraries, not sure if this is good or bad.
      Attached Files

      Comment

      • Brian Bushnell
        Super Moderator
        • Jan 2014
        • 2709

        #18


        For a 2x300bp library, getting 60% merging and a median of 400bp is pretty optimal for merging, actually. Again, whether merging is a good idea depends on the assembler, but this library is a good. You never see 90%+ merging unless the insert sizes came out way too short.
        Attached Files

        Comment

        • Marisa_Miller
          Member
          • Aug 2010
          • 34

          #19
          Originally posted by Brian Bushnell View Post


          For a 2x300bp library, getting 60% merging and a median of 400bp is pretty optimal for merging, actually. Again, whether merging is a good idea depends on the assembler, but this library is a good. You never see 90%+ merging unless the insert sizes came out way too short.
          I think I misunderstood earlier about what to look for in a library to see if it can be merged. I will go ahead and merge them and give the assembly a shot with both merged and unmerged. Thanks again!

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM
          • SEQadmin2
            Cancer Drug Resistance: The Lingering Barrier to Rising Survival
            by SEQadmin2



            Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

            There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
            07-08-2026, 05:17 AM
          • GATTACAT
            Reply to Nine Things a Sample Prep Scientist Thinks About Before Sequencing
            by GATTACAT
            Love this - good data definitely starts from good input, and poor input can only give relatively poor data. I particularly like the mention of Nanodrop/absorbance based methods for quantification. It's such a toss up if you'll get an accurate reading or what amounts to a randomly generated number, and a lot of library/sequencing related issues can be traced back to poor quant.
            07-01-2026, 11:43 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, 07-13-2026, 10:26 AM
          0 responses
          27 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-09-2026, 10:04 AM
          0 responses
          37 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-08-2026, 10:08 AM
          0 responses
          24 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-07-2026, 11:05 AM
          0 responses
          34 views
          0 reactions
          Last Post SEQadmin2  
          Working...