Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • Brian Bushnell
    Super Moderator
    • Jan 2014
    • 2709

    #16
    In that case, mapping the co-normalized reads and assembling the unmapped portion should be fine. I recommend with this strategy that if you have a pair in which only one read maps, to use both of for assembly, since presumably at least one was not assembled, and the other may have mapped to something that it did not actually come from. Of course, if that generates too much data to assemble, you can also just discard half-mapped pairs.

    Comment

    • confurious
      Junior Member
      • Apr 2017
      • 9

      #17
      Hi Brian, I am looking for a tool to merge multiple assemblies together but have not had any luck yet. I am wondering if you are aware of any tools like such? Minimus2 would not be ideal because it's merges only two sets, although I am not sure if I could simply divide my total concatenation file into two files to "fool" it. Thanks!

      Comment

      • Brian Bushnell
        Super Moderator
        • Jan 2014
        • 2709

        #18
        Sorry, I'm not aware of such a tool. We used to use Dedupe followed by Minimus2 to merge multiple assemblies. I don't know the details of how Minimus2 was used, but if it is restricted to two files at a time, you could proceed iteratively:

        assembly1 + assembly2 -> combined12
        combined12 + assembly3 -> combined123
        combined123 + assembly4 -> combined1234

        ...etc

        Comment

        • confurious
          Junior Member
          • Apr 2017
          • 9

          #19
          Hi Brian, sorry for asking for your help again. I am trying to use your bbcountunique.sh but I would like to change the k mer size to a number that is much bigger than 31, say something like 127 because I think 31 is not enough in my case.

          I examined the CalcUniqueness.java file and see the assert statement that makes it smaller than 31. I tried to change it and make it into a class via javac but for some reasons it reports a lot of errors within the java file that is probably due to windows to linux system conversion. However, I am not a guru of JAVA so I am stuck at this point. I could message you my email address if it's easier for you to quickly change that (assuming that's the only thing needed) and send me the updated CalcUniqueness.class file.

          Thanks so much!

          Comment

          • Brian Bushnell
            Super Moderator
            • Jan 2014
            • 2709

            #20
            Sorry, but the reason for the assertion is because the data structures fundamentally won't handle kmers longer than 31, since they are being stored in 64-bit integers. It's possible to modify the code to allow K>31 because I now have some classes that support unlimited-length kmers, but that would take quite a lot of work. You can always disable assertions by adding the flag -da when the program runs - "bbcountunique.sh -da <other arguments>". But it won't give correct results for K>31.

            Generally, I don't see a useful purpose for K>31 with that program - the longer K is, the more errors, which inflate the uniqueness; and K=31 should be sufficient for determining whether you have seen a read before. You won't saturate the K=31 kmer space for read uniqueness plots with Illumina HiSeq machines; that would take 600 million terabases at 2x150bp.
            Last edited by Brian Bushnell; 05-12-2017, 09:45 PM.

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              New Genomics Technologies Take Aim at Long-Standing Limits
              by SEQadmin2


              Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

              We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
              ...
              Yesterday, 10:25 AM
            • SEQadmin2
              How Immunogenomics Decodes Immunity’s Genetic Blueprint
              by SEQadmin2




              The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

              This convergence of genetics, immunology, and computation...
              09-01-2026, 05:41 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Today, 09:51 AM
            0 responses
            10 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-25-2026, 09:06 AM
            0 responses
            32 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-23-2026, 11:05 AM
            0 responses
            27 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-18-2026, 11:37 AM
            1 response
            48 views
            0 reactions
            Last Post pekgio
            by pekgio
             
            Working...