Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • Combined assembly analysis (short reads + long reads)

    Dear NGS Experts,

    I have a question about combined genome assembly.

    We have 75X Hiseq sequencing of an animal species genome (about 3Gb genome size) together with 50X Pacbio Sequel system, now, we would like to make a combined assembly analysis of these 350Gb data. Anybody knows any tools for this kind of analysis?

    Many thanks.

  • #2
    For reference, cross-posted: https://www.biostars.org/p/187524/

    Comment


    • #3
      Originally posted by GenoMax View Post
      For reference, cross-posted: https://www.biostars.org/p/187524/
      Hi GenoMax, thanks a lot. Exactly, if there is any tool which can use the short reads to correct the error on the long reads, that would be great. Such tools should be very useful for combined assembly analysis.

      Comment


      • #4
        Did you see this link: https://github.com/PacificBioscience...Bio-Long-Reads pacBioToCA is the tool you want.

        Comment


        • #5
          Originally posted by GenoMax View Post
          Did you see this link: https://github.com/PacificBioscience...Bio-Long-Reads pacBioToCA is the tool you want.
          Dear GenoMax, thank you very much, that link is quite useful. There is another problem/situation that the genome is with very high heterozygosity, whether the five hybrid assemblers (pacBioToCA, ECTools, SPAdes, Cerulean, dbg2olc) can deal with such situation?

          Comment


          • #6
            You are not going to know until you try. Sounds to me like someone is going to stay busy for a while. Hope you have access to some beefy compute resources since this is going to take a lot of RAM etc.

            Comment


            • #7
              The term you're looking for is "hybrid assembly".

              Comment


              • #8
                Originally posted by GenoMax View Post
                You are not going to know until you try. Sounds to me like someone is going to stay busy for a while. Hope you have access to some beefy compute resources since this is going to take a lot of RAM etc.
                Right, we will try them all. The enough RAM is always important until the transfer rate of sad can reach at least 5GB/s. Thank you again.

                Comment


                • #9
                  I would suggest engaging with PacBio tech support early for this challenging project. Dr. Hall (rhall) from PacBio participates on this forum and he may be a good resource for hints.

                  If you are getting PacBio data from a sequence provider then be sure to ask for raw data files (*.h5). These would be needed for some of the tools we have discussed.

                  Comment


                  • #10
                    There was a thread floating around here somewhere suggesting that you use PacBio reads for assembly and Illumina reads to do error correction, etc. afterwards.

                    Comment

                    Latest Articles

                    Collapse

                    • seqadmin
                      Genetic Variation in Immunogenetics and Antibody Diversity
                      by seqadmin



                      The field of immunogenetics explores how genetic variations influence immune responses and susceptibility to disease. In a recent SEQanswers webinar, Oscar Rodriguez, Ph.D., Postdoctoral Researcher at the University of Louisville, and Ruben Martínez Barricarte, Ph.D., Assistant Professor of Medicine at Vanderbilt University, shared recent advancements in immunogenetics. This article discusses their research on genetic variation in antibody loci, antibody production processes,...
                      11-06-2024, 07:24 PM
                    • seqadmin
                      Choosing Between NGS and qPCR
                      by seqadmin



                      Next-generation sequencing (NGS) and quantitative polymerase chain reaction (qPCR) are essential techniques for investigating the genome, transcriptome, and epigenome. In many cases, choosing the appropriate technique is straightforward, but in others, it can be more challenging to determine the most effective option. A simple distinction is that smaller, more focused projects are typically better suited for qPCR, while larger, more complex datasets benefit from NGS. However,...
                      10-18-2024, 07:11 AM

                    ad_right_rmr

                    Collapse

                    News

                    Collapse

                    Topics Statistics Last Post
                    Started by seqadmin, Today, 11:09 AM
                    0 responses
                    24 views
                    0 likes
                    Last Post seqadmin  
                    Started by seqadmin, Today, 06:13 AM
                    0 responses
                    20 views
                    0 likes
                    Last Post seqadmin  
                    Started by seqadmin, 11-01-2024, 06:09 AM
                    0 responses
                    30 views
                    0 likes
                    Last Post seqadmin  
                    Started by seqadmin, 10-30-2024, 05:31 AM
                    0 responses
                    21 views
                    0 likes
                    Last Post seqadmin  
                    Working...
                    X