Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • yh_gu
    Junior Member
    • Apr 2010
    • 4

    contig assembly

    Hello,

    We've de novo assembled our RNA-Seq reads (about 50 millions 2×75 reads) into contigs by several de novo assemblers with different parameters. Most of the contigs we’ve get were very short due to the poor sequencing quality and the low sequencing depth. The contigs from each assmbler under different parameters varied from eachother, but some of them had overlaps. So these contigs may be assembled into longer contigs. The problem is that we couldn't assemble millions of contigs into supercontigs manually. Moreover, our computer resources were very low (12G RAM, 8 core CPU, 500G spaces). Is anyone knows how to assmeble these contigs into longer contigs with our limited computer resources, and which software could handle the assemble task.

    Thanks

    YH-GU
    Last edited by yh_gu; 08-07-2010, 01:57 AM.
  • Zigster
    Jeremy Leipzig
    • May 2009
    • 117

    #2
    If you have enough memory to assemble reads into contigs then you clearly have enough memory to assemble contigs into supercontigs, as that is an easier feat.

    When an assembler produces contigs that ostensibly overlap and yet remain separate there is some likely path ambiguity that has not been resolved. It sounds like you'll need more paired-end sequence for your assembly to coalesce.
    --
    Jeremy Leipzig
    Bioinformatics Programmer
    --
    My blog
    Twitter

    Comment

    • natstreet
      Member
      • Nov 2009
      • 83

      #3
      The last is certainly true - but can anyone recommend the most suitable software tool for the task? I am currently looking to do something similar and so far have only tried PAVE (which died with an error message that I'm tracking down). In my case I have performed de novo assembly on a number of genotypes of the same species and now I want to merge those together, identify SNPs and see if longer ESTs can be made by merging contigs across the per-genotype assemblies.

      So, what are the current favourite tools for merging large numbers of contigs coming from de novo transcript assemblies of short read data?

      Comment

      • yh_gu
        Junior Member
        • Apr 2010
        • 4

        #4
        Originally posted by Zigster View Post
        When an assembler produces contigs that ostensibly overlap and yet remain separate there is some likely path ambiguity that has not been resolved. It sounds like you'll need more paired-end sequence for your assembly to coalesce.
        The overlaps I've mentioned mainly refer to the contigs that produced by different assembler. So, we want to find a suitable software to assemble them longer.

        Comment

        • jmw86069
          Member
          • Jun 2009
          • 31

          #5
          Velvet seems to work fairly well with contig assemblies in my hands, though as Zigster pointed out, the assembly path ambiguity will ultimately prevent use of as much productive overlaps as you'd suspect because there will be discrepancies across assemblers in just the wrong places per contig.

          It may be interesting to "go conservative" with different assemblers' contigs by trimming away their weakpoints. E.g. maybe try trimming away low quality ends of contigs to minimize including ambiguous sequence spans in your secondary assembly. Otherwise you'd expect to get good overlaps in the middle of the contigs but not good alignments at the ends. But each assembler has its challenge area, so you may want to deal with each one in its own way. At some point we (the collective) should put together some cross-assembler lessons learned, and maybe pre-configurations that help tools like Velvet use each assembler's strengths more natively.

          Velvet does have numerous options to tweak though, which I think gives it promise, and you can try "oases" which is a layer on top of Velvet which is intended to allow for splice variants. Marcel Schulz and Daniel Zerbino seem to have put together a very useful (and timely) toolsuite for this type of work. Kudos to them, and thanks to them as well for providing it as they continue perfecting it.

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM
          • SEQadmin2
            Cancer Drug Resistance: The Lingering Barrier to Rising Survival
            by SEQadmin2



            Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

            There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
            07-08-2026, 05:17 AM
          • GATTACAT
            Reply to Nine Things a Sample Prep Scientist Thinks About Before Sequencing
            by GATTACAT
            Love this - good data definitely starts from good input, and poor input can only give relatively poor data. I particularly like the mention of Nanodrop/absorbance based methods for quantification. It's such a toss up if you'll get an accurate reading or what amounts to a randomly generated number, and a lot of library/sequencing related issues can be traced back to poor quant.
            07-01-2026, 11:43 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, 07-13-2026, 10:26 AM
          0 responses
          27 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-09-2026, 10:04 AM
          0 responses
          37 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-08-2026, 10:08 AM
          0 responses
          24 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-07-2026, 11:05 AM
          0 responses
          35 views
          0 reactions
          Last Post SEQadmin2  
          Working...