Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • mabentley86
    Junior Member
    • May 2011
    • 4

    Initial Assembly Help - Short Reads Large Genome

    Hello everyone,

    I'm thinking about trying my hand at assembling some ~ 300 MBp genomes using 100 bp paired-end Illumina reads (coverage is not huge ~ 8-12X). I'm not new to bioinformatics, but new to this area generally. Does anyone have any preferences for which assemblers work best for this kind of data? I'm not sure if I should attempt to use a ref genome for a guide since the most closely related species is not that closely related.

    Any pointers much appreciated.

    Thanks,

    Michael
  • jimmybee
    Senior Member
    • Sep 2010
    • 119

    #2
    How repetitive is it? How close is the reference genome? Is it highly polymorphic? Diploid/tetraploid/hexaploid?

    Comment

    • mabentley86
      Junior Member
      • May 2011
      • 4

      #3
      Hi Jimmy,

      Thanks for replying. I'm working with Lepitdopteran genomes (butterflies and moths), which for current ref genomes are reported to have fairly high repetitive content. The closest ref genomes are Manduca, Heliconius, Bombyx and Monarch. When I blastn conserved genes against these, the best similarity hits are ~90%, the worst are ~65%. The genomes are diploid.

      My current strategy is to first filter raw paired-end reads for low quality/adapters and then go straight into assembly with SOAPdenovo (47 kmer). I'm running this on a cluster, so will hopefully have some indication of its success in the next day or two. Any further comments welcomed.


      Best,

      Michael

      Comment

      • darked89
        Member
        • Jun 2009
        • 38

        #4
        Few tips:

        1) error correction prior to assembly
        I used Coral (5 iterations) but you may check Quake or Reptile

        2) depending on insert sizes of your libs, you may try to find overlaps between paired ends prior to assembly with FLASH


        3) try more assemblers (Abyss, SGA, etc.), compare the results

        If you feel like experimenting:

        Comment

        • jimmybee
          Senior Member
          • Sep 2010
          • 119

          #5
          Originally posted by mabentley86 View Post
          Hi Jimmy,

          Thanks for replying. I'm working with Lepitdopteran genomes (butterflies and moths), which for current ref genomes are reported to have fairly high repetitive content. The closest ref genomes are Manduca, Heliconius, Bombyx and Monarch. When I blastn conserved genes against these, the best similarity hits are ~90%, the worst are ~65%. The genomes are diploid.

          My current strategy is to first filter raw paired-end reads for low quality/adapters and then go straight into assembly with SOAPdenovo (47 kmer). I'm running this on a cluster, so will hopefully have some indication of its success in the next day or two. Any further comments welcomed.


          Best,

          Michael
          Looks ok. I second darked89's points regarding trying different assemblers. Setup some bash scripts to experiment with different parameters too. I can't comment on error correction too much, but considering your low coverage it wouldnt be a bad option

          Have a look at the assemblathon papers to pick up some good assembly tips too

          Comment

          • Tong.W
            Member
            • Jul 2012
            • 14

            #6
            You had better have more data ,not just more coverage,but also other sequence library data,just like 2000bp insert size or bigger one.A reference may be not a guide for your assembly if they are not so closed,sometimes it may introduce many errors.

            Comment

            • Wallysb01
              Senior Member
              • Feb 2011
              • 286

              #7
              The tips given here are great, but with 8-12x coverage, you're not going to have a usable assembly. With sequencing errors, heterozygousity, and uneven coverage, you might only assemble 1/4 of the genome into contigs >200bp. Maybe you'd be able to get away with it if the genome had few repeats, but it sounds like that's not the case.

              What you may find is that you can map single exons of genes to your contigs, but that will make orthology assignment difficult in many cases. So with that kind of coverage, I'd stick to alignment of your reads to the most closely related species, and try to allow for some extra sequence divergence. From my experience 30x coverage is about the minimum for de novo assembly from NGS data. You could probably get way with traditional Illumina libraries if you did 300, 600 and 1000bp inserts, and avoid the mate pair libraries, at least initially. Though if you want as good of an assembly as possible, you'll need some 2-10kbp mate pair libraries, as well.

              Comment

              • mabentley86
                Junior Member
                • May 2011
                • 4

                #8
                Thanks for all the tips everyone, all very helpful.

                The aim of gathering this kind of data was to build low quality assemblies that could be queried for coding sequences, without worrying so much about synteny or other genome features. This is why different insert sizes weren't used. I guess the aim was to build the cheapest genome that could still prove to be useful.

                As Wallysb mentioned, the problem I'm finding is that it is difficult to assign orthology to hits. This is not such a problem for one of my genes of interest, but several of the others appear to have lineage specific duplicate copies. Here the trouble is first identifying true exon hits, and then piecing them together correctly. The process is also very manual, writing scripts wouldn't really help for this part of the process as I need to eyeball everything and tweak parameters just to be sure I'm getting the right things out.

                I am going to try some of these suggestions and see if things improve. As ever, any further comments or suggestions much appreciated.

                Michael

                Comment

                Latest Articles

                Collapse

                • SEQadmin2
                  Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                  by SEQadmin2


                  Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                  The systematic characterization of the human proteome has
                  ...
                  07-20-2026, 11:48 AM
                • SEQadmin2
                  Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                  by SEQadmin2



                  Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                  ...
                  07-09-2026, 11:10 AM
                • SEQadmin2
                  Cancer Drug Resistance: The Lingering Barrier to Rising Survival
                  by SEQadmin2



                  Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

                  There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
                  07-08-2026, 05:17 AM

                ad_right_rmr

                Collapse

                News

                Collapse

                Topics Statistics Last Post
                Started by SEQadmin2, 07-24-2026, 12:17 PM
                0 responses
                31 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-23-2026, 11:41 AM
                0 responses
                23 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-20-2026, 11:10 AM
                0 responses
                214 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-13-2026, 10:26 AM
                0 responses
                79 views
                0 reactions
                Last Post SEQadmin2  
                Working...