Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • jgibbons1
    Senior Member
    • Oct 2009
    • 135

    #1

    assembly strategies for repetitive dna

    Hi all,

    I was hoping to get some feedback on an assembly strategy of repetitive dna using Illumina 100-PE data.

    I've seen some strategies, such as ALL-PATHS-LG, which utilize multiple libraries of increasingly larger insert size to resolve repetitive regions. For example, assembly of the potato genome (which is ~60% repetitive) used 7 libraries ranging from an insert size of 200 bp - 20,000 bp.

    We're running a pilot on a few samples to see how well the assembly will be, but because of coverage issues, will only be creating 2 libraries per sample.

    So here is my question:

    Is it better to create the libraries with insert sizes that are close in range (ex. 200 bp and 500 bp) or large in range (ex. 200 bp and 5,000 bp)? I can see pros and cons of both, but wanted elicit some advice before going forward.

    Thanks,

    John
  • jgibbons1
    Senior Member
    • Oct 2009
    • 135

    #2
    One more thing while I'm at it...

    Has anyone used Telescoper (DOI:10.1093/bioinformatics/bts399)? If so, I'd be interested in hearing about your experiences with it.

    Comment

    • bckirkup
      Member
      • Jan 2011
      • 17

      #3
      Assembly of repetitive DNA

      The big question: what is repeating/how much is repeating? Sounds vaguely leninist.

      One person who has experience with this is Matt Riley at U. Tennessee, at least in microbial genomes.

      Anyway, if you have tandem repeats of a few bp, then read length is your big factor; if you have repeats of a gene, then you need jumps/paired ends. If you have repeats of gene clusters, you may need something more substantial. 40kb jumps are possible and published; PacBio reads are another option; an Optical Map may be the answer. Of course, you'll need to 'fill in' the map or fix the SNPs in the SMRT reads. Joint assemblies are performed by a number of groups; the folks at NCBI, the FDA, and UMD (Mihai Pop) are familiar with the strategies.

      Ultimately, if you have the worst case scenario, some sort of scale-free nesting of repeats within repeats, you would need all these solutions combined.

      Hope that points you toward some ideas.

      Comment

      • jgibbons1
        Senior Member
        • Oct 2009
        • 135

        #4
        Originally posted by bckirkup View Post
        The big question: what is repeating/how much is repeating? Sounds vaguely leninist.
        Leninist indeed

        Thanks for your response...They are VERY helpful.

        I'm interested in one chromosome which is probably a worst case scenario - variable sized microsatellites, minisatellites, transposable elements, and variable sized rDNA arrays. Quite frankly, it's a mess.

        I'm not looking for a complete assembly, but the chromosome is ~40 Mb and I would like to generate scaffolds large enough to give me something to work with. Using a published illumina dataset (86 bp pe), I wasn't able to assembly anything larger than 1kb, although in non repetitive regions I was getting scaffolds as large as 400 Kb.

        My goal is to predict functional motifs from the assembly (ex. TF binding sites, transposable element content, CNV in satellite sequences) and identify variation in this chromosome across populations.

        Comment

        • HESmith
          Senior Member
          • Oct 2009
          • 512

          #5
          Ugh. (sorry, I meant to say "What a challenging project!")

          Regarding your initial question (close vs. large size differences in the two libraries), the larger difference will be more useful for assembly. The ideal is to have paired end read jumps that span the repeats, which delimits the number of copies in the intervening region. You can calculate transposable element content and satellite CNV by read depth (although their positions will be difficult/impossible to assign).

          Note that coverage issues do not necessarily limit you to two library sizes. For assembly, it would be more useful to have additional jump libraries sequenced at lower depth. IIRC, ALLPATHS-LG sequenced large (10kbp) jump libraries at 1/25th the depth of the shorter-sized inserts. A similar hybrid approach may also be your best bet.

          Good luck!

          Comment

          • jgibbons1
            Senior Member
            • Oct 2009
            • 135

            #6
            Originally posted by HESmith View Post
            Ugh. (sorry, I meant to say "What a challenging project!")
            haha...indeed!

            Thanks for your comments. This is making me thing starting with 2 libraries for the pilot (probably 250 bp and 1 kb) at a higher depth and then running another 2 libraries (5 kb and 10 kb) at a lower depth would be a good strategy.

            Comment

            • piyo2
              Junior Member
              • Oct 2010
              • 4

              #7
              Jumping libraries are a possibility, but the cost and difficulty to make the libraries are a consideration and there may be value in sequencing through the entire region. The current PacBio C2 XL chemistry averages 5kb read length with 10% reads 10kb+, max is usually ~20kb.

              If you would like more information about potentially doing a PacBio library, we can discuss I can help make some introductions to labs in the area to get your challenging region sequenced. You can email me: akieu at pacificbiosciences.com

              Comment

              Latest Articles

              Collapse

              • SEQadmin2
                Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                by SEQadmin2



                CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                Despite this, “CRISPR helped turn genome editing from a specialized technique into
                ...
                07-31-2026, 11:01 AM
              • SEQadmin2
                Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                by SEQadmin2


                Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                The systematic characterization of the human proteome has
                ...
                07-20-2026, 11:48 AM
              • SEQadmin2
                Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                by SEQadmin2



                Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                ...
                07-09-2026, 11:10 AM

              ad_right_rmr

              Collapse

              News

              Collapse

              Topics Statistics Last Post
              Started by SEQadmin2, 07-31-2026, 02:55 AM
              0 responses
              18 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-24-2026, 12:17 PM
              0 responses
              16 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-23-2026, 11:41 AM
              0 responses
              16 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-20-2026, 11:10 AM
              0 responses
              26 views
              0 reactions
              Last Post SEQadmin2  
              Working...