Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • anyone1985
    Member
    • Mar 2009
    • 68

    #1

    How to determine the insert length?

    I have some Solexa pair-end data. But my colleague forgot to tell me the insert length How can I determine the insert length of the data? First, I have no reference genome.
  • westerman
    Rick Westerman
    • Jun 2008
    • 1104

    #2
    The easiest way would be to go back to your colleague and get the insert length. But if you really insist on doing this the hard way then I would suggest doing a de-novo assembly using the reads as if they were fragments (i.e., not as paired ends). This should give you some decent size contigs. Then map the paired ends map onto the contigs. From this you should be able to figure out how far apart are the paired ends that do map. After you obtain the numbers that define the range you can then do an new assembly but this time as a 'paired end' instead of a 'fragment' assembly.

    Comment

    • anyone1985
      Member
      • Mar 2009
      • 68

      #3
      thank you, i think i know what i should do

      Comment

      • system7
        Junior Member
        • Apr 2008
        • 5

        #4
        Originally posted by anyone1985 View Post
        I have some Solexa pair-end data. But my colleague forgot to tell me the insert length How can I determine the insert length of the data? First, I have no reference genome.
        The summary.htm file from Pipeline should have that info in it at the bottom of the file

        Comment

        • Torst
          Senior Member
          • Apr 2008
          • 275

          #5
          Originally posted by westerman View Post
          The easiest way would be to go back to your colleague and get the insert length.
          The problem with this is that the DNA fragment selection step is inexact. You may be aiming for 250 bp, but the average is 220 say, with a standard deviation of 30.

          But if you really insist on doing this the hard way then I would suggest doing a de-novo assembly using the reads as if they were fragments (i.e., not as paired ends). This should give you some decent size contigs. Then map the paired ends map onto the contigs. From this you should be able to figure out how far apart are the paired ends that do map. After you obtain the numbers that define the range you can then do an new assembly but this time as a 'paired end' instead of a 'fragment' assembly.
          This is good advice. If you have a close reference sequence, you can use that instead of de novo contigs. I usually use MAQ to align a SUBSET of the reads in paired-end mode, and MAQ itself will print out the mean and s.d. of the insert size.

          And as another poster said, if this is Illumina GA Pipeline, the Summary HTML files contain an estimate of the insert size which it obtains by using ELAND to map the reads to the reference genome specified in the gerald.cfg file.

          Comment

          • westerman
            Rick Westerman
            • Jun 2008
            • 1104

            #6
            Originally posted by Torst View Post
            The problem with this is that the DNA fragment selection step is inexact. You may be aiming for 250 bp, but the average is 220 say, with a standard deviation of 30.
            Well yes this could be a problem if your colleague only gives a single number then you have problems. I always ask for a minimum and maximum insert length knowing that those numbers are also uncertain. Also sometimes you can have a mixture of libraries with different insert sizes; e.g. average of 500 bp; 3K, 20K. Then one needs to know not only the range but also which reads corresponds to which library.

            And as another poster said, if this is Illumina GA Pipeline, the Summary HTML files contain an estimate of the insert size which it obtains by using ELAND to map the reads to the reference genome specified in the gerald.cfg file.
            Ah, but the original poster said he did not have a reference genome.

            It was an interesting theoretical question -- how does one figure out insert sizes when only given paired ends. A question that I am glad that I do not have to do in practice!

            Comment

            • ashrafi_h
              Junior Member
              • Jan 2010
              • 7

              #7
              How do you use maq to determine the insert size?

              Originally posted by Torst View Post
              The problem with this is that the DNA fragment selection step is inexact. You may be aiming for 250 bp, but the average is 220 say, with a standard deviation of 30.



              This is good advice. If you have a close reference sequence, you can use that instead of de novo contigs. I usually use MAQ to align a SUBSET of the reads in paired-end mode, and MAQ itself will print out the mean and s.d. of the insert size.

              And as another poster said, if this is Illumina GA Pipeline, the Summary HTML files contain an estimate of the insert size which it obtains by using ELAND to map the reads to the reference genome specified in the gerald.cfg file.
              Hi, If you do not have ref sequence, how do you use maq to determine the insert size. Could you please have a sample command line.

              Thanks

              Comment

              • ashrafi_h
                Junior Member
                • Jan 2010
                • 7

                #8
                Originally posted by system7 View Post
                The summary.htm file from Pipeline should have that info in it at the bottom of the file
                Where is this summary.htm that people are talking about?

                Comment

                Latest Articles

                Collapse

                • SEQadmin2
                  Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                  by SEQadmin2



                  CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                  Despite this, “CRISPR helped turn genome editing from a specialized technique into
                  ...
                  07-31-2026, 11:01 AM
                • SEQadmin2
                  Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                  by SEQadmin2


                  Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                  The systematic characterization of the human proteome has
                  ...
                  07-20-2026, 11:48 AM
                • SEQadmin2
                  Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                  by SEQadmin2



                  Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                  ...
                  07-09-2026, 11:10 AM

                ad_right_rmr

                Collapse

                News

                Collapse

                Topics Statistics Last Post
                Started by SEQadmin2, 07-31-2026, 02:55 AM
                0 responses
                14 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-24-2026, 12:17 PM
                0 responses
                15 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-23-2026, 11:41 AM
                0 responses
                13 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-20-2026, 11:10 AM
                0 responses
                24 views
                0 reactions
                Last Post SEQadmin2  
                Working...