Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • cjohnson
    Junior Member
    • Jul 2010
    • 2

    Optimal read lenght for RNASeq

    Has anyone here done any research on optimal read length for doing RNASeq?

    I've been able to find one paper so far that looked at read length (Li et al). Suggesting that shorted reads <35 are best.

    I'm sure like everything it comes down to your experimental question, but I'm looking for some general advice.

    Thanks,
    Last edited by cjohnson; 11-03-2010, 07:59 AM.
  • jdanderson
    Member
    • Sep 2010
    • 45

    #2
    Hello cjohnson,

    I am far from an expert on this matter, but to get the ball rolling on your thread:

    In general, the longer the better and paired ends (especially) are helpful. I think the added value is greater resolution (and statistical power) of smaller genes and as well as lowly expressed genes. I'm sure cost is a concern for you, as it is for everyone (if not, then i envy you), so you may find that adding a little read length is a bit cheaper than running PE's. If you are using a core facility, however, you have to consider how much of the slide you are using (# of lanes) versus how popular a particular read length is with your type of run, PE or SE. I've heard that 40bp SE reads are the minimum you should try to do, and of course the more biological replicates the more powerful your analysis will be.

    Hope this was some modicum of help.

    Cheers,
    Johnathon
    Last edited by jdanderson; 11-03-2010, 04:37 PM.

    Comment

    • Michael.James.Clark
      Senior Member
      • Apr 2009
      • 207

      #3
      I'm interested in this question as well. Haven't done RNAseq myself for a long time, and back then it was 36bp PE. Now I've read 75bp SE is pretty standard and am considering 100bp PE. Anyone have a solid answer?
      Mendelian Disorder: A blogshare of random useful information for general public consumption. [Blog]
      Breakway: A Program to Identify Structural Variations in Genomic Data [Website] [Forum Post]
      Projects: U87MG whole genome sequence [Website] [Paper]

      Comment

      • mnkyboy
        Member
        • Mar 2009
        • 87

        #4
        If it is for mRNA-seq profiling then you will do well with just 54 bp. If it is for whole transcriptome, junctions, and structure longer is better. PE obviously helps there. But you may not get PE support for directional libraries.

        We like 76 directional for whole transcriptome but that is because at 100 we start seeing that we will sequence the adapters sometimes.

        Comment

        • cndewey
          Junior Member
          • Nov 2008
          • 2

          #5
          I'm an author on the Li et al. paper that was cited, so I thought I would add my thoughts.

          If your only goal is to do quantitation and/or differential expression, we've found that the number of reads matters more than read length once you reach a minimum read length of about 25. So if you have a choice between one lane of 50bp SE vs. two lanes of 25bp SE, you should go with the two lane option. Similarly, two lanes of 25bp SE is better than one 25bp PE run. Of course, if you only have one lane to work with, having longer reads won't hurt, unless the number of valid reads goes down.

          Another group has confirmed our results on this:
          Nicolae et al.

          If you want to do reference-based or de novo transcriptome assembly, then longer reads or PE reads will probably be more useful.

          Comment

          • malachig
            Senior Member
            • Aug 2010
            • 117

            #6
            I agree with others here that the greater the length the better and paired is more desirable than single end. For paired-ends I wouldn't go shorter than 40 and for single end, no shorter than 50. I also agree that depending on application, depth may be more important than read length.

            An obvious consideration related to read length is mapability of your reads. If you want to have some numbers to justify your choice you can download the 'mapability' tracks from UCSC and consider how many bases of the transcriptome (i.e. exonic bases) are uniquely mapable at various read lengths. For hg19, this has been pre-calculated for 24-, 36-, 40-, 50-, 75- and 100-mer sequences with up to two mismatches allowed.

            Another suggestion. There are a number of caveats associated with comparing libraries of different length. So whatever you decide for length and paired vs. un-paired, stick with that for all libraries if you can...

            Comment

            • honey
              Senior Member
              • Feb 2010
              • 151

              #7
              I was wondering what parameters should be optimal to get the out put in TopHAT.
              Thanks

              Comment

              Latest Articles

              Collapse

              • SEQadmin2
                Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                by SEQadmin2


                Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                The systematic characterization of the human proteome has
                ...
                07-20-2026, 11:48 AM
              • SEQadmin2
                Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                by SEQadmin2



                Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                ...
                07-09-2026, 11:10 AM
              • SEQadmin2
                Cancer Drug Resistance: The Lingering Barrier to Rising Survival
                by SEQadmin2



                Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

                There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
                07-08-2026, 05:17 AM

              ad_right_rmr

              Collapse

              News

              Collapse

              Topics Statistics Last Post
              Started by SEQadmin2, Today, 12:17 PM
              0 responses
              10 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, Yesterday, 11:41 AM
              0 responses
              11 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-20-2026, 11:10 AM
              0 responses
              23 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-13-2026, 10:26 AM
              0 responses
              37 views
              0 reactions
              Last Post SEQadmin2  
              Working...