Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • days369
    Junior Member
    • Aug 2009
    • 8

    RNA-Seq Quantification and Differential Expression Analysis

    I am currently analyzing RNA-Seq data from bacterial transcriptome. Here I have several questions regarding gene expression quantification and differential expression analysis.

    1. For Partially Overlapped reads
    If I use RPKM for quantification, how should I count the reads that only partially overlaps with the annotated gene regions in the reference genomes. Should I count each read as 1 no matter how long they overlaps with the gene? or to multiply 1 by a weight corresponding to how much they overlaps? Or to discard those reads and only consider the reads that are completely within a gene annotation?

    2. Paired-End reads
    Since my data are paired-end reads, should I consider that the gap between two ends contributes to coverage? Some background is that my RNA-seq data is not from a pure culture of bacteria. It might contain several similar species of bacteria. Their genomes are pretty similar, but not identical.

    3. Differential Expression Analysis Methodology
    a. I‘ve seen some posts discussing about DE methods. T-tests were recommended when there are "many" biological replicates. I am wondering if 5 vs. 5 should be considered as "many" or "OK amount of" replicates?
    b. Say, it is OK to use T-test. Since it is not clear whether it is valid to assume the T statistic follow a t-distribution given the data, is it more appropriate to use permutation method to generate null distribution?

    Those questions have bothered me for a long time. I appreciate any type of helps!!

    Thanks in advance,
    Dezhi
  • days369
    Junior Member
    • Aug 2009
    • 8

    #2
    Can anybody kindly address some of the questions? My post has been ignored

    Comment

    • Simon Anders
      Senior Member
      • Feb 2010
      • 995

      #3
      Hi Dezhi

      to use something like a t test, you need enough replicates to estimate a variance for each gene. With two groups of five samples, you are already entering the regime there this should work well. For comparison, also try a tool that pools information from several genes to get better confidence in variance estimates, such as our DESeq or the Smyth group's edgeR. (Of course, we like to claim that DESeq is better than edgeR, and for only two or three replicates, I do think so, but for five or more replicates, edgeR's "moderation" feature really pays off. So, even though I don't like admitting this, for your set-up, edgeR should work better than DESeq.)

      I wonder whether ten samples are sufficient to get a good permutation null. If you try it, be sure to let us know about your experiences.

      About RPKM: The raw integer number of counts gives you a lot of information about the expected Poisson noise, and this is why I recommend not to use RPKM values in differential expression analysis. Two genes with the same RPKM value can have very different accuracies, if one is a long gene with few counts and the other a short gene with many counts. This is why DESeq and edgeR require unnormalized counts as input.

      Comment

      Latest Articles

      Collapse

      • SEQadmin2
        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
        by SEQadmin2


        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

        The systematic characterization of the human proteome has
        ...
        07-20-2026, 11:48 AM
      • SEQadmin2
        Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
        by SEQadmin2



        Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
        ...
        07-09-2026, 11:10 AM
      • SEQadmin2
        Cancer Drug Resistance: The Lingering Barrier to Rising Survival
        by SEQadmin2



        Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

        There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
        07-08-2026, 05:17 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by SEQadmin2, 07-24-2026, 12:17 PM
      0 responses
      20 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-23-2026, 11:41 AM
      0 responses
      19 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-20-2026, 11:10 AM
      0 responses
      26 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-13-2026, 10:26 AM
      0 responses
      38 views
      0 reactions
      Last Post SEQadmin2  
      Working...