Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • DESeq normalization and sample counts of zero

    First, apologies if this question is posted elsewhere, but I could not find an answer to my question.

    I am using the DESeq normalization procedure on a set of 100 mRNA-Seq samples aligned with STAR and counted with htseq-count. I noticed in the normalization procedure that only genes where all samples have non-zero counts are used in the calculation of size factors. This means that only 15k/52k of my detected genes are used for this calculation, while 21k/52k genes have 5 or fewer zero counts. If I modify the normalization routine to include these extra genes by omitting the zero counts from the geometric mean calculation, the size factors change, where some samples increase, others decrease, and others remain the same. The relative order of the samples by size factor also changes. Overall, the size factors get smaller when increasing numbers of zeros are included, which I think makes sense given that more lowly expressed genes are included but I'm not entirely sure. Analyzing the normalized counts while allowing genes with zeros to be included modestly changes downstream DE results, but I haven't seen any red flags yet so I won't post any examples.

    My question: besides numerical reasons, what is the rationale for omitting genes that have even a single zero count? My feeling is that omitting any samples with zero counts biases the normalization to use only genes with high expression and does not take into account lowly, but confidently, expressed genes. Since I have so many samples, I worry this strategy might too severely penalize genes with outlier counts of zero. Is there danger in modifying the routine to allow, say, at most 5/100 samples to have zeros?

    Thanks.

  • #2
    Check this thread for a discussion of "0" counts: http://seqanswers.com/forums/showthread.php?t=48239

    Comment


    • #3
      Thanks for the reply. I understand that zero counts cannot be confidently distinguished as absent or artifact. For most of the genes that have 5 or fewer zero count samples, the remaining samples have significant (100-1000+) read depth, leading me to believe that these are genes that should be included in the normalization routine. As the number of samples with zero counts increases, the overall counts of the remaining samples decreases, which is what we would expect as we approach the sequencing depth detection limit. This is not the case for the genes with only a few zero counts.

      Comment

      Latest Articles

      Collapse

      • seqadmin
        Non-Coding RNA Research and Technologies
        by seqadmin




        Non-coding RNAs (ncRNAs) do not code for proteins but play important roles in numerous cellular processes including gene silencing, developmental pathways, and more. There are numerous types including microRNA (miRNA), long ncRNA (lncRNA), circular RNA (circRNA), and more. In this article, we discuss innovative ncRNA research and explore recent technological advancements that improve the study of ncRNAs.

        Nobel Prize for MicroRNA Discovery
        This week,...
        10-07-2024, 08:07 AM
      • seqadmin
        Recent Developments in Metagenomics
        by seqadmin





        Metagenomics has improved the way researchers study microorganisms across diverse environments. Historically, studying microorganisms relied on culturing them in the lab, a method that limits the investigation of many species since most are unculturable1. Metagenomics overcomes these issues by allowing the study of microorganisms regardless of their ability to be cultured or the environments they inhabit. Over time, the field has evolved, especially with the advent...
        09-23-2024, 06:35 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by seqadmin, 10-02-2024, 04:51 AM
      0 responses
      103 views
      0 likes
      Last Post seqadmin  
      Started by seqadmin, 10-01-2024, 07:10 AM
      0 responses
      112 views
      0 likes
      Last Post seqadmin  
      Started by seqadmin, 09-30-2024, 08:33 AM
      1 response
      114 views
      0 likes
      Last Post EmiTom
      by EmiTom
       
      Started by seqadmin, 09-26-2024, 12:57 PM
      0 responses
      21 views
      0 likes
      Last Post seqadmin  
      Working...
      X