Seqanswers Leaderboard Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • alabadorf
    Junior Member
    • Jun 2010
    • 3

    DESeq normalization and sample counts of zero

    First, apologies if this question is posted elsewhere, but I could not find an answer to my question.

    I am using the DESeq normalization procedure on a set of 100 mRNA-Seq samples aligned with STAR and counted with htseq-count. I noticed in the normalization procedure that only genes where all samples have non-zero counts are used in the calculation of size factors. This means that only 15k/52k of my detected genes are used for this calculation, while 21k/52k genes have 5 or fewer zero counts. If I modify the normalization routine to include these extra genes by omitting the zero counts from the geometric mean calculation, the size factors change, where some samples increase, others decrease, and others remain the same. The relative order of the samples by size factor also changes. Overall, the size factors get smaller when increasing numbers of zeros are included, which I think makes sense given that more lowly expressed genes are included but I'm not entirely sure. Analyzing the normalized counts while allowing genes with zeros to be included modestly changes downstream DE results, but I haven't seen any red flags yet so I won't post any examples.

    My question: besides numerical reasons, what is the rationale for omitting genes that have even a single zero count? My feeling is that omitting any samples with zero counts biases the normalization to use only genes with high expression and does not take into account lowly, but confidently, expressed genes. Since I have so many samples, I worry this strategy might too severely penalize genes with outlier counts of zero. Is there danger in modifying the routine to allow, say, at most 5/100 samples to have zeros?

    Thanks.
  • GenoMax
    Senior Member
    • Feb 2008
    • 7142

    #2
    Check this thread for a discussion of "0" counts: http://seqanswers.com/forums/showthread.php?t=48239

    Comment

    • alabadorf
      Junior Member
      • Jun 2010
      • 3

      #3
      Thanks for the reply. I understand that zero counts cannot be confidently distinguished as absent or artifact. For most of the genes that have 5 or fewer zero count samples, the remaining samples have significant (100-1000+) read depth, leading me to believe that these are genes that should be included in the normalization routine. As the number of samples with zero counts increases, the overall counts of the remaining samples decreases, which is what we would expect as we approach the sequencing depth detection limit. This is not the case for the genes with only a few zero counts.

      Comment

      Latest Articles

      Collapse

      • seqadmin
        Pathogen Surveillance with Advanced Genomic Tools
        by seqadmin




        The COVID-19 pandemic highlighted the need for proactive pathogen surveillance systems. As ongoing threats like avian influenza and newly emerging infections continue to pose risks, researchers are working to improve how quickly and accurately pathogens can be identified and tracked. In a recent SEQanswers webinar, two experts discussed how next-generation sequencing (NGS) and machine learning are shaping efforts to monitor viral variation and trace the origins of infectious...
        03-24-2025, 11:48 AM
      • seqadmin
        New Genomics Tools and Methods Shared at AGBT 2025
        by seqadmin


        This year’s Advances in Genome Biology and Technology (AGBT) General Meeting commemorated the 25th anniversary of the event at its original venue on Marco Island, Florida. While this year’s event didn’t include high-profile musical performances, the industry announcements and cutting-edge research still drew the attention of leading scientists.

        The Headliner
        The biggest announcement was Roche stepping back into the sequencing platform market. In the years since...
        03-03-2025, 01:39 PM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by seqadmin, 03-20-2025, 05:03 AM
      0 responses
      49 views
      0 reactions
      Last Post seqadmin  
      Started by seqadmin, 03-19-2025, 07:27 AM
      0 responses
      57 views
      0 reactions
      Last Post seqadmin  
      Started by seqadmin, 03-18-2025, 12:50 PM
      0 responses
      50 views
      0 reactions
      Last Post seqadmin  
      Started by seqadmin, 03-03-2025, 01:15 PM
      0 responses
      201 views
      0 reactions
      Last Post seqadmin  
      Working...