Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • simon_seq
    Member
    • Aug 2012
    • 13

    #1

    Using alternative normalization method in DESeq analysis of gene enrichment

    I'm using DESeq to identify differentially expressed genes in a next-gen sequencing dataset. (DESeq: http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3218662/). In my experiment, the normalization of read counts as implemented by DESeq may not perform as well as anticipated. DESeq uses a 'size factor' to achieve a common scale of count values across samples. The size factor for a given library is defined as the median of the ratios of observed counts to the geometric mean of each corresponding target over all samples.

    For my dataset, upper-quartile scaling (using the 75th percentile of data, which often has low read counts, for linear scaling) may improve performance. In my specific case it may even be possible to define a set of genes as an internal scaling standard.

    In order for the statistical test for differential gene expression to work correctly, am I allowed to use an alternative method for data normalization?
  • dpryan
    Devon Ryan
    • Jul 2011
    • 3478

    #2
    Sure, see the CQN package for an example of a different normalization method that can be used with DESeq2, which you should be using rather than DESeq.

    Comment

    • simon_seq
      Member
      • Aug 2012
      • 13

      #3
      Thanks a lot for your quick answer!

      May I ask, what are the formal constraints when changing normalization methods, such that the basic concept of DESeq(2) still works? Ie. assumption of a negative binomial distribution, Fisher's exact testing? As long as I transform the data linearly and get count values out, I'm fine to transform the data?

      Comment

      • simon_seq
        Member
        • Aug 2012
        • 13

        #4
        And to extend my question, would I be fine to apply a LOESS transformation? It appears to me that in some of my conditions, there are very active targets, which then cause a bias in reads. Thus, a LOESS transformation may be relevant. Here are some plots of my data, MA plots given on the right side of the diagonal.

        (Here's the link in case the picture doesn't show: https://www.dropbox.com/s/qssqov0hzr...2_v01.png?dl=0 )
        Last edited by simon_seq; 08-27-2015, 11:00 AM. Reason: Link doesn't work.

        Comment

        • dpryan
          Devon Ryan
          • Jul 2011
          • 3478

          #5
          The trick is to not transform the data at all. Leave the data untouched and supply offsets and such to produce the needed normalization. Again, see how the CQN package and DESeq2 interact.

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
            by SEQadmin2



            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

            Despite this, “CRISPR helped turn genome editing from a specialized technique into
            ...
            07-31-2026, 11:01 AM
          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, 08-06-2026, 07:41 AM
          0 responses
          12 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 08-03-2026, 10:13 AM
          0 responses
          30 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-31-2026, 02:55 AM
          0 responses
          40 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-24-2026, 12:17 PM
          0 responses
          26 views
          0 reactions
          Last Post SEQadmin2  
          Working...