Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • billstevens
    Senior Member
    • Mar 2012
    • 120

    DESeq with SPIA

    Hello all,

    I've been using SPIA (Signal Pathway Impact analysis), and I think its great. I'm surprised I don't see it up on the forum more often, and I thought I would start this thread, one to discuss it and how to use it, and two, how DESeq (which I'm guessing most of us would be using if we were to use SPIA) works with it.

    Let me start off about DESeq. So when I write my res file that I generate from DESeq, I get entries that are Inf or #Name when one gene doesn't have any transcripts. SPIA doesn't know how to handle these and it crashes. So far, I've just been manually changing it (I guess I could use R) to numbers on the higher end. Does anyone else have a more scientific solution?

    Also my annotation gtf is of gene names for human, hg19. SPIA only uses ENTREZ gene IDs. So I convert them using Clone|Gene ID converter (this gets maybe 70% of them). Does anyone have a gtf file that uses just ENTREZ gene IDs or knows where I can get one?
  • dietmar13
    Senior Member
    • Mar 2010
    • 107

    #2
    CAVE: selection bias in RNA-seq

    SPIA and all other gene over-representation analysis methods suffer from the gene length and gene expression height bias of RNA-seq data.

    this means that longer genes and even more higher expressed genes get easier called as statistically significant by nearly all statistical RNA-seq methods. this impairs all methods which use p- or q-value cut-offs.

    see:
    Gene ontology analysis for RNA-seq: accounting for selection bias
    Matthew D. Young, Matthew J. Wakefield, Gordon K. Smyth, Alicia Oshlack
    Genome Biology 2010, 11:R14 (4 February 2010)

    better are gene set enrichment analysis methods, as they don't use p- or q-value cut-offs.

    i also used SPIA with RNA-seq data and set inf values to high numbers (as you) and averaged the expression values of all genes mapped to one ENTREZ ID.

    Comment

    • turnersd
      Senior Member
      • May 2011
      • 115

      #3
      I and my collaborators also find SPIA very useful. Maybe this is naive, but doesn't FPKM fix this issue? If you don't like what cufflinks is doing, can't you use something like eXpress to get FPKMs instead of counts (a la HTSeq-count)

      Comment

      • Wolfgang Huber
        Senior Member
        • Aug 2009
        • 109

        #4
        Originally posted by dietmar13 View Post
        SPIA and all other gene over-representation analysis methods suffer from the gene length and gene expression height bias ... this impairs all methods which use p- or q-value cut-offs....better are gene set enrichment analysis methods, as they don't use p- or q-value cut-offs.
        The fact that over-representation or enrichment analysis methods can be confounded by detection power is not new with RNA-Seq. This has been true for these methods all along. The key point is to use a 'background' set against which you compare that is controlled for that. So don't just use all genes that someone happened find somewhere in a table as background for these types of analyses, but use a matched set of genes that in your experiment would have had roughly equal chance of detection as those that actually did make it to the top of your list.

        Over-representation (i.e. using a fixed cutoff and hypergeometric test or alike) and enrichment analysis (i.e. looking for a trend in the test statistic that is associated with annotation) have fundamentally the same issue. Again, the key issue is choosing the right background set.

        The focus in some of these discussions on gene length is peculiar. Really the total number of counts is the dominant variable here, on which detection power depends.

        Best wishes
        Wolfgang
        Wolfgang Huber
        EMBL

        Comment

        • chadn737
          Senior Member
          • Jan 2009
          • 392

          #5
          Originally posted by Wolfgang Huber View Post
          The fact that over-representation or enrichment analysis methods can be confounded by detection power is not new with RNA-Seq. This has been true for these methods all along. The key point is to use a 'background' set against which you compare that is controlled for that. So don't just use all genes that someone happened find somewhere in a table as background for these types of analyses, but use a matched set of genes that in your experiment would have had roughly equal chance of detection as those that actually did make it to the top of your list.

          Over-representation (i.e. using a fixed cutoff and hypergeometric test or alike) and enrichment analysis (i.e. looking for a trend in the test statistic that is associated with annotation) have fundamentally the same issue. Again, the key issue is choosing the right background set.

          The focus in some of these discussions on gene length is peculiar. Really the total number of counts is the dominant variable here, on which detection power depends.

          Best wishes
          Wolfgang
          I never really understood the focus on gene length, perhaps it is my naivete of the finer aspects of the statistics. I understand the need to adjust for this if comparing expression of gene A to gene B, but not when comparing expression of gene A in condition A to gene A in condition B. The ability to detect DE in smaller genes can be overcome by increasing sequencing depth, so that like you said, "total number of counts is the dominant variable".

          I tried incorporating the methodology of GOseq with very confounding results, but we had also sequenced to depths where we were not seeing significant increases in the detection of lowly expressed genes.

          Comment

          • billstevens
            Senior Member
            • Mar 2012
            • 120

            #6
            What I imagine is that it matters to actually call something differentially expressed because if longer genes give off higher counts, then DESeq requires a lower threshold of fold change to call it differentially expressed, even though it actually had the same level of expression of a shorter gene.



            This is where Simon explains that longer genes can show up as more expressed, but he says this bias cannnot be dealt with by the test for differential expression. He then states it should be taken into account by the gene enrichment test, but I'm not quite understanding how.
            Last edited by billstevens; 09-19-2012, 01:31 PM.

            Comment

            • billstevens
              Senior Member
              • Mar 2012
              • 120

              #7
              Sorry, I think I should have asked that question more explicitly. How can we deal with the bias in counting methods in the gene enrichment?

              Comment

              • billstevens
                Senior Member
                • Mar 2012
                • 120

                #8
                bump?

                Simon? Wolfgang? Bueller?

                Comment

                Latest Articles

                Collapse

                • SEQadmin2
                  Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                  by SEQadmin2



                  Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                  ...
                  07-09-2026, 11:10 AM
                • SEQadmin2
                  Cancer Drug Resistance: The Lingering Barrier to Rising Survival
                  by SEQadmin2



                  Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

                  There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
                  07-08-2026, 05:17 AM
                • GATTACAT
                  Reply to Nine Things a Sample Prep Scientist Thinks About Before Sequencing
                  by GATTACAT
                  Love this - good data definitely starts from good input, and poor input can only give relatively poor data. I particularly like the mention of Nanodrop/absorbance based methods for quantification. It's such a toss up if you'll get an accurate reading or what amounts to a randomly generated number, and a lot of library/sequencing related issues can be traced back to poor quant.
                  07-01-2026, 11:43 AM

                ad_right_rmr

                Collapse

                News

                Collapse

                Topics Statistics Last Post
                Started by SEQadmin2, 07-13-2026, 10:26 AM
                0 responses
                28 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-09-2026, 10:04 AM
                0 responses
                37 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-08-2026, 10:08 AM
                0 responses
                25 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-07-2026, 11:05 AM
                0 responses
                35 views
                0 reactions
                Last Post SEQadmin2  
                Working...