Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • jamal
    Member
    • Jan 2010
    • 10

    huge number of DEGs from cuffdiff

    Hi
    Im working with RNA-seq data from solexa.I have 32 lanes each belonging to one sample and would like to test 16 vs 16.I have a problem when I run cuffdiff on the data I will get large number of DEGs like 3000 or 4000.I have got 72 DEGs out of this data with edgeR.
    I have changed many option but still I have many genes.
    Could you help me to get reasonable number of DEGs?
    Thank you
  • severin
    Genome Informatics Facility
    • Sep 2009
    • 105

    #2
    version

    Originally posted by jamal View Post
    Hi
    Im working with RNA-seq data from solexa.I have 32 lanes each belonging to one sample and would like to test 16 vs 16.I have a problem when I run cuffdiff on the data I will get large number of DEGs like 3000 or 4000.I have got 72 DEGs out of this data with edgeR.
    I have changed many option but still I have many genes.
    Could you help me to get reasonable number of DEGs?
    Thank you

    Hi Jamal, I would start by asking if you have the lastest version of Cuffdiff and to ask what threshold of FDR you are using for edgeR. I know in Cuffdiff they use an FDR threshold of 0.05

    Comment

    • jamal
      Member
      • Jan 2010
      • 10

      #3
      Hi Severin

      Im using the newest version of cufflinks,and I have tried FDR 0.05 and 0.01.also Im trying different options which I think they can make less number of DEGs like minimum number of alignments in a locus,I have set it to even 5000 but still too many genes.that will be great if any body can help.
      Thanks

      Comment

      • Simon Anders
        Senior Member
        • Feb 2010
        • 995

        #4
        Hi Jamal

        as you have very many replicates, you can try the following: Take one of your treatment groups of 16 samples, divide it up at random into two sub-groups of eight samples each. Then, use each of the tools you are considering to perform an eight-versus-eight comparison. As the samples are all replicates, the between group variance should be equal to the within-group variance, the tool should notice that, and hence report only very few results. If cuffdiff again reports a large number, you shouldn't believe it. (We had a similar discussion earlier here on SeqAnswers, and cufdiff failed this test back then; but that was an older version.)

        For edgeR, it is crucial that you use 'tagwise dispersion' to make full use of your many replicates and use a rather low 'prior.n' value (see rule-of-thumb in the vignette). Did you do that? As author of DESeq, I'd like to suggest, of course, that you try our tool as well, but actually, I guess it will give quite similar results as edgeR.

        You haven't told us much about ypur experimental design, but if each sample is a different individual from an outbred population, and the difference between your group is not something very drastic (e.g., if you compare non-related humans, who differ by some not that dramatic trait), it is entirely possible that 16 samples give you way to little power to detect more than the few genes you found.

        Comment

        • 11xinqi
          Member
          • Mar 2011
          • 31

          #5
          Originally posted by jamal View Post
          Hi
          Im working with RNA-seq data from solexa.I have 32 lanes each belonging to one sample and would like to test 16 vs 16.I have a problem when I run cuffdiff on the data I will get large number of DEGs like 3000 or 4000.I have got 72 DEGs out of this data with edgeR.
          I have changed many option but still I have many genes.
          Could you help me to get reasonable number of DEGs?
          Thank you
          Hi

          Did you find out why you get this result? BC I am in the same situation. I have two groups, one is control, the other is treatment with a specific gene knockdown. each group has 3 replicates. I got 3000 genes DEGs from cuffdiff. But only around 70 genes from edgeR. A weird thing is that a lot of DEGs from cuffdiff only expressed in one replicate of one group. For example, gene A is only be expressed in one replicate of control group, but not in other two replicates and treatment group!

          Comment

          • 11xinqi
            Member
            • Mar 2011
            • 31

            #6
            Originally posted by Simon Anders View Post
            Hi Jamal

            as you have very many replicates, you can try the following: Take one of your treatment groups of 16 samples, divide it up at random into two sub-groups of eight samples each. Then, use each of the tools you are considering to perform an eight-versus-eight comparison. As the samples are all replicates, the between group variance should be equal to the within-group variance, the tool should notice that, and hence report only very few results. If cuffdiff again reports a large number, you shouldn't believe it. (We had a similar discussion earlier here on SeqAnswers, and cufdiff failed this test back then; but that was an older version.)

            For edgeR, it is crucial that you use 'tagwise dispersion' to make full use of your many replicates and use a rather low 'prior.n' value (see rule-of-thumb in the vignette). Did you do that? As author of DESeq, I'd like to suggest, of course, that you try our tool as well, but actually, I guess it will give quite similar results as edgeR.

            You haven't told us much about ypur experimental design, but if each sample is a different individual from an outbred population, and the difference between your group is not something very drastic (e.g., if you compare non-related humans, who differ by some not that dramatic trait), it is entirely possible that 16 samples give you way to little power to detect more than the few genes you found.
            Hi, Simon

            I have a question. Would you please help me? I have two groups, one is control, the other is treatment with a specific gene knockdown. each group has 3 replicates. I got 3000 genes DEGs from cuffdiff. But only around 70 genes from edgeR. A weird thing is that a lot of DEGs from cuffdiff only expressed in one replicate of one group. For example, gene A is only be expressed in one replicate of control group, but not in other two replicates and treatment group! Do you know why cuffdiff take these genes only expressed in one replicate as DEGs? And why I get so different result from cuffdiff and edgeR?

            Both groups are cow embryo. There are 3 different embryos in control group. The three treatment group contain 3 embryos with one gene knockdown. Do you think why I get just few DEGs is because each sample is a different individual from an outbred population, and the difference between your group is not something very drastic?

            Joseph

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM
            • SEQadmin2
              Cancer Drug Resistance: The Lingering Barrier to Rising Survival
              by SEQadmin2



              Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

              There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
              07-08-2026, 05:17 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Today, 11:41 AM
            0 responses
            8 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-20-2026, 11:10 AM
            0 responses
            21 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-13-2026, 10:26 AM
            0 responses
            34 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-09-2026, 10:04 AM
            0 responses
            44 views
            0 reactions
            Last Post SEQadmin2  
            Working...