Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • Simon Anders
    Senior Member
    • Feb 2010
    • 995

    #16
    Originally posted by Marianna85 View Post
    With bot the normalization methods I obtain size factors very different:
    -using DESeq 0,095 for one library and 10,85 for the other.
    -using edgeR 0,14 and 7,2 respectively.
    This is because edgeR gives is norm factors relative to the total read count. To get expression values on a scale comparable across sample, you have to divide the counts
    - for DESeq, just by the size factod
    - for edgeR, by the total read counts and by the normalization factors

    Comment

    • Marianna85
      Member
      • Mar 2012
      • 32

      #17
      Hi Simon,
      happy to read your answer

      Originally posted by Simon Anders View Post
      This is because edgeR gives is norm factors relative to the total read count. To get expression values on a scale comparable across sample, you have to divide the counts
      - for DESeq, just by the size factod
      - for edgeR, by the total read counts and by the normalization factors
      Just an example to better understand.
      For DESeq
      gene A raw counts; 5 reads in library 1 - 70 reads in library 2
      library 1:36 million reads size factor: 0.09
      library 2:64 million reads size factor: 10
      gene A normalized counts lib 1=5/0.09 - lib2=70/10

      For edgeR
      gene A raw counts; 5 reads in library 1 - 70 reads in library 2
      library 1:36 million reads size factor: 0.14
      library 2:64 million reads size factor: 7.2
      gene A normalized counts lib 1=(5/36 million)/0.14 - lib2=(70/64million)/7.2

      in this case with a very huge difference in library size, it seems better to normalize with edgeR. Isn't it?

      Thanks a lot.
      I really appreciate your answer.

      Marianna

      Comment

      • Shanrong
        Junior Member
        • Dec 2012
        • 7

        #18
        I am using both edgeR and DESeq to analyze my own dataset. In general, the fold change reported by both are close (should be, right?).

        If your orignal librares have 36 vs 64 million of total reads, your normalization factors from both methods are very weird (at least quite unusual). Please check whether you analyze your dataset properly.

        The idea of normalization behind edgeR and DESeq is very similar to each other (but implementation is different). In practice, I don't see one method is superior than the other. However, it is definitely a mistake if we feed DESeq with the normalization factor from edgeR, and vice versa.

        Comment

        • Jeremy
          Senior Member
          • Nov 2009
          • 190

          #19
          Something is going wrong somewhere, 36M and 64M reads should give normalization factors with less than a 2-fold difference. The normalization factors you listed had a 50 fold difference and suggest a much greater difference in read count.
          Last edited by Jeremy; 03-04-2013, 05:50 PM.

          Comment

          • Marianna85
            Member
            • Mar 2012
            • 32

            #20
            Originally posted by Jeremy View Post
            Something is going wrong somewhere, 36M and 64M reads should give normalization factors with less than a 2-fold difference. The normalization factors you listed had a 50 fold difference and suggest a much greater difference in read count.
            Hi Jeremy,
            in fact I was surprised to obtain such a difference...
            I've not yet understood which is the mistake in the size factor calculation.

            Comment

            • Simon Anders
              Senior Member
              • Feb 2010
              • 995

              #21
              Have you already looked at a scatter plot comparing the counts for the two samples? This should clarify what is going on.

              Comment

              • Marianna85
                Member
                • Mar 2012
                • 32

                #22
                Originally posted by Simon Anders View Post
                Have you already looked at a scatter plot comparing the counts for the two samples? This should clarify what is going on.
                Simon, do you mean the estimateDispersions?

                Comment

                • Marianna85
                  Member
                  • Mar 2012
                  • 32

                  #23
                  This is the script I used

                  CountTable=read.table("decEggs.txt", header=TRUE, row.names=1 )
                  head(CountTable)
                  decDesign = data.frame(row.names = colnames( CountTable ), condition = c( "stripped", "spawned" ), libType = c( "paired-end", "paired-end" ) )
                  decDesign
                  pairedSamples = decDesign$libType == "paired-end"
                  condition = decDesign$condition[ pairedSamples ]
                  library( "DESeq" )
                  cds = newCountDataSet( CountTable, condition )
                  cds = estimateSizeFactors( cds )
                  sizeFactors( cds )
                  head( counts( cds, normalized=TRUE ) )


                  and the size factors have a 100 fold difference...
                  what should I do???

                  Comment

                  • Marianna85
                    Member
                    • Mar 2012
                    • 32

                    #24

                    Comment

                    • Simon Anders
                      Senior Member
                      • Feb 2010
                      • 995

                      #25
                      Originally posted by Marianna85 View Post
                      Simon, do you mean the estimateDispersions?
                      No, I mean a scatter plot of the reads.

                      Try, e.g.,

                      Code:
                      plot( log10( 1 + counts(cds)[1,] ), log10( 1 + counts(cds)[2,] ), pch="." )
                      to plot the raw, unnormalized read counts of the second sample versus the first on a log scale.

                      Comment

                      • Marianna85
                        Member
                        • Mar 2012
                        • 32

                        #26
                        Hi Simon,
                        the plot seems empty...
                        may I change the axis scale?

                        Comment

                        • Simon Anders
                          Senior Member
                          • Feb 2010
                          • 995

                          #27
                          This will be hard to debug via the forum. You may need to get some local help.

                          To try one thing: If you simply type "counts(cds)", you get your table of raw counts (or, if you just want the first 100 lines, try "head( counts(cds), 100 )". Check whether they make sense.

                          Comment

                          • Simon Anders
                            Senior Member
                            • Feb 2010
                            • 995

                            #28
                            Sorry, I made a type. It's

                            Code:
                            plot( log10( 1 + counts(cds)[,1] ), log10( 1 + counts(cds)[,2] ), pch="." )

                            Comment

                            • Marianna85
                              Member
                              • Mar 2012
                              • 32

                              #29
                              of course! I defined the cds rows, not the columns.
                              So this is the plot...



                              something strange in your opinion??

                              Comment

                              • Simon Anders
                                Senior Member
                                • Feb 2010
                                • 995

                                #30
                                ".emf"? That's Windows extended metafile, right? Haven't seen this graphics file format in ten years, and frankly, I have no idea how to open it. Could you use something more common, please, maybe png?

                                Comment

                                Latest Articles

                                Collapse

                                • SEQadmin2
                                  New Genomics Technologies Take Aim at Long-Standing Limits
                                  by SEQadmin2


                                  Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

                                  We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
                                  ...
                                  Today, 10:25 AM
                                • SEQadmin2
                                  How Immunogenomics Decodes Immunity’s Genetic Blueprint
                                  by SEQadmin2




                                  The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

                                  This convergence of genetics, immunology, and computation...
                                  09-01-2026, 05:41 AM

                                ad_right_rmr

                                Collapse

                                News

                                Collapse

                                Topics Statistics Last Post
                                Started by SEQadmin2, 09-25-2026, 09:06 AM
                                0 responses
                                27 views
                                0 reactions
                                Last Post SEQadmin2  
                                Started by SEQadmin2, 09-23-2026, 11:05 AM
                                0 responses
                                24 views
                                0 reactions
                                Last Post SEQadmin2  
                                Started by SEQadmin2, 09-18-2026, 11:37 AM
                                1 response
                                45 views
                                0 reactions
                                Last Post pekgio
                                by pekgio
                                 
                                Started by SEQadmin2, 09-16-2026, 10:23 AM
                                1 response
                                55 views
                                0 reactions
                                Last Post pekgio
                                by pekgio
                                 
                                Working...