Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • claire_
    Junior Member
    • Nov 2015
    • 2

    #1

    Finding significantly different ChIP-seq peak length?

    Hello,

    I have a simple data analysis question.

    Say I have a protein ChIP-seq data comparing three replicates of wild-type and three replicates of mutants. I have normalized the data using spikes, mapped reads, etc and let's say for the sake of the argument the data is normalized.

    A typical pipeline would then call peak then count how many reads in those peaks then compare WT vs mutant with statistical tests (e.g. using MAnorm or DESeq2 or other softwares that's built to find differentially "expressed" ChIP-seq)

    The ideal scenario is like this below for example, both WT and MUT peaks overlap pretty nicely in a hotspot such that we can use the method above to find significantly up/down ChIP-seq peaks by using the signal.



    However, what if it's not the signal that we want to compare, but the length? Let's say this protein upon mutation are more spread around the hotspots instead of in the hotspots. Therefore the overall signal does not change, but the shape or length does. Worse, the length change is not consistent but can be to the left, or to the right, or both, as such:



    The wild-type peak is very consistent in shape and length and only mutants change.

    The goal is really to test whether there is a significant length increase/decrease compared to wild type. Whether it's go to the right, left, or both doesn't matter. What statistical test would be the best?

    What I did was I first call peak, then merged peaks that are close together in the 6 samples (I call these hotspots). Then for each sample I sum their peak length in each merged peak group. When I looked at the distribution, the length for each sample really follows poisson/nbinom therefore I used DESeq2 to find significantly different length. Would this be an acceptable method?
  • dariober
    Senior Member
    • May 2010
    • 311

    #2
    Hi- It's a good question...! When you use methods like Deseq you compress all the information in a peak in a single number: The count of reads (or the length, in your case). Have a look at this paper and associated R package "MMDiff: quantitative testing for shape changes in ChIP-Seq data sets".

    Dario

    Comment

    • dpryan
      Devon Ryan
      • Jul 2011
      • 3478

      #3
      For something like this you might want to test for differences in the distribution, such as with a KS test.
      edit: MMDiff looks MUCH more interesting!

      Comment

      • dariober
        Senior Member
        • May 2010
        • 311

        #4
        Originally posted by dpryan View Post
        For something like this you might want to test for differences in the distribution, such as with a KS test.
        edit: MMDiff looks MUCH more interesting!
        Trying MMdiff out for both Chip-Seq and BS-Seq data is on my todo list, but still haven't got around it!

        Comment

        • claire_
          Junior Member
          • Nov 2015
          • 2

          #5
          Thanks for the MMdiff suggestion! Will definitely try it out!!

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            New Genomics Technologies Take Aim at Long-Standing Limits
            by SEQadmin2


            Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

            We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
            ...
            Yesterday, 10:25 AM
          • SEQadmin2
            How Immunogenomics Decodes Immunity’s Genetic Blueprint
            by SEQadmin2




            The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

            This convergence of genetics, immunology, and computation...
            09-01-2026, 05:41 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, 09-25-2026, 09:06 AM
          0 responses
          29 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 09-23-2026, 11:05 AM
          0 responses
          24 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 09-18-2026, 11:37 AM
          1 response
          46 views
          0 reactions
          Last Post pekgio
          by pekgio
           
          Started by SEQadmin2, 09-16-2026, 10:23 AM
          1 response
          55 views
          0 reactions
          Last Post pekgio
          by pekgio
           
          Working...