Folks,
I'd like to hear your advice, suggestion and practice on how to statistically identify difference between two sequencing datasets. This applies to many analysis I encountered.
Suppose I have only one control and one experimental sample. I only have one sequencing data on each sample and the sequencing data can be RNA-Seq, ChIP-seq, FAIRE-seq etc, I want to compare the two sequencing data and identify either those most differentially expressed genes for RNA-Seq, or those regions that show the most differentially transcription factor binding between two samples, or those most differentially methylated regions between two samples etc... given I only have one sequencing data on each sample, what would you do to get those differentially "changed" sets of genes/regions?
I know some people calculate rpkm difference for the same gene/region and take the top difference as the most differentially "changed" sets of genes/regions. But this doesn't have statistical significance I guess. So how do you go about doing this in a statistically sound way?
I hope I stated my questions clearly. Thanks!
I'd like to hear your advice, suggestion and practice on how to statistically identify difference between two sequencing datasets. This applies to many analysis I encountered.
Suppose I have only one control and one experimental sample. I only have one sequencing data on each sample and the sequencing data can be RNA-Seq, ChIP-seq, FAIRE-seq etc, I want to compare the two sequencing data and identify either those most differentially expressed genes for RNA-Seq, or those regions that show the most differentially transcription factor binding between two samples, or those most differentially methylated regions between two samples etc... given I only have one sequencing data on each sample, what would you do to get those differentially "changed" sets of genes/regions?
I know some people calculate rpkm difference for the same gene/region and take the top difference as the most differentially "changed" sets of genes/regions. But this doesn't have statistical significance I guess. So how do you go about doing this in a statistically sound way?
I hope I stated my questions clearly. Thanks!
Comment