Unconfigured Ad

**Brian Bushnell** · 03-11-2014, 11:00 PM

Are you sure they're PCR duplicates rather than real duplicates? If so, how do you know?

**ffinkernagel** · 03-12-2014, 12:32 AM

You know because a library of a million 'effective reads' (after dedup) distributed in the largish genome of a mammal isn't useful, and no antibody is so good that you'll only get the enriched regions. And if they were real, you'd still have many starts in a small region, not a start here, one there, one over yonder...

I basically stand my ground. People come to me because I have the experience in these data sets, and my advice is: find the error and repeat the experiment.

Do not waste time with this unusable data set. Mostly they come around when I start repeating the old 'if your positive control is no different from your negative control you can draw no conclusions from this experiment' mantra.

If they're desperate or dense enough not to grasp the above point, it's best to cut your losses and suggest they find somebody else to look at their data.

**Anomilie** · 03-12-2014, 03:02 PM

Thanks ffinkernagel for sharing your experience.

To answer your question, Brian Bushnell, the way that I assessed PCR duplicates is though the following:

1) In FASTQC look at the duplication level tab, this will give a rough estimate.

Result 90-95% duplication level.

2) On the file containing aligned reads (usually bam file) calculate the fraction of non-redundant reads (NRF in Encode guidelines) by calculating the number of unique genomic positions/all uniquely mapped read.

Result: 12-17% of reads are non-redundant

3) Sort your aligned reads (bam file) according to chromosomal location and perform samtools rmdup and picard MarkDuplicates (with option REMOVE_DUPLICATES=TRUE). Calculate the percentage of reads remaining from the original and assess if there are any difference between samtools and Picard.

Result: between 95- 99% of reads removed

4) Visualize your aligned reads in a genome viewer such as IGV. If the reads stack up on top of each other, with black spaces in between stacks, rather than diagonally overlapping reads, you have PCR duplicates

**Chacal** · 01-07-2015, 11:10 PM

Anomilie,

I do not get point 4 of your last post. Could you provide a visual for vertically stacked reads versus diagonally stacked reads from IGV? I am doing ATAC and I get lots of duplicate reads 50-90% despite all efforts to optimise cell numbers and PCR cycles for library amplification.

Thanks you.

**Brian Bushnell** · 01-08-2015, 10:12 AM

Point 4 is the most important one, or in my opinion, the only one really able to indicate a difference between PCR and real duplicates. If the coverage is randomly distributed (in terms of start coordinates), then duplication events are implied to be real duplicates that naturally result from high sequencing depth. However, if you have a few stacks in which all the reads line up perfectly with the same start location, and few or no reads starting between the stacks, this implies PCR duplicates.

**Chipper** · 01-08-2015, 12:35 PM

The percentage of duplicated reads is not meaningful without knowing the total number. It is possible to have an excellent ChIP to yield only a few million unique reads, if you sequence this on one lane most reads will be duplicates. And if you have one good replicate and one that failed then the IDR will only give you the common false positives. Just call peaks on unique starts and look at the wiggle in a genome browser.

Topics	Statistics	Last Post
Whole-Genome Sequencing Traces Faroe Islands Ancestry to a North Atlantic Founder Population by SEQadmin2 Started by SEQadmin2, Yesterday, 06:09 AM	0 responses 16 views 0 reactions	Last Post by SEQadmin2 Yesterday, 06:09 AM
Sequencing the Two-Toed Sloth Genome Reveals Jumping Genes Tied to Its Extreme Metabolism by SEQadmin2 Started by SEQadmin2, 06-09-2026, 11:58 AM	0 responses 37 views 0 reactions	Last Post by SEQadmin2 06-09-2026, 11:58 AM
A New Method Makes Hantavirus Genome Analysis Faster and More Accessible by SEQadmin2 Started by SEQadmin2, 06-05-2026, 10:09 AM	0 responses 42 views 0 reactions	Last Post by SEQadmin2 06-05-2026, 10:09 AM
A New Single-Cell Method Maps DNA-Protein Interactions by SEQadmin2 Started by SEQadmin2, 06-04-2026, 08:59 AM	0 responses 49 views 0 reactions	Last Post by SEQadmin2 06-04-2026, 08:59 AM

Unconfigured Ad

Failed chip-seq experiments

Comment

Comment

Comment

Comment

Comment

Comment

Latest Articles

ad_right_rmr

News