Seqanswers Leaderboard Ad

**swbarnes2** · 03-20-2013, 10:38 AM

You don't need to sort after the rmdup.

But yes, sometimes a library really is that repetitive. Maybe you have way more reads than your small genome needs. Maybe your library prep got screwed up.

Rule #1 Don't disbelieve your data just because you don't like what its telling you.

If you like, use Picard MarkDuplicates instead, because it will only mark, and not remove the duplicates, and then you can eyeball and see the repetition. Or look at the first sort file in IGV, or something like that, and see if you can see the duplicates.

**mmmm** · 03-14-2014, 04:08 AM

using flagstat- what does this mean?
"76 reads have mapped to a different chromosome than their pair" (BWA was used for reads mapping to the reference bacterial genome)

**dpryan** · 03-14-2014, 04:34 AM

The origin of that is comparison of the RNAME and RNEXT columns. If they're not the same, then this is incremented. I would have to presume that the bacterial genome you aligned against contained not just the main genome but also some plasmids (or perhaps it was all just a bunch of contigs).

**mmmm** · 03-14-2014, 04:45 AM

yes- there is a plasmid but it showed only 1 SNP

I have got these stats for mapping to the chromosome and plasmid- does "76 + 0 with mate mapped to a different chr" refere to repitative sequences that are shared between plasmid and chromosome

3753203 + 0 in total (QC-passed reads + QC-failed reads)
0 + 0 duplicates
3660519 + 0 mapped (97.53%:-nan%)
3753203 + 0 paired in sequencing
1876530 + 0 read1
1876673 + 0 read2
3644736 + 0 properly paired (97.11%:-nan%)
3653854 + 0 with itself and mate mapped
6665 + 0 singletons (0.18%:-nan%)
76 + 0 with mate mapped to a different chr
74 + 0 with mate mapped to a different chr (mapQ>=5)

**dpryan** · 03-14-2014, 04:50 AM

Given that 74/76 have MAPQ>=5 it would seem that these are unlikely to be that repetitive. However, they might be technical artefacts, since there are so few of them.

**mmmm** · 03-14-2014, 04:56 AM

thanks- if I want to determine the (0.18%) of unmapped reads- what is the simplest way for this?- should I use velvet for denovo assembly (kmer 31) then annotate the output?

**dpryan** · 03-14-2014, 05:06 AM

~2.5% and yeah that'd be one approach. You might have a look at the reads first to see if they just didn't align due to low quality or something like that.

**mmmm** · 03-14-2014, 07:07 AM

thank you- what do you mean by looking at the reads (do you mean viewing the bam file using IGV)

**dpryan** · 03-14-2014, 10:56 AM

They're unmapped, so you would just "zcat unmapped.fq.gz | less" or even just run the unmapped reads through fastqc. The general idea is to first see if they're even worth assembling (if they're all repetitive regions or low quality then there's no reason to bother).

Topics	Statistics	Last Post
A Closer Look at the Enigmatic Genomes of Oikopleura dioica by seqadmin Started by seqadmin, 05-10-2024, 06:35 AM	0 responses 19 views 0 likes	Last Post by seqadmin 05-10-2024, 06:35 AM
Advanced Epigenome Editing Platform Explores Gene Regulation Mechanisms by seqadmin Started by seqadmin, 05-09-2024, 02:46 PM	0 responses 22 views 0 likes	Last Post by seqadmin 05-09-2024, 02:46 PM
Telomere Maintenance by PARP1: A New Perspective in Cancer Research by seqadmin Started by seqadmin, 05-07-2024, 06:57 AM	0 responses 21 views 0 likes	Last Post by seqadmin 05-07-2024, 06:57 AM
Enhanced Neoantigen Detection: Introducing NeoHunter by seqadmin Started by seqadmin, 05-06-2024, 07:17 AM	0 responses 21 views 0 likes	Last Post by seqadmin 05-06-2024, 07:17 AM

Seqanswers Leaderboard Ad

Announcement

flagstat

Comment

Comment

Comment

Comment

Comment

Comment

Comment

Comment

Comment

Latest Articles

ad_right_rmr

News