Thanks, Brian. This is where I am showing my ignorance I am sure, but how did the reads become so short? Looking at what I pulled out of the sam file, they are full-length (300bp) reads for the first few matches, but then become those little buggers are well.
Unconfigured Ad
Collapse
X
-
-
BWA-mem produces 'chimeric alignments'. This is actually a really neat feature in some cases, and a big pain in other cases - in my opinion, it should be disabled by default.
If you look at the sam lines you posted, most of them have a bitflag (the second column) of over 2048. That indicates they are chimeric. BWA-mem appears to do multiple local alignments on reads, such that if there is a really good match for the first 20% somewhere, that will be presented as a single line in the sam file, and if there is a really good match for the middle 40%, that will be displayed as a different line, etc. So a single read could generate a huge number of lines in the sam file. The goal is to correctly map reads that are chimeric (such as reads from a cancer sample with two chromosomes randomly fused together). But apparently, it does not work well in extreme-GC genomes; most mappers are designed for human and mouse genomes, which have approximately 50% GC, as they constitute the majority of genetic research. But since I work at a place that strictly deals with microbial, plant, and fungal genomes, BBMap (which was originally designed for human) is now developed for and tested on a much wider array of organisms than most.
BWA's chimeric alignments are local and hard-clipped. For example, this cigar string from the second line you posted - "221H79M" - means that the first 221 bases were ignored and only the last 79 bases are included in the alignment. Of course, this will wreak havoc with something like fastqc, where all reads are weighted equally regardless of length. Rather than a length filter (which will unnecessarily exclude reads that had been adapter- or quality-trimmed), I think you should simply use samtools to filter out reads with the chimeric flag marked.Last edited by Brian Bushnell; 08-02-2014, 12:00 PM.
Comment
-
Nucacidhunter has a nice description in this thread: http://seqanswers.com/forums/showthread.php?t=43071Originally posted by Genomics101 View PostThanks very much, GenoMax. Indeed, it is MiSeq data, but I never had this problem with MiSeq before (that was with 250bp PE reads, these are 300s). Can you tell me more the particular pathology with MiSeq? Is this a problem with library construction? And, goodness, what is an adapter lawn?
Comment
Latest Articles
Collapse
-
by SEQadmin2
Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.
We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing...-
Channel: Articles
-
ad_right_rmr
Collapse
News
Collapse
| Topics | Statistics | Last Post | ||
|---|---|---|---|---|
|
Started by SEQadmin2, 09-29-2026, 09:51 AM
|
0 responses
34 views
0 reactions
|
Last Post
by SEQadmin2
09-29-2026, 09:51 AM
|
||
|
Started by SEQadmin2, 09-25-2026, 09:06 AM
|
0 responses
44 views
0 reactions
|
Last Post
by SEQadmin2
09-25-2026, 09:06 AM
|
||
|
Started by SEQadmin2, 09-23-2026, 11:05 AM
|
0 responses
34 views
0 reactions
|
Last Post
by SEQadmin2
09-23-2026, 11:05 AM
|
||
|
Started by SEQadmin2, 09-18-2026, 11:37 AM
|
1 response
51 views
0 reactions
|
Last Post
by pekgio
09-21-2026, 02:04 AM
|
Comment