I was wondering how to interpret the Kmer content graph in fastqc. I have attached it. It seems from my graph that 50% of all reads have the 6 kmers listed ? Is this a normal graph ?
Seqanswers Leaderboard Ad
Collapse
Announcement
Collapse
No announcement yet.
X
-
Is this RNA-seq data?
Following thread may be useful (though it refers to a MiSeq run the issue is applicable to illumina sequencing in general): http://seqanswers.com/forums/showthread.php?t=30448
One more: http://seqanswers.com/forums/showthread.php?t=17219Last edited by GenoMax; 08-13-2013, 11:57 AM.
-
Thank you very much, GenoMax, for the links. They are very informative.
This is RNA-Seq. It seems the first 12 or so bases are due to 'random' priming, but I was also wondering about why the lines on the graph stay up at ~50% for the length of the graph ? Random priming would explain the first 12 or so bases, but why the steady % for the rest of the sequence ?
Comment
-
Originally posted by mattanswers View PostThank you very much, GenoMax, for the links. They are very informative.
This is RNA-Seq. It seems the first 12 or so bases are due to 'random' priming, but I was also wondering about why the lines on the graph stay up at ~50% for the length of the graph ? Random priming would explain the first 12 or so bases, but why the steady % for the rest of the sequence ?
Comment
-
Thanks again for your help, GenoMax.
My sequence length is only 50 bases and the quality is very good.
From what I read on the linked site, it seems that I have 6 kmers that are 50-fold enriched throughout the length of my sequence. But what does this mean in terms of sample quality ?
If I have 25-30 million reads and there is a 50 fold enrichment of these kmers (most likely I would guess from the adaptor) then how many sequences does that affect ? So, if there were 100,000 sequences in which had adaptor sequence at various positions other than the end of the sequence what would the fold-enrichment be ? 100,000 affected sequences may be enough to make the fold-enrichment high, but they are only a small percentage of the total. On the other hand, if I had a much smaller number of total sequences, then the fold-enrichment may be a problem. So, I guess I want to know how to relate fold-enrichment and total number of sequences in order to tell if the fold-enrichment is a problem or just from an insignificant part of the total.
Comment
Latest Articles
Collapse
-
by seqadmin
The complexity of cancer is clearly demonstrated in the diverse ecosystem of the tumor microenvironment (TME). The TME is made up of numerous cell types and its development begins with the changes that happen during oncogenesis. “Genomic mutations, copy number changes, epigenetic alterations, and alternative gene expression occur to varying degrees within the affected tumor cells,” explained Andrea O’Hara, Ph.D., Strategic Technical Specialist at Azenta. “As...-
Channel: Articles
07-08-2024, 03:19 PM -
ad_right_rmr
Collapse
News
Collapse
Topics | Statistics | Last Post | ||
---|---|---|---|---|
Started by seqadmin, 07-25-2024, 06:46 AM
|
0 responses
9 views
0 likes
|
Last Post
by seqadmin
07-25-2024, 06:46 AM
|
||
Started by seqadmin, 07-24-2024, 11:09 AM
|
0 responses
28 views
0 likes
|
Last Post
by seqadmin
07-24-2024, 11:09 AM
|
||
Started by seqadmin, 07-19-2024, 07:20 AM
|
0 responses
161 views
0 likes
|
Last Post
by seqadmin
07-19-2024, 07:20 AM
|
||
Started by seqadmin, 07-16-2024, 05:49 AM
|
0 responses
127 views
0 likes
|
Last Post
by seqadmin
07-16-2024, 05:49 AM
|
Comment