Originally posted by aaronrjex
View Post
Unconfigured Ad
Collapse
X
-
This is a very helpful description of BGI's methods but I have one question. How do you determine M, what you call the mean k-mer coverage, from the k-mer frequency histogram produced by soapdenovo's pregraph program (the file with the *.kmerFreq extension), or with any other method? The first post in this thread quotes the methods from the Panda genome paper as saying M is the peak k-mer coverage (I guess on a normal distribution), and that is what is challenging to calculate.
-
Hello folks,
I am trying to understand the genome size estimation method, and here is the part that is not clear to me.
BGI is estimating coverage depth by peak of frequency histogram, but in highly repetitive genomes the peak shifts down. E.g. check the second chart here, where we generated simulated reads from sea urchin genome at 50X coverage, but the histogram peak came at ~40 due to repeats.
If that is true, how does BGI estimate genome size so precisely? In later section of bat paper, they claimed that the assembly size is close to the 'estimated size', which is puzzling given the impreciseness in genome size estimation.
Maybe I am missing something. Please help !!
---------------
Edit. I am rerunning the above simulation to make sure everything is done correctly. Results will be reported here.Last edited by samanta; 04-10-2013, 04:59 PM.
Comment
-
Lizhenyu from BGI explained where I made the error in thinking. When a read has 100 nucleotides and is split into 21-mers, the read will produce only 80 k-mers, not 100. Here is his full response, which agrees with aaronrjex's post above.
"Hi,
I think you mixed the concepts of base coverage depth and kmer coverage depth.
When you refered 50X genome coverage, it meant base coverage which is obtained by
total_base_num/genome_size=(read_num*read_length)/genome_size.
Similarly, the kmer coverage depth, the peak value in kmer frequency curve, is calculated by
total_kmer_num/genome_size=read_num*(read_length-kmer_size+1)/genome_size.
So the relationship between base coverage depth and kmer coverage depth is:
kmer_coverage_depth = base_coverage_depth*(read_length-kmer_size+1)/read_length.
In your case, kmer_coverage_depth = 50 * (100 - 21 + 1)/100 = 40, which is exactly
the peak value in you plot.
best,"
Comment
Latest Articles
Collapse
-
by SEQadmin2
The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.
This convergence of genetics, immunology, and computation...-
Channel: Articles
09-01-2026, 05:41 AM -
ad_right_rmr
Collapse
News
Collapse
| Topics | Statistics | Last Post | ||
|---|---|---|---|---|
|
Started by SEQadmin2, 09-03-2026, 10:22 AM
|
0 responses
22 views
0 reactions
|
Last Post
by SEQadmin2
09-03-2026, 10:22 AM
|
||
|
Started by SEQadmin2, 09-02-2026, 12:32 PM
|
0 responses
21 views
0 reactions
|
Last Post
by SEQadmin2
09-02-2026, 12:32 PM
|
||
|
Started by SEQadmin2, 08-24-2026, 10:32 AM
|
0 responses
52 views
0 reactions
|
Last Post
by SEQadmin2
08-24-2026, 10:32 AM
|
||
|
Started by SEQadmin2, 08-20-2026, 11:17 AM
|
0 responses
51 views
0 reactions
|
Last Post
by SEQadmin2
08-20-2026, 11:17 AM
|
Comment