Seqanswers Leaderboard Ad

**krobison** · 05-31-2013, 07:10 AM

#distinct kmers / 2 should be the genome size with a few important caveats

1) Including erroneous kmers will inflate the count, so typically would count only those kmers with a count of >=2

2) Repeat regions will be collapsed

3) regions that just don't show up will be missed, again underestimating. With high G+C genome, there may be regions simply missing from Illumina or with very low coverage.

Ray produces the kmer statistics in a way that is easy to parse & generate these estimates.

Assemblies are often a bit too large due to missed overlaps. If you convert these histograms to genome size estimates, how big a range is covered?

Even without a reference, the taxonomy of the bug may suggest a range -- though you could well have something outside that range.

**Krish_143** · 05-31-2013, 07:28 AM

Hi krobison,

when i estimted the genome size using kmer information (histogram, kmer Peaks)
ESti_Gsize: 2.8mb (at Kmer 31)
Assembled Gsize using SoapDenovo : 5.7mb (Draft)

I will check with the Ray and very thanks krobison for the quick response.

**rchikhi** · 06-02-2013, 02:13 PM

I sometimes observe that SOAPdenovo contigs (not scaffolds) tend to assemble more than the genome size. Did you run a Velvet assembly, and if so, what was the assembly size?

Topics	Statistics	Last Post
Expanding the Horizons of Cellular Research with the Single Cell Atlas by seqadmin Started by seqadmin, 04-25-2024, 11:49 AM	0 responses 20 views 0 likes	Last Post by seqadmin 04-25-2024, 11:49 AM
Genetic Variants and Diabetes Risk in Childhood Cancer Survivors by seqadmin Started by seqadmin, 04-24-2024, 08:47 AM	0 responses 20 views 0 likes	Last Post by seqadmin 04-24-2024, 08:47 AM
Cancer Metastasis: A Deep Dive into Cellular Plasticity by seqadmin Started by seqadmin, 04-11-2024, 12:08 PM	0 responses 62 views 0 likes	Last Post by seqadmin 04-11-2024, 12:08 PM
Proteogenomic Profiles Offer New Clues in Prostate Cancer by seqadmin Started by seqadmin, 04-10-2024, 10:19 PM	0 responses 61 views 0 likes	Last Post by seqadmin 04-10-2024, 10:19 PM

Seqanswers Leaderboard Ad

Announcement

Estimating the bacterial genome size using Kmer frequency

Comment

Comment

Comment

Latest Articles

ad_right_rmr

News