Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • huan
    Member
    • Oct 2010
    • 56

    Is it possible to evaluate genome size with sequel data?

    Now we are doing the denovo assembly of marine organism with whole genome sequcing using sequel system. As we all know, the DNA extraction from marine organism is very difficult because of pollution and degradation. So is there any way to evaluate the genome size, heterozygus rate or genome repeat with DNA sequel data?
    happy
  • Markiyan
    Senior Member
    • Sep 2010
    • 126

    #2
    Use multipass pacbio reads for self error correction and Kmer counting.

    First try filtering out the multipass reads, and using those for kmer counting and self error correction.

    Make sure to remove any mitochondrial/symbionts reads before doing the kmer counting. (Identify and complete the respective genome(s) first).

    Get some good quality PCR-free illumina 2x250 reads or (BGIseq data if it works in your hands) and use it to confirm the kmer counting/self error correction/etc.

    Short reads are very helpful for getting the contaminant(s)/symbionts genomes to a good draft stage and for filtering them out from the main dataset.
    Usually such approach has to be done in the iterative fashion (with increasing amount of the input data after each iteration).

    Comment

    • luc
      Senior Member
      • Dec 2010
      • 469

      #3
      Markiyan has alluded to it already; Pacbio data are not suitable for genome size estimates based on kmer analyses. The error rates of the uncorrected raw data are too high.

      Comment

      • rhall
        Senior Member
        • Aug 2012
        • 324

        #4
        While a kmer analysis is going to be difficult with the raw pacbio data, it is possible to estimate the (effective) genome size from overlap statistics, either for the raw reads, the error corrected preassembled reads or by mapping the raw reads to the assembled contigs.
        Run an initial assembly using a small seed read length, then plot the preassembled read overlap histogram.


        Comment

        • huan
          Member
          • Oct 2010
          • 56

          #5
          I really appreciate for your help! I will have a try!
          happy

          Comment

          Latest Articles

          Collapse

          • mylaser
            Reply to Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by mylaser
            The world of online gaming has grown tremendously over the past few years, giving players access to exciting sports, casino games, and interactive entertainment from the comfort of their homes. Among the platforms gaining attention, Kheloyaar has become a trusted destination for users seeking a fast, secure, and engaging gaming experience.
            Whether you're a first-time visitor or an existing user, understanding the features of Kheloyar and the Kheloyaar login process can help you enjoy everything...
            Today, 12:33 AM
          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            Yesterday, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, Yesterday, 11:10 AM
          0 responses
          8 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-13-2026, 10:26 AM
          0 responses
          30 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-09-2026, 10:04 AM
          0 responses
          39 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-08-2026, 10:08 AM
          0 responses
          25 views
          0 reactions
          Last Post SEQadmin2  
          Working...