Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • ZoeG
    Member
    • Jun 2013
    • 31

    how to utilize unmapped paired-end data

    Hi, guys, I am thinking the way to utilize those unmapped reads.
    When I mapped PE100 data using TopHat, some reads were threw out as unmapped reads.
    When I looked into these unmapped reads, though a lot are garbage or mis-matched reads, I found some of them contained information. Thus, I am thinking to utilize these unmapped data. I gave TopHat a specific GTF and asked the data mapped to this GTF with -G option.
    However, I found that though with a good mapping rate (>80%), only a small percentage of reads are properly aligned (~20%). I think that's because the unmapped data of paired-end reads ( left vs right) are not synchronized.
    Can I synchronize the unmapped left and right reads with some recommended tools?
    Or do you think if this idea worth time spending?
    Thanks,
  • Cofactor Genomics
    Registered Vendor
    • Jan 2010
    • 52

    #2
    Hi ZoeG,
    I wonder what your ultimate goal is in the project.

    My team uses these unmapped reads many times in a way to build confidence in a possible insertion sequence between the reference and the sequenced organism by identifying one leg of the pair mapping to the reference and another that does not. This information in concert with other key coverage characteristics provides these insertion results.

    However, if you are ultimately after SNPs and small INDELs the inclusion of this data is not worth it... it only casts doubt on your more confident mappings.

    My best.

    Jarret Glasscock
    Cofactor Genomics

    Comment

    • ZoeG
      Member
      • Jun 2013
      • 31

      #3
      Thanks for the advice, Jarret.
      I am thinking whether it is possible to map the long genes first and then, map relatively small genes in the unmapped reads. I hope that the mis-matches for short genes could be decreased in this way.
      But there may be a risk of a loss of reads containing short genes ?

      How to align the unmapped reads is also a question for me since I am relatively new to the field.


      Originally posted by Cofactor Genomics View Post
      Hi ZoeG,
      I wonder what your ultimate goal is in the project.

      My team uses these unmapped reads many times in a way to build confidence in a possible insertion sequence between the reference and the sequenced organism by identifying one leg of the pair mapping to the reference and another that does not. This information in concert with other key coverage characteristics provides these insertion results.

      However, if you are ultimately after SNPs and small INDELs the inclusion of this data is not worth it... it only casts doubt on your more confident mappings.

      My best.

      Jarret Glasscock
      Cofactor Genomics
      http://www.cofactorgenomics.com

      Comment

      Latest Articles

      Collapse

      • SEQadmin2
        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
        by SEQadmin2


        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

        The systematic characterization of the human proteome has
        ...
        07-20-2026, 11:48 AM
      • SEQadmin2
        Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
        by SEQadmin2



        Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
        ...
        07-09-2026, 11:10 AM
      • SEQadmin2
        Cancer Drug Resistance: The Lingering Barrier to Rising Survival
        by SEQadmin2



        Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

        There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
        07-08-2026, 05:17 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by SEQadmin2, 07-20-2026, 11:10 AM
      0 responses
      14 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-13-2026, 10:26 AM
      0 responses
      31 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-09-2026, 10:04 AM
      0 responses
      42 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-08-2026, 10:08 AM
      0 responses
      27 views
      0 reactions
      Last Post SEQadmin2  
      Working...