Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • hlwright
    Member
    • Feb 2011
    • 30

    Newbler runMapping via command line

    Hello everyone, I am new to this forum and this is my first post so I hope someone can help me.

    I have some 454 transcriptomic data which I am trying to analyse using Newbler mapping to the human GRCh37.61 cDNA fasta reference. I am having to run Newbler via the command line at the moment as I do not have enough RAM to launch it via Java. However, it seems to be running quite well this way so I am not too bothered about not being able to launch it via Java.

    However, I am exploring the command line options to try and improve the number of reads which are fully/partially mapping. I am still getting a large number of reads which are classified as repeats and I wondered if anyone had any tips on how to improve the quality of my mapping. (The default settings gave me 11% fully mapped, 6% partially mapped, 32% unmapped, 41% repeat, 6.5% chimeric, 3.5% too short).

    I have tried decreasing the seed length from 16 to 10 and this greatly decreased the number of unmapped reads, but increased the number of repeats (almost 50%). I have also changed the repeat score threshold from default (12) to 0 which has improved it a bit more and has greatly increased the number of contigs generated. I am now playing with the minimum overlap length but am getting more chimeric reads.

    I am really just arbitrarily changing these numbers and could sit here from now until Christmas doing this, so I wondered if anyone had any advice or tips they could give me.

    Before you ask why I am not using the assembler, well I just don't think I have enough reads to get a good assembly. My dataset contains around 50,000 reads per sample. What do you think?

    Any advice would be very much appreciated. Thank you in advance.

    Helen
  • flxlex
    Moderator
    • Nov 2008
    • 412

    #2
    Reads marked 'Repeat' map equally well to multiple locations in the reference. The settings you are trying are not going to change that...

    The only thing I can think of is to have more stringent alignment requirements, so that perhaps these reads start mapping uniquely (i.e. reads from different paralogues mapping to just one of the copies). This can be done by

    - increasing the minimum overlap length -ml, default is 40 bases, but you can go up to higher numbers, or even better, use '-ml 90%' to force at least 90% of the length of the read to map (or try 95%).
    - increasing the minimum overlap identity, -mi, default 90, but you could try '-mi 95' (no % here).

    On the other hand, you might get less reads mapped this way....

    Good luck anyways!

    Comment

    • hlwright
      Member
      • Feb 2011
      • 30

      #3
      Thank you for replying so quickly. I have been exploring many options with Newbler mapping.

      Unfortunately, the options you suggested did not improve the number of reads mapped. However, I think I may have worked out the problem. I am using a cDNA fasta reference as I have transcriptome reads. I have had a look at some of the reads which are 'unmapped' and a quick BLAST of a couple shows these are ribosomal RNAs (and as such will not be in my cDNA fasta file).

      I wonder if anyone else has noticed this in the past? Do you know of a fasta file containing rRNA sequences that I could concatenate with my cDNA reference to maybe annotate my 'unmapped' reads?

      Thank you
      Helen

      Comment

      Latest Articles

      Collapse

      • SEQadmin2
        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
        by SEQadmin2


        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

        The systematic characterization of the human proteome has
        ...
        07-20-2026, 11:48 AM
      • SEQadmin2
        Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
        by SEQadmin2



        Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
        ...
        07-09-2026, 11:10 AM
      • SEQadmin2
        Cancer Drug Resistance: The Lingering Barrier to Rising Survival
        by SEQadmin2



        Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

        There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
        07-08-2026, 05:17 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by SEQadmin2, Today, 11:41 AM
      0 responses
      9 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-20-2026, 11:10 AM
      0 responses
      21 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-13-2026, 10:26 AM
      0 responses
      34 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-09-2026, 10:04 AM
      0 responses
      44 views
      0 reactions
      Last Post SEQadmin2  
      Working...