Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • wilflugo
    Junior Member
    • Dec 2014
    • 8

    'N-Blocks' or Any Nucleotide Alignments Best Practices

    How 'N-blocks' in a genome are normally handled on a Local and/or global alignment computation?
    I have been reading about this but seems I can't find what the 'best' practice should be. I see some tools that remove all N-blocks from a chromosome before doing the alignments and others assign a penalty whenever a nucleotide is compared against a 'N' nucleotide.

    As far as I understand (and I may be wrong) 'N' is an unknown nucleotide that has not being decoded. Thus in theory could be any one. However, if I assigned a match on a large N-block region, all individual reads are mapped incorrectly(?) into that region since the alignment see an exact match with all the highest scores into that region.

    Should N-blocks be simply trimmed out of the chromosome? Is there any other methodology to follow when doing alignments between reads and chromosomes (Assembly)?

    Sorry if this is common knowledge, but turns out that searching for 'N-blocks' is not a wise search key since it matches almost everything that is not relevant to this topic.

    thanks.
  • Brian Bushnell
    Super Moderator
    • Jan 2014
    • 2709

    #2
    Aligners generally ignore large blocks of Ns, or replace them with random sequence (which is not expected to match a read). When you align reads to chromosomes, just leave the Ns there, they won't hurt anything. If you remove them, your coordinates will get messed up.

    Comment

    • wilflugo
      Junior Member
      • Dec 2014
      • 8

      #3
      Thanks!. Haven't thought about the random sequence. Thank makes sense and it is probably the way I will be going.

      However, I am a little bit curious, how can you ignore N blocks when using optimal local alignment algorithms like Smith-Waterman where scores are based on match/mismatches on a pair of nucleotides?

      You could score 'zero' on all Ns, but depending on the similarity matrix used, zero could be a valid penalty value. I guess that ignoring a block need to be addressed case by case and depending on the statistical significance of the similarity matrix.

      thanks again.

      Comment

      • Brian Bushnell
        Super Moderator
        • Jan 2014
        • 2709

        #4
        BBMap, for example, assigns a positive score for a match, negative score for a mismatch, and 0 if either the reference or read are N. For nucleotide alignment similarity matrices are not used; those are for amino acid alignment.

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
          by SEQadmin2


          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

          The systematic characterization of the human proteome has
          ...
          07-20-2026, 11:48 AM
        • SEQadmin2
          Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
          by SEQadmin2



          Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
          ...
          07-09-2026, 11:10 AM
        • SEQadmin2
          Cancer Drug Resistance: The Lingering Barrier to Rising Survival
          by SEQadmin2



          Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

          There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
          07-08-2026, 05:17 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, Today, 12:17 PM
        0 responses
        9 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, Yesterday, 11:41 AM
        0 responses
        10 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-20-2026, 11:10 AM
        0 responses
        23 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-13-2026, 10:26 AM
        0 responses
        37 views
        0 reactions
        Last Post SEQadmin2  
        Working...