Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • CIGAR strings and 'N' symbols

    Just wondering if anybody can help me to understand the following example:


    ref: AC GTACGT
    r1 : ACCGTACGT
    r2 : AC......T

    Will the CIGARs be:

    2M1I6M
    2M6N1M

    OR will it be:
    2M1I6M
    2M5N1M



    In other words, if there are alignment positions with possible insertions in the reference, will the skipped positions (Ns) take into account these possible insertions?


    Also, how do mappers determine segments with skipped N positions?

    Thanks!

  • #2
    Originally posted by pieffe View Post
    Just wondering if anybody can help me to understand the following example:


    ref: AC GTACGT
    r1 : ACCGTACGT
    r2 : AC......T

    Will the CIGARs be:

    2M1I6M
    2M6N1M

    OR will it be:
    2M1I6M
    2M5N1M



    In other words, if there are alignment positions with possible insertions in the reference, will the skipped positions (Ns) take into account these possible insertions?


    Also, how do mappers determine segments with skipped N positions?

    Thanks!
    Code:
    ref: AC GTACGT
    r1 :  ACCGTACGT
    r2 :  AC......T
    In my opinion r2 is not valid. The "." is meant to represent skipping a base within the read that correspond to a reference base, not an insertion. The inserted base should be represented with an "N" if it the inserted base is unknown. Therefore r2 should be "ACN.....T" and the cigar would be 2M1I5N1M.

    You may want to send an email to the samtools-help mailing list for clarification.

    Comment


    • #3
      Thank you for your prompt reply. I think what you says makes sense. However, in the SAM specifications, there is an example where the alignment is spliced and the read does not have a 'N' symbol at all, and yet the cigar has 32N in it.

      I tried to use samtools to dig into this, but got even more confused. I guess I am going to ask to the authors of samtools for clarifications. I will post what I find out.

      Thanks again!

      Comment


      • #4
        Originally posted by pieffe View Post
        Thank you for your prompt reply. I think what you says makes sense. However, in the SAM specifications, there is an example where the alignment is spliced and the read does not have a 'N' symbol at all, and yet the cigar has 32N in it.

        I tried to use samtools to dig into this, but got even more confused. I guess I am going to ask to the authors of samtools for clarifications. I will post what I find out.

        Thanks again!
        The "N" symbol in the read indicates a missing base, the "N" symbol in the cigar indicates a skipped reference base. The splice alignment has no "N" bases but skips reference bases, hence the "N" in the cigar string.
        Last edited by nilshomer; 03-11-2010, 12:53 PM.

        Comment


        • #5
          Hi Nils,

          I think I was too quick to email you a reply. I re-read your answer and you probably are right. The example in the SAM paper won't need N symbols in the read, since there are no insertions in reference.
          I guess that, if I am not misunderstanding, whenever a read has a I symbol in the cigar, other reads will either have P symbols in the cigar or they will have a N symbol in the read and a I symbol in the cigar.

          Thanks again,
          PF

          Comment

          Latest Articles

          Collapse

          • seqadmin
            Recent Advances in Sequencing Analysis Tools
            by seqadmin


            The sequencing world is rapidly changing due to declining costs, enhanced accuracies, and the advent of newer, cutting-edge instruments. Equally important to these developments are improvements in sequencing analysis, a process that converts vast amounts of raw data into a comprehensible and meaningful form. This complex task requires expertise and the right analysis tools. In this article, we highlight the progress and innovation in sequencing analysis by reviewing several of the...
            05-06-2024, 07:48 AM
          • seqadmin
            Essential Discoveries and Tools in Epitranscriptomics
            by seqadmin




            The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist...
            04-22-2024, 07:01 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by seqadmin, Yesterday, 06:57 AM
          0 responses
          11 views
          0 likes
          Last Post seqadmin  
          Started by seqadmin, 05-06-2024, 07:17 AM
          0 responses
          16 views
          0 likes
          Last Post seqadmin  
          Started by seqadmin, 05-02-2024, 08:06 AM
          0 responses
          19 views
          0 likes
          Last Post seqadmin  
          Started by seqadmin, 04-30-2024, 12:17 PM
          0 responses
          24 views
          0 likes
          Last Post seqadmin  
          Working...
          X