Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • pieffe
    replied
    Hi Nils,

    I think I was too quick to email you a reply. I re-read your answer and you probably are right. The example in the SAM paper won't need N symbols in the read, since there are no insertions in reference.
    I guess that, if I am not misunderstanding, whenever a read has a I symbol in the cigar, other reads will either have P symbols in the cigar or they will have a N symbol in the read and a I symbol in the cigar.

    Thanks again,
    PF

    Leave a comment:


  • nilshomer
    replied
    Originally posted by pieffe View Post
    Thank you for your prompt reply. I think what you says makes sense. However, in the SAM specifications, there is an example where the alignment is spliced and the read does not have a 'N' symbol at all, and yet the cigar has 32N in it.

    I tried to use samtools to dig into this, but got even more confused. I guess I am going to ask to the authors of samtools for clarifications. I will post what I find out.

    Thanks again!
    The "N" symbol in the read indicates a missing base, the "N" symbol in the cigar indicates a skipped reference base. The splice alignment has no "N" bases but skips reference bases, hence the "N" in the cigar string.
    Last edited by nilshomer; 03-11-2010, 12:53 PM.

    Leave a comment:


  • pieffe
    replied
    Thank you for your prompt reply. I think what you says makes sense. However, in the SAM specifications, there is an example where the alignment is spliced and the read does not have a 'N' symbol at all, and yet the cigar has 32N in it.

    I tried to use samtools to dig into this, but got even more confused. I guess I am going to ask to the authors of samtools for clarifications. I will post what I find out.

    Thanks again!

    Leave a comment:


  • nilshomer
    replied
    Originally posted by pieffe View Post
    Just wondering if anybody can help me to understand the following example:


    ref: AC GTACGT
    r1 : ACCGTACGT
    r2 : AC......T

    Will the CIGARs be:

    2M1I6M
    2M6N1M

    OR will it be:
    2M1I6M
    2M5N1M



    In other words, if there are alignment positions with possible insertions in the reference, will the skipped positions (Ns) take into account these possible insertions?


    Also, how do mappers determine segments with skipped N positions?

    Thanks!
    Code:
    ref: AC GTACGT
    r1 :  ACCGTACGT
    r2 :  AC......T
    In my opinion r2 is not valid. The "." is meant to represent skipping a base within the read that correspond to a reference base, not an insertion. The inserted base should be represented with an "N" if it the inserted base is unknown. Therefore r2 should be "ACN.....T" and the cigar would be 2M1I5N1M.

    You may want to send an email to the samtools-help mailing list for clarification.

    Leave a comment:


  • pieffe
    started a topic CIGAR strings and 'N' symbols

    CIGAR strings and 'N' symbols

    Just wondering if anybody can help me to understand the following example:


    ref: AC GTACGT
    r1 : ACCGTACGT
    r2 : AC......T

    Will the CIGARs be:

    2M1I6M
    2M6N1M

    OR will it be:
    2M1I6M
    2M5N1M



    In other words, if there are alignment positions with possible insertions in the reference, will the skipped positions (Ns) take into account these possible insertions?


    Also, how do mappers determine segments with skipped N positions?

    Thanks!

Latest Articles

Collapse

  • seqadmin
    Best Practices for Single-Cell Sequencing Analysis
    by seqadmin



    While isolating and preparing single cells for sequencing was historically the bottleneck, recent technological advancements have shifted the challenge to data analysis. This highlights the rapidly evolving nature of single-cell sequencing. The inherent complexity of single-cell analysis has intensified with the surge in data volume and the incorporation of diverse and more complex datasets. This article explores the challenges in analysis, examines common pitfalls, offers...
    06-06-2024, 07:15 AM
  • seqadmin
    Latest Developments in Precision Medicine
    by seqadmin



    Technological advances have led to drastic improvements in the field of precision medicine, enabling more personalized approaches to treatment. This article explores four leading groups that are overcoming many of the challenges of genomic profiling and precision medicine through their innovative platforms and technologies.

    Somatic Genomics
    “We have such a tremendous amount of genetic diversity that exists within each of us, and not just between us as individuals,”...
    05-24-2024, 01:16 PM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by seqadmin, 06-14-2024, 07:24 AM
0 responses
12 views
0 likes
Last Post seqadmin  
Started by seqadmin, 06-13-2024, 08:58 AM
0 responses
14 views
0 likes
Last Post seqadmin  
Started by seqadmin, 06-12-2024, 02:20 PM
0 responses
17 views
0 likes
Last Post seqadmin  
Started by seqadmin, 06-07-2024, 06:58 AM
0 responses
186 views
0 likes
Last Post seqadmin  
Working...
X