Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • tristan dubos
    Member
    • Dec 2015
    • 39

    #1

    question about mpileup

    Hi all ! My first post on this web site witch helped me a lot in many times before !!

    I got a question about mpileup, i took an exemple from : http://samtools.sourceforge.net/pileup.shtml :

    seq2 156 A 11 .$......+2AG.+2AG.+2AGGG <975;:<<<<<

    If i understood right the column 4 represent the the number of reads covering the site but in reality there is not 11 but 14 reads covering in reality. I mean if you are looking for mutation like INDEL on this position after filtering by covering you could miss some interesting mutation site isn't it ?

    My second question : i saw that .^] in read bases (column 5) and in all the cases it correspond of the start of the reads is it normal ? (there is no base before then why there is a quality ? )

    Thanks
  • Jessica_L
    Senior Member
    • Feb 2010
    • 117

    #2
    looking at the example you posted, the 4th column is the number of reads, in this case 11. I'm not sure what you mean about 14 reads in reality. I don't see that anywhere in the example, but I only skimmed. Maybe I missed something.

    To your second question: The mapping quality referred to by the ] character is not a base call quality score, it represents the mapping quality, which is a measure of how well the read aligns/matches the reference. I think this link explains it better than I can: https://www.biostars.org/p/8371/

    Comment

    • tristan dubos
      Member
      • Dec 2015
      • 39

      #3
      Thank you for the answer. I can explain with more details :
      Always in this case

      seq2 156 A 11 .$......+2AG.+2AG.+2AGGG <975;:<<<<<

      We have have : 9 . + 2 G = 11
      But the 3 reads with the insert of AG are not recorded, that is why i m saying there is 14 reads covering this position in reality . I m asking this question because if any variant caller use this column 4 for filtering the covering depth of the position it can t find INDELS ...

      Ok i understand better the meaning of the ^ thank you again !

      Comment

      • Jessica_L
        Senior Member
        • Feb 2010
        • 117

        #4
        My understanding of the base read format is that the insertions of AG exists on three of the 11 reads that have already been counted-- the insertion is between this reference position (156) and the next position in the reference sequence (157). The reads on which those insertions appear are counted here, in the example they are all matches to the reference-- dots. I read them as ".+2AG" which is a match to the reference, plus a 2bp insertion consisting of AG. Treating that read as "." and "+2AG" is double counting it.

        If you're saying that filtering on read depth would potentially cause you to toss out some indels, I'd agree with that statement. Filtering can always cost you the ability to see a novel variant, but it's the tradeoff for less noisy data. That's not the same as the variant caller not being able to find indels at all, though.

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
          by SEQadmin2



          CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

          Despite this, “CRISPR helped turn genome editing from a specialized technique into
          ...
          07-31-2026, 11:01 AM
        • SEQadmin2
          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
          by SEQadmin2


          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

          The systematic characterization of the human proteome has
          ...
          07-20-2026, 11:48 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, Yesterday, 10:35 AM
        0 responses
        7 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 08-06-2026, 07:41 AM
        0 responses
        25 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 08-03-2026, 10:13 AM
        0 responses
        45 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-31-2026, 02:55 AM
        0 responses
        48 views
        0 reactions
        Last Post SEQadmin2  
        Working...