Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • adrianl
    Junior Member
    • Sep 2012
    • 5

    #1

    Varscan frequency question

    I've been trying out varscan for indel calling recently but I have an issue with the output, rather the information in the output... The field named Freq doesnt seem to agree with the different Reads fields in my output. Here is a couple of examples (I've formatted the output so that it is a bit more reader-friendly):

    No code has to be inserted here.I'd have thought the frequency would be Reads2/(Reads1+Reads2), the reads supporting the variant over all the reads. This is not the case for most of my entries. They are all ball-park-close but few are spot on (I've included an entry that seems correct, the last one).

    For the first row of my example: 13/14 = 0.9286 != 0.8667
    However, what is 0.8667 is this: 13/15.

    Are the reads wrong? Are the frequencies wrong? Are they calculated with different views on what is a supporting read? Something is going on that I cant seem to be able to figure out, or find searching forums etc.

    I'm using VarScan v2.3.2, mpileup2indel with a few parameters:
    --min-var-freq 0.001
    --min-avg-qual 30
    --min-reads2 10
    --strand-filter 1
    --p-value 0.9

    I'm sure theres a simple answer, does anyone have it?
    Cheers
    //Adrian
  • Graham Etherington
    Member
    • Apr 2010
    • 22

    #2
    Hi Adrian,
    I'm also new to VarScan. I wonder if these figures are affected by the flag:
    "--min-avg-qual Minimum base quality at a position to count a read [15]"
    The Freq of 94.12% given in your second example can be derived from 32/34, so perhaps there are 34 reads over that position, but only 33 that have base qualities of 15 or above. You could eyeball the pileup and see if this is true or try using the
    VarScan readcounts tools with parameters:
    --min-coverage 0
    --min-base-qual 0

    I'd be interested to see how you get on.
    Cheers,
    Graham

    Comment

    • adrianl
      Junior Member
      • Sep 2012
      • 5

      #3
      Thank you Graham, for your input!

      I ran readcounts with parameters:
      --min-coverage 0
      --min-base-qual 30 (Since this is what I ran mpileup2idel with)

      And I think I found the answer to my question.

      No code has to be inserted here.The entries with "deviating" frequencies had additional variants, in the example cases one read each. Adding this read to the total gives us the same frequency as printed in the output file (13/15 and 32/34). Some of my positions had several additional variants in the pileup file making the total even more "wrong" when added together just from the mpileup2indel output file. At least now that I understand it there is no problem anymore, I can trust these values a bit more.

      So in the end the answer was quite simple, I guess.

      Thanks again
      //Adrian


      Edit: I may have broken the forum boundaries with my huge table...

      Comment

      • dkoboldt
        Member
        • Mar 2009
        • 62

        #4
        Hello Adrian and Graham,

        Thank you for bringing this up and looking into the issue. Adrian, would you mind sending the raw pileup for the two positions that you mentioned?

        Feel free to use VarScan's support forum if you have other issues or questions.

        Note that correctly counting reads (supporting or refuting) for indels is difficult using the first-pass alignments in a BAM file. Optimally, you would use VarScan to discover the indels, and then use realignment (GATK) or indel haplotype remapping (DINDEL) to obtain more accurate read counts and variant allele frequencies.

        Yours,

        Dan Koboldt

        Comment

        • Yamol
          Junior Member
          • Dec 2014
          • 5

          #5
          Originally posted by adrianl View Post
          Thank you Graham, for your input!

          I ran readcounts with parameters:
          --min-coverage 0
          --min-base-qual 30 (Since this is what I ran mpileup2idel with)

          And I think I found the answer to my question.

          No code has to be inserted here.The entries with "deviating" frequencies had additional variants, in the example cases one read each. Adding this read to the total gives us the same frequency as printed in the output file (13/15 and 32/34). Some of my positions had several additional variants in the pileup file making the total even more "wrong" when added together just from the mpileup2indel output file. At least now that I understand it there is no problem anymore, I can trust these values a bit more.

          So in the end the answer was quite simple, I guess.

          Thanks again
          //Adrian


          Edit: I may have broken the forum boundaries with my huge table...
          Can you tell how this result come out? With which program and parameters?

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
            by SEQadmin2



            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

            Despite this, “CRISPR helped turn genome editing from a specialized technique into
            ...
            07-31-2026, 11:01 AM
          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, Yesterday, 10:13 AM
          0 responses
          14 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-31-2026, 02:55 AM
          0 responses
          28 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-24-2026, 12:17 PM
          0 responses
          21 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-23-2026, 11:41 AM
          0 responses
          20 views
          0 reactions
          Last Post SEQadmin2  
          Working...