Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • Filter variants by DP4 (total)?

    Greetings. I have found that in my particular data set that the most sensitive factor for true variants versus false positives is actually the DP4 (the total number) rather than the QUAL score or depth (DP). Is there a program or little script to filter by the total DP4? Thanks!

  • #2
    filtering a VCF with javascript

    I wrote a tool named vcffilterjs to filter a VCF using a javascript expression: See: https://github.com/lindenb/jvarkit#-...ascript-rhino-

    Example:
    wth
    Code:
    function accept()
    	{
    	var DP4=variant.getAttribute("DP4");
    	if(DP4==null || DP4.size()!=4) return false;
    	return  DP4.get(0)+DP4.get(1)+DP4.get(2)+DP4.get(3)<10;
    	}
    accept();
    Code:
    $ curl "https://raw.github.com/CBMi-BiG/snpEff/master/tests/vcf_homo.vcf" | \
    java -jar dist/vcffilterjs.jar SCRIPT_FILE=filter.js 2> /dev/null
    Code:
    (...)
    #CHROM	POS	ID	REF	ALT	QUAL	FILTER	INFO	FORMAT	s_1_ACAGTGA_sort.bam
    Y	3720217	.	A	G	8.65	.	AC=2;AF1=1;AN=2;CI95=0.5,1;DP=2;DP4=0,0,0,1;FQ=-30;G3=4.415e-15,5.291e-06,1;MQ=38;SF=5	GT:GQ:PL	0/0:61:60,6,0
    Y	3721230	.	C	G	21.80	.	AC=2;AF1=1;AN=2;CI95=0.5,1;DP=2;DP4=0,0,0,2;FQ=-33;G3=1.456e-17,8.564e-07,1;MQ=29;SF=3	GT:GQ:PL	1/1:61:60,6,0
    Y	3744605	.	C	A	3.98	.	AC=2;AF1=1;AN=2;CI95=0.5,1;DP=2;DP4=0,0,0,2;FQ=-33;G3=1.468e-15,8.599e-07,1;MQ=19;SF=2	GT:GQ:PL	0/0:61:60,6,0
    Y	9945223	.	ATTT	ATTTT	19.80	.	AC=4;AF1=1;AN=4;CI95=0.5,1;DP=2;DP4=0,0,0,2;FQ=-40.5;G3=2.906e-18,8.564e-07,1;INDEL;MQ=45;SF=0,2	GT:GQ:PL	1/1:61:60,6,0

    Comment


    • #3
      @lindenb looks like just what I need, I'll let you know how it goes!

      Comment

      Latest Articles

      Collapse

      • seqadmin
        Non-Coding RNA Research and Technologies
        by seqadmin




        Non-coding RNAs (ncRNAs) do not code for proteins but play important roles in numerous cellular processes including gene silencing, developmental pathways, and more. There are numerous types including microRNA (miRNA), long ncRNA (lncRNA), circular RNA (circRNA), and more. In this article, we discuss innovative ncRNA research and explore recent technological advancements that improve the study of ncRNAs.

        Nobel Prize for MicroRNA Discovery
        This week,...
        10-07-2024, 08:07 AM
      • seqadmin
        Recent Developments in Metagenomics
        by seqadmin





        Metagenomics has improved the way researchers study microorganisms across diverse environments. Historically, studying microorganisms relied on culturing them in the lab, a method that limits the investigation of many species since most are unculturable1. Metagenomics overcomes these issues by allowing the study of microorganisms regardless of their ability to be cultured or the environments they inhabit. Over time, the field has evolved, especially with the advent...
        09-23-2024, 06:35 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by seqadmin, Yesterday, 02:44 PM
      0 responses
      7 views
      0 likes
      Last Post seqadmin  
      Started by seqadmin, 10-11-2024, 06:55 AM
      0 responses
      14 views
      0 likes
      Last Post seqadmin  
      Started by seqadmin, 10-02-2024, 04:51 AM
      0 responses
      110 views
      0 likes
      Last Post seqadmin  
      Started by seqadmin, 10-01-2024, 07:10 AM
      0 responses
      117 views
      0 likes
      Last Post seqadmin  
      Working...
      X