I have a question about syndip dataset : https://github.com/lh3/CHM-eval . I'm struggling to find the syndip vcf.
In the release ( https://github.com/lh3/CHM-eval/releases ), we have a file named : rep2.37.broad.hc.raw.vcf.gz, that i don't know what it is. And we have a file named CHM-evalkit-20180222.tar wich contain full.37m.vcf and other files ( bed, eval ...). So i did my search and according to this file: they mentionned that full.37m.vcf is the truth dataset. ( https://www.biorxiv.org/content/bior...1/456103-1.pdf Page 16).
The problem is that the file rep2.37.broad.hc.raw.vcf.gz contain variants with MQ, DP, GQ ... that i need to extract. But the full.37m.vcf doesn't contain this information.. ( just Chrom pos ref alt and QUAL.)
So i tried to intersect rep2.37.broad.hc.raw.vcf.gz with full.37m.vcf and take the variant that present in two files, with the DP MQ GQ in rep2.37.broad.hc.raw.vcf.gz. Is that okay ? Since I don't know what is rep2.37.broad.hc.raw.vcf.gz.
And i also noticed that the QUAL in the full.37m.vcf is always 30 .. Is it normal ? Thank's
Seqanswers Leaderboard Ad
Collapse
Announcement
Collapse
No announcement yet.
X
Latest Articles
Collapse
-
by seqadmin
The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist...-
Channel: Articles
04-22-2024, 07:01 AM -
-
by seqadmin
Proteins are often described as the workhorses of the cell, and identifying their sequences is key to understanding their role in biological processes and disease. Currently, the most common technique used to determine protein sequences is mass spectrometry. While still a valuable tool, mass spectrometry faces several limitations and requires a highly experienced scientist familiar with the equipment to operate it. Additionally, other proteomic methods, like affinity assays, are constrained...-
Channel: Articles
04-04-2024, 04:25 PM -
ad_right_rmr
Collapse
News
Collapse
Topics | Statistics | Last Post | ||
---|---|---|---|---|
Started by seqadmin, Today, 08:47 AM
|
0 responses
12 views
0 likes
|
Last Post
by seqadmin
Today, 08:47 AM
|
||
Started by seqadmin, 04-11-2024, 12:08 PM
|
0 responses
60 views
0 likes
|
Last Post
by seqadmin
04-11-2024, 12:08 PM
|
||
Started by seqadmin, 04-10-2024, 10:19 PM
|
0 responses
59 views
0 likes
|
Last Post
by seqadmin
04-10-2024, 10:19 PM
|
||
Started by seqadmin, 04-10-2024, 09:21 AM
|
0 responses
54 views
0 likes
|
Last Post
by seqadmin
04-10-2024, 09:21 AM
|