Hiseq 2000 paired-end capture data analysis problem-too many variants!

lazyworm

Junior Member

Join Date: May 2010

Posts: 2
- Share
- Tweet
#1

Hiseq 2000 paired-end capture data analysis problem-too many variants!

08-11-2010, 09:44 AM

Hi,
We are trying to analysis a Hiseq 2000 paired-end whole exome capture sequencing data. The quality of the data is very good. We get an average depth of coverage around 120x. The fastq files looks perfect. We used bwa for paired end alignment. Picard to remove duplicates and Samtools for variant calling. The problem we have now is that there are too many SNV and indel variants from this data, around 140,000 SNVs and Indels after filtration (mapping quality>=45, read depth>=10 and standard varFilter in Samtools). I just wonder if somebody else on this board are doing similar data analysis. How many SNV and Indel you got? Can BWA and Samtools be used on Hiseq data? Or if there are some other software we should try? Any other information we should know about Hiseq data?

Thanks
Tags: None
Lee Sam

Member

Join Date: Oct 2008

Posts: 57
- Share
- Tweet
#2

08-11-2010, 10:03 AM

Have you done verification against dbSNP? Have you filtered down your candidates to just those within known exons after alignment (often times PE reads "splash over" into intronic regions where variation is likely more liberally tolerated)?

EDIT: I'm actually really curious to hear about how many reads and read length and how many lanes you ran the sample on. We just got our HiSeq2k installed last week and we're running our first samples though it. Details would be fantastic.

Last edited by Lee Sam; 08-11-2010, 12:29 PM.
Comment

Previous template Next

Essential Discoveries and Tools in Epitranscriptomics

by seqadmin

The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist...
- Channel: Articles
04-22-2024, 07:01 AM
Current Approaches to Protein Sequencing

by seqadmin

Proteins are often described as the workhorses of the cell, and identifying their sequences is key to understanding their role in biological processes and disease. Currently, the most common technique used to determine protein sequences is mass spectrometry. While still a valuable tool, mass spectrometry faces several limitations and requires a highly experienced scientist familiar with the equipment to operate it. Additionally, other proteomic methods, like affinity assays, are constrained...
- Channel: Articles
04-04-2024, 04:25 PM

Topics	Statistics	Last Post
A Close Examination at Probiotic-Related Bacteremia by seqadmin Started by seqadmin, 05-02-2024, 08:06 AM	0 responses 16 views 0 likes	Last Post by seqadmin 05-02-2024, 08:06 AM
Expanded Genetic Insights into Blood Pressure Regulation by seqadmin Started by seqadmin, 04-30-2024, 12:17 PM	0 responses 20 views 0 likes	Last Post by seqadmin 04-30-2024, 12:17 PM
The Role of Enhancers in Defining Cell Fate by seqadmin Started by seqadmin, 04-29-2024, 10:49 AM	0 responses 25 views 0 likes	Last Post by seqadmin 04-29-2024, 10:49 AM
Expanding the Horizons of Cellular Research with the Single Cell Atlas by seqadmin Started by seqadmin, 04-25-2024, 11:49 AM	0 responses 28 views 0 likes	Last Post by seqadmin 04-25-2024, 11:49 AM

Seqanswers Leaderboard Ad

Announcement

Hiseq 2000 paired-end capture data analysis problem-too many variants!

Comment

Latest Articles

ad_right_rmr

News