Hi guys I have a problem that I cannot seem to get around regarding Paired end vs Single end data. I have a data set of ~1000 genes with an average length of 900bp per gene.
I wish to compare the various read lengths and their affect on de novo assembly of multiple genes in bot SE and PE data.
I have simulated various length data sets using the ART NGS simulator. The command that I used for the PE data was :
./art_illumina --paired -i genes.fa -l 85 -f 200 -m 900 -s 150 -o outfile.fa
where reads of length 85 will be generated with fold coverage of 200 and mean DNA fragment size of 900 with a standard deviation of 150.
I have used SOAPdenovo to assemble all data however, more of the initial genes are recovered using SE data. I would have expected PE data to produce better results, it is quite poor in comparison to SE.
I think I may be doing something wrong in the simulation as the final assembled contigs only ever get to a max size of about 400bp for PE.
Does anyone know what the problem might be?
Thank you!
I wish to compare the various read lengths and their affect on de novo assembly of multiple genes in bot SE and PE data.
I have simulated various length data sets using the ART NGS simulator. The command that I used for the PE data was :
./art_illumina --paired -i genes.fa -l 85 -f 200 -m 900 -s 150 -o outfile.fa
where reads of length 85 will be generated with fold coverage of 200 and mean DNA fragment size of 900 with a standard deviation of 150.
I have used SOAPdenovo to assemble all data however, more of the initial genes are recovered using SE data. I would have expected PE data to produce better results, it is quite poor in comparison to SE.
I think I may be doing something wrong in the simulation as the final assembled contigs only ever get to a max size of about 400bp for PE.
Does anyone know what the problem might be?
Thank you!
Comment