Hi all,
I'm begginer in the next generation sequencing and recently I got my first data. I have some issues with them because there is just one study which had similar approach so far. Briefly I'm working with phylogeny of plants which includes polyploidy events. We have few markers and in some species (e.g. tetraploids) there may be few copies of one gene (marker). To obtain all possible copies we used 454 and got around 100 sequences per species. I tried to used BAPS (bayesian clustering program used for population studies mostly) to decide whether there is one copy or two of each gene. Although it is quite labourious basically it works and I couldn't figure out anything better so far. But then there is one thing which I would like to automate a little bit and make it more objective:
Even when I know that there is one or two copies in my 100 reads there are still sequences which are different only by 1 or 2 bp, usually it's easily visible that it is artifact especially when it is present only once in the dataset. I wonder wheter there is any software that could collapse sequences for me and exlude this artifacts? Let's say that I would like to have sequences that are present with minimum 20% in the dataset. The truth that in most cases in the end I need just 2 sequences out of 100...
Most programs with collapsing option just collapse everything which is identical to one haplotype and doesn't include any onformation about presence in the data or similarities. And they leave some sequences with more than one difference which makes things even more complicated.
Is there any program that I could use for my analysis or do I have to do this by hand? If anybody has any idea I would be grateful!
Thank's a lot in advance for any help!
I'm begginer in the next generation sequencing and recently I got my first data. I have some issues with them because there is just one study which had similar approach so far. Briefly I'm working with phylogeny of plants which includes polyploidy events. We have few markers and in some species (e.g. tetraploids) there may be few copies of one gene (marker). To obtain all possible copies we used 454 and got around 100 sequences per species. I tried to used BAPS (bayesian clustering program used for population studies mostly) to decide whether there is one copy or two of each gene. Although it is quite labourious basically it works and I couldn't figure out anything better so far. But then there is one thing which I would like to automate a little bit and make it more objective:
Even when I know that there is one or two copies in my 100 reads there are still sequences which are different only by 1 or 2 bp, usually it's easily visible that it is artifact especially when it is present only once in the dataset. I wonder wheter there is any software that could collapse sequences for me and exlude this artifacts? Let's say that I would like to have sequences that are present with minimum 20% in the dataset. The truth that in most cases in the end I need just 2 sequences out of 100...
Most programs with collapsing option just collapse everything which is identical to one haplotype and doesn't include any onformation about presence in the data or similarities. And they leave some sequences with more than one difference which makes things even more complicated.
Is there any program that I could use for my analysis or do I have to do this by hand? If anybody has any idea I would be grateful!
Thank's a lot in advance for any help!
Comment