Dear all
I was wondering if anyone could help me in obtaining a concatenamer of sequences in the way showed below.
I have several multifasta files relative a genes sequences (ABC, GHJ…) in different organisms (>182680572, >749299147…)
Gene ABC
>182680572
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATA
>749299147
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCA
>584117620
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCTCACGACCTAA
>985743106
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTCCCGCCAGAAGATGTTCATACTTGGTTAAGACCTTTACAAGCCGACCAACGTGGTGACAGTGTCGTCCTTTACGCACCGAATCCCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGACGTCTTCGGGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGG
GENE GHJ
>182680572
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATA
>749299147
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCA
>584117620
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCTCACGACCTAA
>985743106
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTCCCGCCAGAAGATGTTCATACTTGGTTAAGACCTTTACAAGCCGACCAACGTGGTGACAGTGTCGTCCTTTACGCACCGAATCCCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGACGTCTTCGGGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGG
Then I want to obtain for each organism a concatened sequence of the genes in the same order for each organisms like below:
>182680572
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATAATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATA
>749299147
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCAATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCA
>584117620
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCTCACGACCTAAATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCTCACGACCTAA
>985743106
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTCCCGCCAGAAGATGTTCATACTTGGTTAAGACCTTTACAAGCCGACCAACGTGGTGACAGTGTCGTCCTTTACGCACCGAATCCCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGACGTCTTCGGGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTCCCGCCAGAAGATGTTCATACTTGGTTAAGACCTTTACAAGCCGACCAACGTGGTGACAGTGTCGTCCTTTACGCACCGAATCCCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGACGTCTTCGGGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGG
Does anyone knows how to do it with a perl/python script or bioinformatic software?
I was wondering if anyone could help me in obtaining a concatenamer of sequences in the way showed below.
I have several multifasta files relative a genes sequences (ABC, GHJ…) in different organisms (>182680572, >749299147…)
Gene ABC
>182680572
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATA
>749299147
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCA
>584117620
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCTCACGACCTAA
>985743106
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTCCCGCCAGAAGATGTTCATACTTGGTTAAGACCTTTACAAGCCGACCAACGTGGTGACAGTGTCGTCCTTTACGCACCGAATCCCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGACGTCTTCGGGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGG
GENE GHJ
>182680572
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATA
>749299147
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCA
>584117620
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCTCACGACCTAA
>985743106
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTCCCGCCAGAAGATGTTCATACTTGGTTAAGACCTTTACAAGCCGACCAACGTGGTGACAGTGTCGTCCTTTACGCACCGAATCCCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGACGTCTTCGGGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGG
Then I want to obtain for each organism a concatened sequence of the genes in the same order for each organisms like below:
>182680572
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATAATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATA
>749299147
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCAATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCA
>584117620
ATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCTCACGACCTAAATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTTCCACCAGAAGATGTTCATACTTGGTTGAGACCTTTACAAGCTGACCAACGCGGTGACAGTGTCATCCTTTACGCACCCAATACCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGGCGTCTTCGAGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGCTCACGACCTAA
>985743106
ATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTCCCGCCAGAAGATGTTCATACTTGGTTAAGACCTTTACAAGCCGACCAACGTGGTGACAGTGTCGTCCTTTACGCACCGAATCCCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGACGTCTTCGGGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGGATGACAACATTGATGGAATCTTGGTCCCGTTGCCTGGAACGTCTTGAAACTGAATTCCCGCCAGAAGATGTTCATACTTGGTTAAGACCTTTACAAGCCGACCAACGTGGTGACAGTGTCGTCCTTTACGCACCGAATCCCTTTATCATTGAACTAGTAGAAGAGCGATACTTAGGACGTCTTCGGGAATTGTTATCCTATTTTTCAGGAATACGTGAAGTAGTCCTTGCAATTGG
Does anyone knows how to do it with a perl/python script or bioinformatic software?
Comment