how to simply extract the sequences of a gene list (~1000) in FASTA format from a sequence database (~400MB) in FASTA format generated by MAQ?
Seqanswers Leaderboard Ad
Collapse
Announcement
Collapse
No announcement yet.
X
-
Are you starting with a ~400MB FASTA file, containing ~1000 sequences, and you just want the list of sequence identifiers ("gene names")?
Try something like this at the Unix command line:
grep "^>" my_database.fasta
That string "^>" is a regular expression meaning look for any lines starting ("^") with the greater than symbol.
-
Originally posted by johnsequence View PostThanks for reply. I actually need the IDs (headers) and sequences in FASTA format.
How is your list of identifiers stored? e.g. a text file with one id per line?
I would suggest you write a simple script, e.g. using Perl (perhaps with BioPerl) or Python (perhaps with Biopython), or your preferred script language.
Or, if you are happier just working at the command line, you can probably do this with EMBOSS seqret.
Comment
-
You can use a couple of the utilities in the BLAST package from NCBI. Take your large FASTA file and create a BLAST database from it using formatdb. Then retrieve just the sequences you want from the BLASTdb using the fastacmd tool.
Code:%> formatdb -i <your.FASTA.file> -p F -n <your.blast.db> %> fastacmd -d <your.blast.db> -i <your.ID.file> > <output.file>
Comment
Latest Articles
Collapse
-
by seqadmin
The sequencing world is rapidly changing due to declining costs, enhanced accuracies, and the advent of newer, cutting-edge instruments. Equally important to these developments are improvements in sequencing analysis, a process that converts vast amounts of raw data into a comprehensible and meaningful form. This complex task requires expertise and the right analysis tools. In this article, we highlight the progress and innovation in sequencing analysis by reviewing several of the...-
Channel: Articles
05-06-2024, 07:48 AM -
ad_right_rmr
Collapse
News
Collapse
Topics | Statistics | Last Post | ||
---|---|---|---|---|
Started by seqadmin, Today, 10:28 AM
|
0 responses
8 views
0 likes
|
Last Post
by seqadmin
Today, 10:28 AM
|
||
Started by seqadmin, Today, 07:35 AM
|
0 responses
11 views
0 likes
|
Last Post
by seqadmin
Today, 07:35 AM
|
||
Started by seqadmin, Yesterday, 02:06 PM
|
0 responses
8 views
0 likes
|
Last Post
by seqadmin
Yesterday, 02:06 PM
|
||
Started by seqadmin, 05-14-2024, 07:03 AM
|
0 responses
28 views
0 likes
|
Last Post
by seqadmin
05-14-2024, 07:03 AM
|
Comment