Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • thmourikis
    replied
    Thanks a lot Richard! I really appreciate that!

    Best,
    Thanos

    Leave a comment:


  • Richard Finney
    replied
    awk -v n=1 -v p=0 '/^\/\//{p++;if(((p%1000)==0)&&(p!=0)){close("out"n);n++;next}} {print > "out"n}' yourfilename.gbk

    splits at 1000 records.

    Leave a comment:


  • thmourikis
    replied
    Hi Peter and thank you for your immediate reply.

    I currently use Perl (not very experienced though). I guess I can try to alter Richard's awk command and implement it in a Perl script for renaming etc.

    Thank you once again.

    Leave a comment:


  • maubp
    replied
    Thanos - which scripting languages do you know? GenBank records end with a // line (which is what Richard's awk command exploits) so it is very simple to split up a file into sub-files named however you like using Perl, Python or Ruby.

    Leave a comment:


  • thmourikis
    replied
    Hi all,

    I have the same problem but I want to split the file every 1000 entries. My file has 500,000 records and I want 500 files of 1000 records each. Any suggestions?

    Thanks in advance.
    Thanos

    Leave a comment:


  • Richard Finney
    replied
    split genbank files using awk

    awk -v n=1 '/^\/\//{close("out"n);n++;next} {print > "out"n}' yourfilename.gbk

    Split yourfilename.gbk into multiple files by splitting at "//" (end of record) line.
    Last edited by Richard Finney; 03-20-2012, 09:51 AM.

    Leave a comment:


  • nickloman
    replied
    Come on Peter, I caught you napping again

    Leave a comment:


  • maubp
    replied
    If you want one file per record, try EMBOSS seqret and the -ossingle_outseq option.

    EDIT: That probably does the same as EMBOSS seqretsplit suggested by Nick while I was writing this.


    Do you just want to break it up into batches, say 10 records in each file? Or, do you have a particular order in mind (which could involve either sorting or random access).
    Last edited by maubp; 03-19-2012, 11:28 AM. Reason: Nick posted at same time

    Leave a comment:


  • nickloman
    replied
    I haven't tried it out but 'seqretsplit' from the EMBOSS package might do what you want. Otherwise it's a quick script in Bioperl or Biopython, e.g. in BioPython (untested)

    Run like python splitgbk.py < input.gbk

    Will create a file for each entry in the current directory.

    -- splitgbk.py

    Code:
    from Bio import SeqIO
    import sys
    
    for rec in SeqIO.parse(sys.stdin, "genbank"):
       SeqIO.write([rec], open(rec.id + ".gbk", "w"), "genbank")

    Leave a comment:


  • joscarhuguet
    started a topic splitting big genbank file

    splitting big genbank file

    I have a big gbk file containing multiple gbks, is there any simple way to split this big gbk into small gbks Thanks.

Latest Articles

Collapse

  • seqadmin
    Recent Advances in Sequencing Analysis Tools
    by seqadmin


    The sequencing world is rapidly changing due to declining costs, enhanced accuracies, and the advent of newer, cutting-edge instruments. Equally important to these developments are improvements in sequencing analysis, a process that converts vast amounts of raw data into a comprehensible and meaningful form. This complex task requires expertise and the right analysis tools. In this article, we highlight the progress and innovation in sequencing analysis by reviewing several of the...
    05-06-2024, 07:48 AM
  • seqadmin
    Essential Discoveries and Tools in Epitranscriptomics
    by seqadmin




    The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist...
    04-22-2024, 07:01 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by seqadmin, 05-14-2024, 07:03 AM
0 responses
23 views
0 likes
Last Post seqadmin  
Started by seqadmin, 05-10-2024, 06:35 AM
0 responses
44 views
0 likes
Last Post seqadmin  
Started by seqadmin, 05-09-2024, 02:46 PM
0 responses
58 views
0 likes
Last Post seqadmin  
Started by seqadmin, 05-07-2024, 06:57 AM
0 responses
44 views
0 likes
Last Post seqadmin  
Working...
X