Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • GFF editing-Perl, awk

    I would like to get the help of computer experts, in editing my GFF file

    Code:
    ##fsdfsd
    ##sdd-ver gen 5.5.7
    Xm Gen CDS     1       148     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     149     193     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     194     279     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     280     412     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     413     499     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     500     702     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen extracted region        1       148     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="437225 <- 437372";ID=Bm
    Xm Gen extracted region        149     193     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="436969 <- 437013";ID=Bm
    Xm Gen extracted region        194     279     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="435418 <- 435503";ID=Bm
    Xm Gen extracted region        280     412     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="435209 <- 435341";ID=Bm
    Xm Gen extracted region        413     499     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="434376 <- 434462";ID=Bm
    Xm Gen extracted region        500     702     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="434084 <- 434286";ID=Bm
    Replace (Xm Gen CDS) rows with values present in (Xm Gen extracted region) rows. i.e first row ($4 column:1 is replaced with 437225 value and $5 column:148 is replaced with 437372, in row 2 ($4 column:149 replaced with 436969, $5 column:193 replaced with 437013 and so on) and print output like below

    Code:
    ##gff-version 2
    ##source-version geneious 5.5.7
    Xm Gen CDS     437225       437372     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     436969     437013     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     435418     435503     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     435209     435341     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     434376     434462     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen CDS     434084     434286     .       +       .       Name=;created by=User;modified by=User;ID=Bm
    Xm Gen extracted region        1       148     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="437225 <- 437372";ID=Bm
    Xm Gen extracted region        149     193     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="436969 <- 437013";ID=Bm
    Xm Gen extracted region        194     279     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="435418 <- 435503";ID=Bm
    Xm Gen extracted region        280     412     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="435209 <- 435341";ID=Bm
    Xm Gen extracted region        413     499     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="434376 <- 434462";ID=Bm
    Xm Gen extracted region        500     702     .       +       .       Name=Extracted region from gi|371442828|gb|JH557032.1|;Extracted interval="434084 <- 434286";ID=Bm

    I have used grep & sed commands as below.

    head -2 file
    join <(grep "Xm Gen CDS" file | cat -n) \
    <(grep "Xm Gen extracted region" file | cat -n) | \
    sed 's/^[0-9]* //;s/CDS [0-9]*\s[0-9]*\(\s.*interval="\([0-9]*\)\s<-\s\([0-9]*\)\)/CDS\t\2\t\3\t\1/;s/ Xm Gen extracted.*//'
    grep "Xm Gen extracted region" file


    But I would like convert this to shell script like get input from the user and give as desired by the user.
    Any perl script is appreciated.

Latest Articles

Collapse

  • seqadmin
    Choosing Between NGS and qPCR
    by seqadmin



    Next-generation sequencing (NGS) and quantitative polymerase chain reaction (qPCR) are essential techniques for investigating the genome, transcriptome, and epigenome. In many cases, choosing the appropriate technique is straightforward, but in others, it can be more challenging to determine the most effective option. A simple distinction is that smaller, more focused projects are typically better suited for qPCR, while larger, more complex datasets benefit from NGS. However,...
    10-18-2024, 07:11 AM
  • seqadmin
    Non-Coding RNA Research and Technologies
    by seqadmin




    Non-coding RNAs (ncRNAs) do not code for proteins but play important roles in numerous cellular processes including gene silencing, developmental pathways, and more. There are numerous types including microRNA (miRNA), long ncRNA (lncRNA), circular RNA (circRNA), and more. In this article, we discuss innovative ncRNA research and explore recent technological advancements that improve the study of ncRNAs.

    Nobel Prize for MicroRNA Discovery
    This week,...
    10-07-2024, 08:07 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by seqadmin, Yesterday, 06:09 AM
0 responses
10 views
0 likes
Last Post seqadmin  
Started by seqadmin, 10-30-2024, 05:31 AM
0 responses
13 views
0 likes
Last Post seqadmin  
Started by seqadmin, 10-24-2024, 06:58 AM
0 responses
22 views
0 likes
Last Post seqadmin  
Started by seqadmin, 10-23-2024, 08:43 AM
0 responses
52 views
0 likes
Last Post seqadmin  
Working...
X