Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • Raw counts of 12 column bed file (against multiple BAM)

    Hello Everyone,

    I am sharing a perl script to get raw read counts from mulitple BAM files, for a given bed file. The script requires bedtools and samtools to be preinstalled in the system.

    I hope it is useful for users. An alternative solution exists in python
    (http://seqanswers.com/forums/showthread.php?t=18182)
    Attached Files

  • #2
    thanks so much for ur script.
    But you didn't give the usage info.

    I hv a .gtf file download from EMBL for D.mel now, should i convert it to .bed in advance?
    and the usage is?
    perl bed_count.pl annotated.bed dir_of_my_bamfiles tmpcount/ my.exon my.count

    Comment


    • #3
      Dear Jason,

      A much quicker solution is to

      1. Convert GTF to BED. A nice script is found here http://ea-utils.googlecode.com/svn/t...lipper/gtf2bed

      2. Convert the BED file to exon coordinates (BEDTools package).
      bed12toBed6 -i filename.bed > filename.exons

      3. Get read counts from multiple BAM files (BEDTools package)
      multiBamCov -bams aln.1.bam aln.2.bam -bed filename.exons > filename.counts
      (Remember to build index of each BAM file using samtools index, else multiBamCov wont work)

      4. Load the count file in R environment.

      filename <- read.delim("filename.counts", header=FALSE,sep = "\t", stringsAsFactors =FALSE)

      5. Sum the exons into gene counts using apply function in R
      input <- filename[,7:ncol(filename)]
      names <- filename$V4

      filename.gene.count<-apply(input,2,function(x){tapply(x,names,function(y){sum(y)})})

      6. Write the count file as a text file
      write.table(filename.gene.count, file = "filename.gene.count", quote = FALSE, row.names=FALSE, col.names=FALSE)

      The order of the counts for samples will be the same order in which you give the BAM files. My previous script is easier to implement but little slower. If you open the script in a text editor you can give the name of the directory of BAM files and the BED file. Hope this helps.

      Comment


      • #4
        Dear swaraj,

        It helps me a lot. Thanks again^^
        But i hv encountered a now problem now when adopting the new solution. I hv got 23017 exons in total while 14797 genes are mapped to the genome by Tophat & Cufflinks. and 14797 is a proper count for fruit fly. I notice that ur solution use exon coordinates(FBtrxxxx instead of FBgnxxxxx).
        I am confused now. what else do i need to do to process the matrix of raw counts i just got before importing it into R packages like edgeR for DEG analysis?

        Comment


        • #5
          In step 2 I coverted transcript BED file into exon file. Here the name of each exon is the transcript name which means transcript name is reduddant. In step 5 this same file is converted back into transcript file using R. All reads in exons with same transcript name are summed up. I would not go into using exon names. Just keep the same name (gene name or trasncript name) for each exon of a feature. This makes it quicker in R to sum the counts of all exons.

          Comment


          • #6
            ya, i know and i do exactly as ur pipeline. might be i mixed the concept of exon and transcript. i thought they were the same thing.(both named "FBtrxxx")
            So, this is not the reason why i got almost twice counts of genes as i expected.
            Could you figure out any other reason? Thank you!

            Comment


            • #7
              Or,
              what do i need to do to get raw counts for each gene?
              Thank you!

              Comment


              • #8
                Originally posted by jason_ARGONAUTE View Post
                Or,
                what do i need to do to get raw counts for each gene?
                Thank you!
                If you have a GTF file, just use htseq-count.

                Comment


                • #9
                  I hv tried htseq-count, really hard.
                  But it cannt be installed on our lab server.
                  Thank you all the same^^

                  Comment

                  Latest Articles

                  Collapse

                  • seqadmin
                    Non-Coding RNA Research and Technologies
                    by seqadmin


                    Non-coding RNAs (ncRNAs) do not code for proteins but play important roles in numerous cellular processes including gene silencing, developmental pathways, and more. There are numerous types including microRNA (miRNA), long ncRNA (lncRNA), circular RNA (circRNA), and more. In this article, we discuss innovative ncRNA research and explore recent technological advancements that improve the study of ncRNAs.

                    [Article Coming Soon!]...
                    Yesterday, 08:07 AM
                  • seqadmin
                    Recent Developments in Metagenomics
                    by seqadmin





                    Metagenomics has improved the way researchers study microorganisms across diverse environments. Historically, studying microorganisms relied on culturing them in the lab, a method that limits the investigation of many species since most are unculturable1. Metagenomics overcomes these issues by allowing the study of microorganisms regardless of their ability to be cultured or the environments they inhabit. Over time, the field has evolved, especially with the advent...
                    09-23-2024, 06:35 AM
                  • seqadmin
                    Understanding Genetic Influence on Infectious Disease
                    by seqadmin




                    During the COVID-19 pandemic, scientists observed that while some individuals experienced severe illness when infected with SARS-CoV-2, others were barely affected. These disparities left researchers and clinicians wondering what causes the wide variations in response to viral infections and what role genetics plays.

                    Jean-Laurent Casanova, M.D., Ph.D., Professor at Rockefeller University, is a leading expert in this crossover between genetics and infectious...
                    09-09-2024, 10:59 AM

                  ad_right_rmr

                  Collapse

                  News

                  Collapse

                  Topics Statistics Last Post
                  Started by seqadmin, 10-02-2024, 04:51 AM
                  0 responses
                  14 views
                  0 likes
                  Last Post seqadmin  
                  Started by seqadmin, 10-01-2024, 07:10 AM
                  0 responses
                  25 views
                  0 likes
                  Last Post seqadmin  
                  Started by seqadmin, 09-30-2024, 08:33 AM
                  1 response
                  31 views
                  0 likes
                  Last Post EmiTom
                  by EmiTom
                   
                  Started by seqadmin, 09-26-2024, 12:57 PM
                  0 responses
                  20 views
                  0 likes
                  Last Post seqadmin  
                  Working...
                  X