Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • adrian
    Member
    • Oct 2009
    • 90

    #1

    shorter time for building bowtie index

    Hi:
    I am not sure if this question has been asked before.

    I have 120 RNA-Seq fasta files. I am using tophat2 for aligning to human genome.

    I am queing this job on a 70 node cluster. I am using bowtie genome files downloaded from Illumina genome website.

    Homo_sapiens/Ensembl/GRCh37/Sequence/Bowtie2Index/genome

    I see for every fastq file alignment, tophat building index file from genes.fa..

    It takes more than 30 minutes for every file.

    Can this time consuming process be avoided since we already have index files for bowtie.

    Following is my command line:

    tophat2 -p 16 \
    -G /Refs/Homo_sapiens/Ensembl/GRCh37/Annotation/Genes/genes.gtf \
    -o tpht \
    /Refs/Homo_sapiens/Ensembl/GRCh37/Sequence/Bowtie2Index/genome \

    myfile_1.fastq myfile_2.fastq

    Thanks
  • blancha
    Senior Member
    • May 2013
    • 367

    #2
    I don't know the answer to your question. Probably not.

    You can however get faster TopHat runs simply by adding the option --no-novel-juncs.

    Comment

    • gringer
      David Eccles (gringer)
      • May 2011
      • 845

      #3
      Yes, in the most recent version(s) of tophat, you can pre-generate the indexes with the 'transcriptome-index' option:

      Comment

      • blancha
        Senior Member
        • May 2013
        • 367

        #4
        @gringer

        Very interesting, and useful. Thanks for the information.

        Comment

        • adrian
          Member
          • Oct 2009
          • 90

          #5
          Thank you so much! that is helpful.

          However when I use the following I don't see files - known.gff, known.fa, known.fa.tlst, known.fa.ver and the known.* Bowtie index files in the directory.

          following is my command:

          tophat2 -p 16 \
          -G /Refs/Homo_sapiens/Ensembl/GRCh37/Annotation/Genes/genes.gtf \
          -o tpht \
          /Refs/Homo_sapiens/Ensembl/GRCh37/Sequence/Bowtie2Index/genome \
          --transcriptome-index=known
          myfile_1.fastq myfile_2.fastq

          Is there something wrong. Where would the indexed known files will be deposited?
          thanks
          Adrian

          Comment

          • GenoMax
            Senior Member
            • Feb 2008
            • 7142

            #6
            To prepare transcriptome indexes you should not include the -o tpht or the sequence files. This is a special one time tophat run (Check the example in the link gringer included)

            Your command would be:
            Code:
            $ tophat2  \
            -G /Refs/Homo_sapiens/Ensembl/GRCh37/Annotation/Genes/genes.gtf \
            --transcriptome-index=IF_you_want_a_directory_name/known \
            /Refs/Homo_sapiens/Ensembl/GRCh37/Sequence/Bowtie2Index/genome
            This needs to be a non-threaded run so you have to omit the -p 16 as well.
            Last edited by GenoMax; 06-30-2014, 09:59 AM. Reason: Corrected error in the genome_base_name

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
              by SEQadmin2



              CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

              Despite this, “CRISPR helped turn genome editing from a specialized technique into
              ...
              07-31-2026, 11:01 AM
            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Today, 07:41 AM
            0 responses
            9 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 08-03-2026, 10:13 AM
            0 responses
            24 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-31-2026, 02:55 AM
            0 responses
            38 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-24-2026, 12:17 PM
            0 responses
            25 views
            0 reactions
            Last Post SEQadmin2  
            Working...