Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • imsharmanitin
    Postdoc Cancer Bioinformatics
    • Dec 2014
    • 17

    #1

    transcriptome_index Bowtie2 Index

    Dear all,

    As far as understand Tophat2 creates an transcriptome_index which also includes Bowtie2 Index files using GTF and the genome file. Then why do we need to supply Bowtie2 Index to create transcriptome index for the first time.

    I tried to create transcriptome index without supplying bowtie2 index files and it gave following error.

    tophat2 -G ../data/Homo_sapiens.GRCh37.75.gtf --transcriptome-index ./transcriptome_index/ ../data/Homo_sapiens.GRCh37.75.dna.primary_assembly.fa

    [2015-06-30 12:25:22] Building transcriptome files with TopHat v2.0.13
    -----------------------------------------------
    [2015-06-30 12:25:22] Checking for Bowtie
    Bowtie version: 2.2.4.0
    [2015-06-30 12:25:24] Checking for Bowtie index files (genome)..
    Error: Could not find Bowtie 2 index files (../data/Homo_sapiens.GRCh37.75.dna.primary_assembly.fa.*.bt2)



    However, after that I created bowtie2 index files and ran same process and no error was produced.
  • GenoMax
    Senior Member
    • Feb 2008
    • 7142

    #2
    Tophat2 uses the entire human genome index files and then creates a "transcriptome-only" subset using information in the GTF file. You only need to do this once, as you discovered.

    Comment

    • imsharmanitin
      Postdoc Cancer Bioinformatics
      • Dec 2014
      • 17

      #3
      Originally posted by GenoMax View Post
      Tophat2 uses the entire human genome index files and then creates a "transcriptome-only" subset using information in the GTF file. You only need to do this once, as you discovered.

      so this will be the workflow:

      1) generate whole genome index using bowtie2
      Homo_sapiens.GRCh37.75.dna.primary_assembly.fa
      Homo_sapiens.GRCh37.75.gtf


      2) the using tophat 2 to create a "transcriptome-only" subset using the whole genome index flie created in step 1

      3) point the tophat 2 to the directory containing "transcriptome-only" subset using

      --transcriptome-index <directory containing "transcriptome-only" subset>

      for subsequent runs

      Comment

      • GenoMax
        Senior Member
        • Feb 2008
        • 7142

        #4
        That is correct. When you are using --transcriptome-only option consider additional options (e.g. -T) that become relevant (scroll down to find the --transcriptome-only option section): https://ccb.jhu.edu/software/tophat/manual.shtml#toph.

        Comment

        • imsharmanitin
          Postdoc Cancer Bioinformatics
          • Dec 2014
          • 17

          #5
          Originally posted by GenoMax View Post
          That is correct. When you are using --transcriptome-only option consider additional options (e.g. -T) that become relevant (scroll down to find the --transcriptome-only option section): https://ccb.jhu.edu/software/tophat/manual.shtml#toph.
          If i understood correctly, when we use the --transcriptome-index without -T then the reads will be first mapped to the transcriptome-index and the reads which fail to match will be mapped to genome (just like using only -G option)

          on the other hand if i use -T then mapping will happen only with transcriptome and report only those mappings as genomic mappings.

          if i want to map to genome only then use of -G and -T should be avoided.

          Also as far i have read and understood mapping to both transcriptome and genome will be an ideal approach i.e. use of -T should be avoided. Am I right?

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            New Genomics Technologies Take Aim at Long-Standing Limits
            by SEQadmin2


            Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

            We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
            ...
            Yesterday, 10:25 AM
          • SEQadmin2
            How Immunogenomics Decodes Immunity’s Genetic Blueprint
            by SEQadmin2




            The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

            This convergence of genetics, immunology, and computation...
            09-01-2026, 05:41 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, Today, 09:51 AM
          0 responses
          8 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 09-25-2026, 09:06 AM
          0 responses
          31 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 09-23-2026, 11:05 AM
          0 responses
          27 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 09-18-2026, 11:37 AM
          1 response
          47 views
          0 reactions
          Last Post pekgio
          by pekgio
           
          Working...