Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • hbt
    Member
    • Jan 2011
    • 20

    #1

    Bowtie and Tophat

    We are currently trying to analyse Solid RNA-seq data using tophat and/or bowtie.
    With default settings tophat is able to align about 35% of reads, whereas bowtie on default is able to align 53%!!
    I'm assuming that tophat first runs bowtie then tried to align reads that span exon splice junctions, how is it therefore aligning less reads that bowtie alone!?
    With some parameter adjustments (i.e. --best flag) we are able to get up to 70% reads mapped using bowtie alone but in doing this we are not able to map the splice junction reads and the --best flag cannot be set for bowtie when running it through tophat.
    Does anyone know if it is possible to get tophat to output only the splice juntion reads (i.e. everything bowtie would not have aligned) so they can be added to the bowtie output, or is there a fairly simple way of altering the bowtie parameters when running it within tophat?

    As you might have guessed, I am only just begining my journey into the worderful world of bioinformatics so go easy on me!

    Any help or advise would be greatly received!!

    Huw
  • hlwright
    Member
    • Feb 2011
    • 30

    #2
    We also can't understand why tophat is less able to map reads than bowtie despite it using bowtie for the initial mapping?

    Comment

    • adumitri
      Member
      • Jan 2010
      • 27

      #3
      Hi,

      Hopefully, this is the right place to post my question:

      I ran TopHat v1.2.0 on single reads (40 ntds) generated with Illumina GA IIx. These were the used options:

      tophat --max-multihits 2\
      --segment-mismatches 2\
      --library-type fr-unstranded\
      -p 4\
      -o sample_tophat_out\
      hg19 sample.fastq

      After getting the results, I tried to collect some statistics about the runs. Interestingly, in the generated log folder for each of the samples, there are two files that contain statistics data such as:

      ==> fileq5okX8.log <==
      # reads processed: 28183908
      # reads with at least one reported alignment: 1081639 (3.84%)
      # reads that failed to align: 26764052 (94.96%)
      # reads with alignments suppressed due to -m: 338217 (1.20%)
      Reported 1219857 alignments to 1 output stream(s)

      ==> fileueW7Yw.log <==
      # reads processed: 28183908
      # reads with at least one reported alignment: 20657663 (73.30%)
      # reads that failed to align: 3984728 (14.14%)
      # reads with alignments suppressed due to -m: 3541517 (12.57%)
      Reported 23408344 alignments to 1 output stream(s)

      There is no documentation on TopHat's website on what these two files actually represent. As you can see, the summaries look pretty different - I am not sure why there are two files in the first place, nor do I understand why there is such a large difference between the % of aligned reads.

      There are also two additional files that seem to contain relevant data: reports.log and prep_reads.log. Does anyone know what the results presented in all these files are?

      Thank you so much!
      Alexandra

      Comment

      • ttnguyen
        Member
        • Mar 2010
        • 41

        #4
        Looks like the files in log folder are not much informative. It is quite easy to collect statistics about the mapping by looking at 'accepted_hits.bam'.

        Comment

        • sterding
          Member
          • Sep 2010
          • 36

          #5
          This post helps a lot to the question here:

          http://biostar.stackexchange.com/que...ings-in-tophat

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
            by SEQadmin2



            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

            Despite this, “CRISPR helped turn genome editing from a specialized technique into
            ...
            07-31-2026, 11:01 AM
          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, 08-06-2026, 07:41 AM
          0 responses
          15 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 08-03-2026, 10:13 AM
          0 responses
          31 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-31-2026, 02:55 AM
          0 responses
          41 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-24-2026, 12:17 PM
          0 responses
          26 views
          0 reactions
          Last Post SEQadmin2  
          Working...