Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • chimb
    Junior Member
    • Aug 2014
    • 2

    #1

    Tuxedo Pipeline Issue - Multiple Gene Hits per Transcript, High FPKM

    Hello all!

    ** I already posted this in the Bioinformatics forum -- I'm not sure which one it should belong to-- my apologies. Admins, feel free to delete/merge my post as necessary **

    - I've been analyzing some RNA-seq data using the Tuxedo pipeline and have been getting some peculiar results, which are especially noticeable in the tables of significant genes (and their differential expression data) I've attached.

    - Some biological background: the experiment is looking at bacteria-bacteria interaction effects between Streptococcus sanguinis (Ss) and Porphyromonas gingivalis (Pg). There are numerous conditions and comparisons that were made using Cuffdiff, but the data I've attached is based on the comparison between the conditions:

    Wild-type Ss (SK36) grown in isolation (sample_1)
    --vs--
    Wild-type Ss cultured with wild-type Pg (sample_2)

    In this case, the cuffdiff run utilizes the Ss read alignments and uses the merged transcriptome of Ss across both conditions.

    - In Sk--Sk_Pg_sig_genes.txt, I ran the data through the whole Tuxedo Pipeline using Trapnell et. al's protocol from Nature. Tophat, Cufflinks, Cuffmerge, Cuffdiff, cummeRbund -- all default commands/options. In cummeRbund, I used the getSig(), getGenes(), diffData() and featureNames() functions to merge together a table of the significantly diff-expressed genes (alpha=0.05), their differential expression data and their short names. Two peculiar things:

    - Some transcripts report hits with multiple genes each (many gene_short_name's per transcript)

    - FPKM (value_1 and value_2) are extremely high for some transcripts ~ 3089410 for one of them, which can't be possible.


    - My PI and I suspected that tophat may be finding splice junctions that do not exist (I did not include "--no-novel-juncs" in my initial tophat runs). I imagine this would link together disparate stretches of DNA as a single transcript and garner multiple gene hits. That, or perhaps many genes overlapping across the same stretches of DNA in different reading frames (though I'd imagine cufflinks would account for that?).

    - I tried running the whole pipeline again, but skipped the tophat step (which includes read fragmentation and splice junction discovery). I ran bowtie2 alone for the bare-read alignments, converted the output SAM to BAM, sorted it and fed it through cufflinks and the rest of the pipeline as normal. The result is (using the same extraction methods in cummeRbund): sig_genes_Sk-Sk_Pg_bt2.txt

    ~ Still, getting multiple gene hits per transcript.. and still getting extremely high FPKM values

    **************************************************

    - Have any of you experienced the same sort of problems? What might be causing this? Any suggestions for alternate methods for alignment, transcript construction or visualization? ... I realize the Tuxedo pipeline was designed with eukaryotic systems in mind so I'm not sure if it is, in whole or in part, unsuitable for prokaryotes.

    Any input would be greatly appreciated!

    Thanks!

Latest Articles

Collapse

  • SEQadmin2
    Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
    by SEQadmin2



    CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

    Despite this, “CRISPR helped turn genome editing from a specialized technique into
    ...
    Yesterday, 11:01 AM
  • SEQadmin2
    Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
    by SEQadmin2


    Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

    The systematic characterization of the human proteome has
    ...
    07-20-2026, 11:48 AM
  • SEQadmin2
    Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
    by SEQadmin2



    Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
    ...
    07-09-2026, 11:10 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by SEQadmin2, Yesterday, 02:55 AM
0 responses
9 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 07-24-2026, 12:17 PM
0 responses
12 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 07-23-2026, 11:41 AM
0 responses
12 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 07-20-2026, 11:10 AM
0 responses
24 views
0 reactions
Last Post SEQadmin2  
Working...