Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • reventropy
    Junior Member
    • Apr 2014
    • 7

    #1

    Cuffcompare with a single experiment

    I am running RNA-seq analysis on a paired-end deep sequencing data set with no replicates. We are interested in finding novel gene and transcript isoforms in addition to variant info. Grooming and Tophat alignment went well and I’ve processed the .bam output through cufflinks in RABT mode with –GTF-guide. I then take the .gtf output from this and run cuffcompare with the reference .gtf and .fasta.

    I am experiencing confusion related to the last step and was hoping that somebody with more experience than I could help to clarify a few things.

    Firstly, most of the references I have read regading cuffcompare indicate that it is used for multiple replicates or experiments: “Used to Track Cufflinks transcripts across multiple experiments (e.g. across a time course)”. Is it common to use cuffcompare on a single experiment in order to find novel isoforms?

    Secondly, there are some entries in the output from cuffcompare that aren’t making sense to me. What does it mean when I see an "=" class code with a zero FMI? How about a "j" class code with a FMI of 100? Based on the definition of FMI (fraction of major isoform), these scenarios don't seem possible.

    Thirdly, if I want an fpkm score for a known gene, is it common to sum all transcript fpkms belonging to that gene with an "=" class code?







    Thanks so much for any help, and let me know if I can/should provide more information!



    -Jeremy
  • mikep
    Member
    • Feb 2011
    • 45

    #2
    Originally posted by reventropy View Post
    Firstly, most of the references I have read regading cuffcompare indicate that it is used for multiple replicates or experiments: “Used to Track Cufflinks transcripts across multiple experiments (e.g. across a time course)”. Is it common to use cuffcompare on a single experiment in order to find novel isoforms?
    Depends on your definition of "common". There's no technical reason you can't (I certainly have). Usually people use the cuffcompare output as the guide file for cuffdiff. The former gives you the union set of transcripts, the latter then looks for differential expression in those transcripts.

    Secondly, there are some entries in the output from cuffcompare that aren’t making sense to me. What does it mean when I see an "=" class code with a zero FMI? How about a "j" class code with a FMI of 100? Based on the definition of FMI (fraction of major isoform), these scenarios don't seem possible.
    cuffcompare outputs all the transcripts it finds, or is told are real (exist in the guide file). "=" transcripts exist in the guide file, so are output even if there's no support for their existence. It's not clear why you think a j class transcript cannot have an FMI of 100.


    Thirdly, if I want an fpkm score for a known gene, is it common to sum all transcript fpkms belonging to that gene with an "=" class code?
    Summing fpkms is fine, but you should include novel transcripts, or rerun cufflinks without novel transcript finding.

    Comment

    • reventropy
      Junior Member
      • Apr 2014
      • 7

      #3
      Thanks a lot mikep!

      It's not clear why you think a j class transcript cannot have an FMI of 100.
      This is probably owing to my flawed reasoning.

      I was operating under the assumption that major isoforms come from the annotation file and cannot be novel. If I see an FMI of 100 and a "j" class code then should I assume that Cufflinks identified the man isoform as being novel, i.e., a novel gene?

      Thanks again for addressing my questions so that I can proceed with more confidence.

      -Jeremy

      Comment

      • mikep
        Member
        • Feb 2011
        • 45

        #4
        Originally posted by reventropy View Post
        If I see an FMI of 100 and a "j" class code then should I assume that Cufflinks identified the man isoform as being novel, i.e., a novel gene?

        -Jeremy
        Your interpretation is correct.

        Thanks again for addressing my questions so that I can proceed with more confidence.
        I would be very careful being confident in novel isoforms from cufflinks, it has a pretty high error rate. You haven't mentioned which organism you are working with but if it is human or one of the model organisms you might be better off with using just the existing annotation. If it is a few genes you care about I'd load the cufflinks output & BAM file into a genome viewer and have a look at the actual reads.

        Comment

        • reventropy
          Junior Member
          • Apr 2014
          • 7

          #5
          I would be very careful being confident in novel isoforms from cufflinks, it has a pretty high error rate. You haven't mentioned which organism you are working with but if it is human or one of the model organisms you might be better off with using just the existing annotation. If it is a few genes you care about I'd load the cufflinks output & BAM file into a genome viewer and have a look at the actual reads.
          I'll definitely keep that in the front of my mind. The sequencing is human. I have been using IGV, but am still training my eye. We're only interested in coding genes so I will be filtering the cuffdiff output, but we would like to catch any novel transcripts or gene isoforms in this subset.

          -Jeremy

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
            by SEQadmin2



            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

            Despite this, “CRISPR helped turn genome editing from a specialized technique into
            ...
            07-31-2026, 11:01 AM
          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, Yesterday, 10:13 AM
          0 responses
          14 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-31-2026, 02:55 AM
          0 responses
          29 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-24-2026, 12:17 PM
          0 responses
          22 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-23-2026, 11:41 AM
          0 responses
          21 views
          0 reactions
          Last Post SEQadmin2  
          Working...