Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • poisson200
    Member
    • Feb 2010
    • 63

    #1

    To cuffcompare or not to cuffcompare

    Dear Bioinformaticians,
    The experiment; RNA-seq, transcriptomics; we simply want to find differentially expressed genes between 2 Illumina samples, case vs. control. No novel transcripts needed, just want to find differentially expressed known genes and/or known transcript isoforms.

    Do I make the right assumption that there is no need to use cufflinks or cuffcompare?
    E.g. Cuffcompare does;
    1) Compare your assembled transcripts to a reference annotation
    2) Track Cufflinks transcripts across multiple experiments
    ????

    Our proposed approach to finding known differentially expressed genes/isoforms:
    1) Run Tophat with a Refseq GTF file on both case and control
    2) Run cuffdiff on both sam files generated from 1 and the GTF file used in 1 as well.

    Please let me know if this is or is not a valid approach?

    Thank you,

    J.
    Last edited by poisson200; 07-29-2010, 04:16 AM.
  • poisson200
    Member
    • Feb 2010
    • 63

    #2
    There are 49 views but nobody knows the answer?


    Or is this is a stupid question?

    Comment

    • Olgy
      Junior Member
      • Aug 2010
      • 1

      #3
      I would really want to help you, but you know that my knowlage about programming are very small

      Olgy Wolgy

      Comment

      • Enrico Palazzo
        Junior Member
        • Jul 2010
        • 9

        #4
        Originally posted by poisson200 View Post
        Dear Bioinformaticians,
        The experiment; RNA-seq, transcriptomics; we simply want to find differentially expressed genes between 2 Illumina samples, case vs. control. No novel transcripts needed, just want to find differentially expressed known genes and/or known transcript isoforms.

        Do I make the right assumption that there is no need to use cufflinks or cuffcompare?
        E.g. Cuffcompare does;
        1) Compare your assembled transcripts to a reference annotation
        2) Track Cufflinks transcripts across multiple experiments
        ????

        Our proposed approach to finding known differentially expressed genes/isoforms:
        1) Run Tophat with a Refseq GTF file on both case and control
        2) Run cuffdiff on both sam files generated from 1 and the GTF file used in 1 as well.

        Please let me know if this is or is not a valid approach?

        Thank you,

        J.
        Hi J,
        I'm new to this field and this is just a guess:
        I think you have to run cuffcompare on the two GTF files produced in step 2 and your reference annotation file:
        cuffcompare -r Mus_musculus.test.gtf -R -o prefix transcripts1.gtf transcripts2.gtf

        Then you run cuffdiff on the combined output GTF file (from cuffcompare) and the 2 SAM files which had been generated by TopHat in step 1:
        cuffdiff -o diff_out/ combined.gtf accepted_hits1.sam accepted_hits2.sam

        DE genes/isoforms will be in 0_1_gene_exp.diff/0_1_isoform_exp.diff
        but it seems that there is no output for coding sequences as mentioned in an other thread http://seqanswers.com/forums/showthread.php?t=4989. If you have trouble with duplicate entries in your reference annotation filles: check this: http://seqanswers.com/forums/showthread.php?t=3493

        Comment

        • poisson200
          Member
          • Feb 2010
          • 63

          #5
          Thanks Olgy,
          I am sure your intelligence and knowledge of other things is vast.

          ===============================

          Dear Enrico,
          thanks for your answer. I since found out this is the answer to my questions, which are (very similar to yours with the novel and known case)

          1) Known transcripts only; Find differentially expressed known Refseq transcripts between case and control data (no novel transcripts wanted);
          a. Run TopHat on both fastq read files using a GFF file of Refseq
          b. Run cuffdiff on both sam outputs, using a Refseq GTF file

          2) Known and novel transcripts; Find differentially expressed transcripts, both novel and known Refseq
          a. Run TopHat on both fastq (without GFF)
          b. Run cufflinks without GTF file on each SAM file
          c. Run cuffcompare with Refseq.GTF sam1.gtf and sam2.gtf
          d. Run cuffdiff on the combined.GTF and sam files

          Thanks for the link too.

          Kind regards,

          J

          Comment

          • ChrisL
            Member
            • Nov 2009
            • 14

            #6
            No need for Cufflinks or Cuffcompare if you are happy to stick within an existing genome annotation like RefSeq, UCSC or Ensembl. Just take the sam files from Tophat and go straight to Cuffdiff. There is a little trick to get a gtf file that is suitable to use with Cuffdiff with tss id, etc. Just use Cuffcompare and feed it the reference annotation (RefSeq, Ensembl, etc) twice and it will give you a gtf file that is suitable.

            Hope this helps.

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              New Genomics Technologies Take Aim at Long-Standing Limits
              by SEQadmin2


              Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

              We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
              ...
              Yesterday, 10:25 AM
            • SEQadmin2
              How Immunogenomics Decodes Immunity’s Genetic Blueprint
              by SEQadmin2




              The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

              This convergence of genetics, immunology, and computation...
              09-01-2026, 05:41 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Today, 09:51 AM
            0 responses
            9 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-25-2026, 09:06 AM
            0 responses
            31 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-23-2026, 11:05 AM
            0 responses
            27 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-18-2026, 11:37 AM
            1 response
            47 views
            0 reactions
            Last Post pekgio
            by pekgio
             
            Working...