Cuffset number of genes too high

aprice67

Member

Join Date: Nov 2012

Posts: 49
- Share
- Tweet
#1

Cuffset number of genes too high

05-25-2017, 11:19 AM

Hi,

So I'm implementing a pretty standard tuxedo pipeline on paired-end mouse data. I went along as follows.

Using the standard mouse build (37) from Ensemble and the associated annotation data I did the following.

cutadapt -> tophat2 -> cufflinks -> cuffmerge -> cuffquant -> cuffdiff -> cummerbund.

Tophat2 is giving me roughly 85% overall mapping, 5-10% multimap. Cufflinks was run on around 100 files and those were merged with cuffmerge. All of the bam files were then run with cuffquant using the gtf from cuffmerge.

To test out an initial dataset I just used two sample conditions (2 replicates each) to do cuffdiff. 8 files. (4 control, 4 corresponding experimental case).

Now, I load this up in cummeRbund and see the following:

CuffSet instance with:
2 samples
52615 genes
220501 isoforms
106928 TSS
49476 CDS
52615 promoters
106928 splicing
21223 relCDS

So this is my first time working with mouse or cufflinks pipeline, but somehow these numbers don't feel right. So I checked it out and at least I found that mouse has only around 23,000 genes, so thats wrong for sure.

Could someone explain to me what sort of numbers I should be seeing here and why at least cuffdiff/cummRbund is showing approximately double the number of genes that should exist in the genome I'm looking at?

I appreciate any feedback anyone can provide.
Tags: counts, cufflinks, cummerbund, tophat2, tuxedo

Previous template Next

Recent Advances in Sequencing Analysis Tools

by seqadmin

The sequencing world is rapidly changing due to declining costs, enhanced accuracies, and the advent of newer, cutting-edge instruments. Equally important to these developments are improvements in sequencing analysis, a process that converts vast amounts of raw data into a comprehensible and meaningful form. This complex task requires expertise and the right analysis tools. In this article, we highlight the progress and innovation in sequencing analysis by reviewing several of the...
- Channel: Articles
05-06-2024, 07:48 AM

Topics	Statistics	Last Post
Ancient Viral Sequences in Human Brain Linked to Psychiatric Disorders by seqadmin Started by seqadmin, Today, 07:35 AM	0 responses 5 views 0 likes	Last Post by seqadmin Today, 07:35 AM
New Milestone for COSMIC with Extensive Cancer Mutation Data by seqadmin Started by seqadmin, Yesterday, 02:06 PM	0 responses 8 views 0 likes	Last Post by seqadmin Yesterday, 02:06 PM
The Role of Spliceosomes in RNA Splicing and Genome Evolution by seqadmin Started by seqadmin, 05-14-2024, 07:03 AM	0 responses 27 views 0 likes	Last Post by seqadmin 05-14-2024, 07:03 AM
A Closer Look at the Enigmatic Genomes of Oikopleura dioica by seqadmin Started by seqadmin, 05-10-2024, 06:35 AM	0 responses 47 views 0 likes	Last Post by seqadmin 05-10-2024, 06:35 AM

Seqanswers Leaderboard Ad

Announcement

Cuffset number of genes too high

Latest Articles

ad_right_rmr

News