Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts

  • vkartha
    replied
    Hello,
    I am working with RNASEq data and we are to use 2 sets of 3 bam files each, i.e. 2 different conditions and 3 replicates per condition (which were obtained using Tophat2) with Cufflinks, Cuffmerge and finally Cuffdiff.

    Cuffmerge requires a txt file with the full paths to the (in this case 6) transcripts.gtf files which was obtained on running cufflinks on each of the bam files. It also requires a reference genome fasta file and a gtf annotation file.

    I chose to use gencode's latest version 'gencode.v12.annotation.gtf' with cuffmerge. I however am getting the following error, and I don't have an answer for it.


    [Wed Jul 4 00:53:50 2012] Preparing output location ./merged_asm/
    [Wed Jul 4 00:55:46 2012] Converting GTF files to SAM
    [00:57:34] Loading reference annotation.
    Error: duplicate GFF ID 'ENST00000361547.2' encountered!
    [FAILED]
    Error: could not execute gtf_to_sam
    Traceback (most recent call last):
    File "/share/apps/assembly/bin/cuffmerge", line 576, in ?
    sys.exit(main())
    File "/share/apps/assembly/bin/cuffmerge", line 554, in main
    sam_input_files = convert_gtf_to_sam(gtf_input_files)
    File "/share/apps/assembly/bin/cuffmerge", line 287, in convert_gtf_to_sam
    sam_out = gtf_to_sam(line)
    File "/share/apps/assembly/bin/cuffmerge", line 247, in gtf_to_sam
    exit(1)
    TypeError: 'str' object is not callable


    Does anyone know what this means, or why it is happening? Can it be fixed? Any help would be greatly appreciated. Thanks

    - Vinay

    Leave a comment:


  • manianslab
    replied
    Okay, here is my attempt at a qucik perlscript to do this (see attachment). Hope some perl expert can make it less clumsy!
    Attached Files

    Leave a comment:


  • manianslab
    replied
    Unfortunately, cuffdiff does not seem to carry over the original transcript ids from the merged.gtf file in to the output files. However, you can either use Excel (vlookup) or a perlscript or awk to populate an additional field in your 'diff' output files by searching merged.gtf. The original transcript ids are listed as 'oId' in the merged.gtf created by cuffmerge.

    I will post a script if I find time to get one done.

    hope this helps!

    Leave a comment:


  • andremho
    replied
    Hi, does anyone have an answer to this question?

    I am fairly new to this game, but I am running a small RNA-seq study on cancer samples.
    I have mapped the reads using tophat, and quantified transcript/gene expression using cufflinks and the -g option with ENSEMBL annotation .gtf file.

    However, I also don't find many of the expected expressed gene IDs in my genes.fpkm_tracking file, like TP53 etc. But when i do a manual inspection of the assembly using IGV, I clearly see that these genes are highly covered by reads....

    Also, in my isoforms.fpkm_tracking file I can find some of the transcript Ids for these genes, but as i said not their gene IDs..

    Can anyone explain this/give a solution to my problem?

    Leave a comment:


  • goudurix
    replied
    Hi all,
    Epi, did you find a workaround ? I ran into the same problem. Running cufflinks with -g options, some original gene_ids were converted to new identifiers sharing the CUFF prefix.
    Il2ra (NM_008367) became CUFF.265243...
    Regards

    Leave a comment:


  • epi
    replied
    Thanks for commenting here, it is good to know I am not the only one in the this boat. I have pinged Cole again, hope he writes something here.

    Leave a comment:


  • AsoBioInfo
    replied
    Hi, I also encountered the same issue. The ./cufflinks also gave me some *.CUFF id's.
    I used -g option, as I am interested in novel transcripts. How can you get all novel transcripts or genes?

    Leave a comment:


  • Lee Sam
    replied
    Has anyone figured this issue out? I'm running a small study with cancer samples and I don't even see TP53 annotated in the genes.fpkm_tracking using either the Illumina-provided UCSC or Ensembl files (with chromosome names edited for compatibility). I'm running using the "-g" option as well and I was expecting that existing genes observed by cufflinks would have entries with the annotated gene names.
    Last edited by Lee Sam; 04-05-2012, 10:05 AM.

    Leave a comment:


  • epi
    replied
    Originally posted by Cole Trapnell View Post
    Hmm, that's annoying... can you show us how you're running cuffmerge? One limitation of the current version of Cuffmerge is that it doesn't do any of its own ORF prediction. So new isoforms of coding genes won't get assigned a new p_id or attached to an existing one. New genes will not be declared as coding or non-coding (this is a hard problem in general). However, they will get grouped under a tss_id. The thing is, Cuffmerge will reassign tss_ids just like it reassigns transcript_ids, because those also have to be unique. The only thing it really tries to propogate from the reference in terms of metadata is the gene short name.
    Hi Cole, Thanks again for replying. While your comment clarifies some things, it also raises some more question. First here is my command

    Code:
    cuffmerge -o test1 -g genes.gtf -p 10 -s ~/path/mm9 assemblies1.txt
    Assemblies1.txt contains absolute path to cufflinks gtf
    Code:
    path-to-dir/transcripts_sample1.gtf
    path-to-dir/transcripts_sample2.gtf
    .
    .
    .
    The behavior you described is what I am observing as well. As I mentioned earlier, I am using -g and not -G to allow known as well as novel genes. Should I not expect it to preserve transcript, TSS and pids for the known genes.

    Also I am still unclear why there are new predictions (with CUFF. IDs) that are identical to existing annotations. I have manually looked at one of the test runs and all new predictions (with CUFF. IDs) were original genes/transcripts in the source GTF annotation (I must have looked only about 20-25 ). So far I have not encountered any novel predictions which do not correspond to existing annotation. In some cases, original annotation still exists and in some cases not. Similar to the example about Trib3 I posted above.

    Leave a comment:


  • Cole Trapnell
    replied
    Originally posted by epi View Post
    Actually, apart form a specific case, the real question here is how come cufflinks is not using the transcript_ID, p_ID and in some cases gene_ID from the original GFF as it is supposed to.
    Hmm, that's annoying... can you show us how you're running cuffmerge? One limitation of the current version of Cuffmerge is that it doesn't do any of its own ORF prediction. So new isoforms of coding genes won't get assigned a new p_id or attached to an existing one. New genes will not be declared as coding or non-coding (this is a hard problem in general). However, they will get grouped under a tss_id. The thing is, Cuffmerge will reassign tss_ids just like it reassigns transcript_ids, because those also have to be unique. The only thing it really tries to propogate from the reference in terms of metadata is the gene short name.

    Leave a comment:


  • epi
    replied
    Thanks you all for replies, and I apologize if I was panicking more than needed. Perhaps I should explain my question a bit further.

    As I understand, -g adds new transcripts "in addition to" existing ones, which is my goal under this analysis.

    From the manual:
    Output will include all reference transcripts as well as any novel genes and isoforms that are assembled
    While for many other genes, it is using the original gene name, for Trib3, it is converting it to CUFF.9451, and after cuffmerge to CUFF.7092.2. This is the source of all confusion, from manual and published protocol it is not clear why it should do that. After all, the annotation of Trib3 transcript and its exons is identical to CUFF.9451. They why this behaves like a novel gene now.

    From jcgrenier comments, I went back and found that it is indeed correct that transcript_id are incremented in ascending order, which has made a chance identity of IDs between chr2 and chrY, could very well be other cases which I don't know of.

    Actually, apart form a specific case, the real question here is how come cufflinks is not using the transcript_ID, p_ID and in some cases gene_ID from the original GFF as it is supposed to.
    Thanks again for comments.
    Last edited by epi; 03-12-2012, 08:02 AM.

    Leave a comment:


  • Luyi Tian
    replied
    There is a nature protocol published several days before, which may help you:
    "Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks"

    Leave a comment:


  • Cole Trapnell
    replied
    As far as I can tell, Cufflinks is working correctly and as described in the manual. As noted above, you want -G, not -g.

    Leave a comment:


  • jcgrenier
    replied
    For the tss_id, there is probably a bug, because it's numerated in ascending order but the TSS number as nothing to do with the gene itself.


    ***Update***
    Cuffmerge merge together multiple transcripts.gtf files coming from different analyses. The TSS information is not present in that file. Maybe it doesn't use the information contained in the given reference GTF file and just rename the tss_id...
    Last edited by jcgrenier; 03-09-2012, 12:24 PM.

    Leave a comment:


  • jcgrenier
    replied
    Hi,

    The option "-g" that you are using is usually used to discover new transcripts. So the CUFF.* names that you got there are for those newly found transcripts.
    If you only want to do it on known transcripts, use the "-G" option.

    That is, I don't know why CuffMerge gives us certain oId with CUFF.* names.

    Leave a comment:

Latest Articles

Collapse

  • SEQadmin2
    New Genomics Technologies Take Aim at Long-Standing Limits
    by SEQadmin2


    Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

    We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
    ...
    09-28-2026, 10:25 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by SEQadmin2, 09-29-2026, 09:51 AM
0 responses
43 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-25-2026, 09:06 AM
0 responses
52 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-23-2026, 11:05 AM
0 responses
43 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-18-2026, 11:37 AM
1 response
51 views
0 reactions
Last Post pekgio
by pekgio
 
Working...