Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • pbarros
    Junior Member
    • Jul 2012
    • 7

    #1

    StringTie - isoforms not overlapping

    Hi,
    I am using StringTie for transcriptome reconstruction and identification of new isoforms.
    While I was exploring the "[file]-transcripts.gtf" output file from stringtie I found something that intrigued me... in the example below I show three isoforms resulting from the same gene ("STRG.14686") and the last one was present in the reference annotation. However the start and end coordinates do not match. The first two isoforms end at 394180 bp and 389422 bp, respectively, while the third starts at 396257 bp...


    scaffold_96 StringTie transcript 383404 394180 1000 + .gene_id "STRG.14686"; transcript_id "STRG.14686.1"; cov "72.829010"; FPKM "6.506798"; TPM "8.058224";
    scaffold_96 StringTie transcript 383404 389422 1000 + .gene_id "STRG.14686"; transcript_id "STRG.14686.2"; cov "61.675678"; FPKM "5.510321"; TPM "6.824155";
    scaffold_96 StringTie transcript 396257 398001 1000 + .gene_id "STRG.14686"; transcript_id "STRG.14686.3"; reference_id "scaffold_96.g39603.t1"; ref_gene_id "scaffold_96.g39603"; cov "2963.938721"; FPKM "264.808624"; TPM "327.947357";
    Why is StringTie "clustering" these isoforms in the same gene?
    Last edited by pbarros; 04-26-2017, 03:22 AM.
  • sdriscoll
    I like code
    • Sep 2009
    • 436

    #2
    first of all let me say that I agree with your interpretation. the third transcript does not overlap the first two and it has given all three the same 'gene_id' value.

    the only thing that comes to mind is that the assembly from stringtie, or cufflinks for that matter, is an attempted explanation, and often a simplification, of the alignment data. you may learn more about this by looking at the alignments in an alignment browser such as UCSC or IGV.

    while i doubt this, stringtie could be very, very smart and have assembled the third isoform from reads that multimapped between it and the first two transcripts which would imply they are all the same gene but one that is repeated in more than one position.
    /* Shawn Driscoll, Gene Expression Laboratory, Pfaff
    Salk Institute for Biological Studies, La Jolla, CA, USA */

    Comment

    • pbarros
      Junior Member
      • Jul 2012
      • 7

      #3
      thank you for the input sdriscoll ... maybe I was overthinking this

      cheers,
      pedro

      Comment

      Latest Articles

      Collapse

      • SEQadmin2
        Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
        by SEQadmin2



        CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

        Despite this, “CRISPR helped turn genome editing from a specialized technique into
        ...
        Today, 11:01 AM
      • SEQadmin2
        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
        by SEQadmin2


        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

        The systematic characterization of the human proteome has
        ...
        07-20-2026, 11:48 AM
      • SEQadmin2
        Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
        by SEQadmin2



        Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
        ...
        07-09-2026, 11:10 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by SEQadmin2, Today, 02:55 AM
      0 responses
      7 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-24-2026, 12:17 PM
      0 responses
      12 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-23-2026, 11:41 AM
      0 responses
      12 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-20-2026, 11:10 AM
      0 responses
      24 views
      0 reactions
      Last Post SEQadmin2  
      Working...