Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • apredeus
    Senior Member
    • Jul 2012
    • 151

    #1

    Classify genes as expressed or not expressed

    Hello all,

    this is probably a very obvious question, but I've never dealt with this sort of a problem, so I hope you all can point me in the right direction.

    Imagine we have an array or annotated and quantified RNA-seq experiment. There are about ~24k genes, with normalized numerical expression value (or FPKM) assigned to them.

    What is the most statistically sound way to automatically classify genes as "expressed" and "not expressed"? People often use empirical cutoff for this, e.g. FPKM of 1, but that's not what I'm interested in.

    Thank you for any inputs.
  • cariboudoug
    Member
    • Oct 2013
    • 17

    #2
    I was wondering this as well... I have multi-species/ multi-individual RNA-seq data so its easy to bin genes into on and off if they are expressed highly in all individuals of 1 species and very little in another... The difficulty is that many genes might be expressed stochastically and at low levels although still functional. If both species have low FPKM then I have trouble distinguishing mapping errors and actual low expression.

    I was thinking of just choosing a cut-off based on the distribution of FPKMs that makes sense to me

    Comment

    • Bukowski
      Senior Member
      • Jan 2010
      • 388

      #3
      How can you ever say something is 'not expressed'? Absence of evidence is not evidence of absence, and this is especially true with RNA-Seq data, because it's a sampling technique. Unless you're doing ridiculously deep sequencing on your samples any cut off you put in place is pretty arbitrary.

      Comment

      • apredeus
        Senior Member
        • Jul 2012
        • 151

        #4
        Yes, I see your point. Furthermore, if I remember the ENCODE papers correctly, they came to the conclusion that good fraction of mRNAs are present in some cells of the same type, and absent in others. Thus we would only see certain average, which could be quite low..

        On the other hand, especially for microarrays, there is quite big (numerical) difference between something obviously expressed, and something that's not. From the practical standpoint (I need it for gene expression clustering) it would have been useful to restrict the gene array to only significantly expressed ones. How would one go about it? I do have some ideas, but I'm curious what do others think.

        Comment

        • bruce01
          Senior Member
          • Mar 2011
          • 160

          #5
          "Absence of evidence is not evidence of absence , and this is especially true with RNA-Seq data, because it's a sampling technique"

          I agree wholeheartedly, but there is still a distribution for that sample you take, and I like to think that taking a cut-off from that distribution is 'less arbitrary' than the "FPKM<1" route.

          I put up a script on github after being asked about a method I use in a seminar. If you want to have a look I would appreciate ideas about this. A workable solution, I thought, though I am frequently wrong!

          Comment

          • krapulaxdoctor
            Member
            • May 2015
            • 22

            #6
            Hi, I hope to revive this thread, because I need some advice:
            I got a debate with colleagues about the very same question: which gene is expressed or not. When I mentioned that I used an FPKM cutoff to select genes I further analyzed, they demanded statistics to prove if my selected genes are significant.
            Could someone propose a statistical approach to show whether my genes are expressed and if that is "significant"?

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
              by SEQadmin2



              CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

              Despite this, “CRISPR helped turn genome editing from a specialized technique into
              ...
              07-31-2026, 11:01 AM
            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, 08-06-2026, 07:41 AM
            0 responses
            20 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 08-03-2026, 10:13 AM
            0 responses
            33 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-31-2026, 02:55 AM
            0 responses
            43 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-24-2026, 12:17 PM
            0 responses
            26 views
            0 reactions
            Last Post SEQadmin2  
            Working...