Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • mtiwaridros
    Member
    • Feb 2014
    • 20

    #1

    Candidate gene selection

    Hi

    My query is generic in nature and pertains to the downstream analysis of RNA-seq results.
    Usually, most differential expression (DE) analyses (of RNA-seq data) with different tools (edgeR/RSEM/DESeq etc) results in a list of differentially expressed genes (with logFC and corresponding FDR's).
    To be able to impart biological significance to such a list, one ususally identifies some candidates on which to work in vivo or in vitro.
    Enrichment with GO terms is the ususal approach that is employed to refine such a list.
    My concerns are two-fold:
    1. I have a long list of DE genes (~ 2500), and am having difficulty refining it (to 5-10 candidates).
    2. The 'post-DE-list' selection with a pre-determined target (GO terms) shall not yield novel genes that might otherwise be important, but are DE.
    Can anyone kindly suggest ways to get a pre-refined list of DE genes by specifying mapping parameters or without the need for GO selection?

    Any ideas/suggestions are appreciated.

    Thanks.
    Best
    M
  • mbblack
    Senior Member
    • Aug 2009
    • 245

    #2
    How exactly did you define the final list of genes that were considered to be "differentially expressed" in your selection of genes for enrichment? You will get by far your most robust list by simultaneously applying a statistical threshold (commonly, for example, FDR<0.05) and a magnitude threshold (e.g. {Log2FC>+1 OR Log2FC<-1}.

    If you select DGE gene lists based solely on a statistical threshold, or solely on a magnitude threshold, those lists will typically have fewer genes that will validate independently (e.g. say, by qPCR) than a list derived from both thresholds applied simultaneously.

    I usually use a threshold setting of (FDR<0.05 AND [log2ratio>+0.5849625 OR Log2ratio<-0.5849625]). Genes passing that threshold are my Differentially Expressed Genes, regardless of what is in that list. That list can then be used for enrichment to see which groups within the list appear to have some common functional role.

    But your list of DGE is whatever it is, based on your thresholds for selecting it - that determines what was or was not differentially expressed. Then you can look to see if there is any functional signal in that list, and whether that functional signal has any relevance to your particular study.

    I would never go into an enrichment analysis already excluding any ontology category outright, as that pre-supposes you will see enrichment for that selected category, and a hypergeometric analysis of lists is dependent on what is in the actually initial lists. So, you should define your list of differentially expressed genes, use them for your various enrichment analyses, and only then sort those enrichment results to see if your target ontology element(s) was actually enriched.

    When you find the ontology element you are interested in or expected to see, then you can look to see which of your differentially expressed genes actually contributed to the significant enrichment of that list. Those then would become your target genes for further work.
    Michael Black, Ph.D.
    ScitoVation LLC. RTP, N.C.

    Comment

    • mtiwaridros
      Member
      • Feb 2014
      • 20

      #3
      Hi Michael

      Thanks for your reply.
      My route to get DE genes has been: Paired-end reads>Bowtie>HTSeq>DESeq.
      On the normalized data, I have simultaneously applied the statistical (FDR<0.05) and magnitude thresholds (FC>2x).
      Subsequently, I get around 2200 DE genes.

      My original query meant to prune-down this list of DE genes, as most of the normal enrichment tools that I have used (DAVID, STRING etc) either do not work with such a large list or give vague results. Consequently, I am not able to figure out the functional relevance from this list.

      Can you suggest some relevant (1) GSEA methods, and/or (2) ways to prune down the original list (pre-/post-mapping).

      Thanks.
      Best
      M

      Comment

      • mbblack
        Senior Member
        • Aug 2009
        • 245

        #4
        With just a big list from a single condition or contrast, there are the obvious things you can do which is to increase your threshold stringency. Try an FDR<0.025 or FDR<0.01, or increase the minimum fold change that you applied. All such thresholds are subjective anyhow, so you are free to push or relax stringency as you wish, as long as you can defend the choice you end up using. Certainly, if increasing the statistical or the magnitude of change threshold yields informative results, nobody will criticize you for using higher stringency cutoffs.

        You can also rank order the list you have now by fold change and pick the top 5% or 10% from the up and down regulated ends of that rank ordered list. Or use any other creative means to sub-sample that list by some rational process.

        GSEA is another option, as long is there is a suitable gene set to compare to - or you can create your own if there is nothing already available that you think a suitable gene set to compare your list with.
        Michael Black, Ph.D.
        ScitoVation LLC. RTP, N.C.

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
          by SEQadmin2


          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

          The systematic characterization of the human proteome has
          ...
          07-20-2026, 11:48 AM
        • SEQadmin2
          Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
          by SEQadmin2



          Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
          ...
          07-09-2026, 11:10 AM
        • SEQadmin2
          Cancer Drug Resistance: The Lingering Barrier to Rising Survival
          by SEQadmin2



          Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

          There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
          07-08-2026, 05:17 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, 07-24-2026, 12:17 PM
        0 responses
        10 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-23-2026, 11:41 AM
        0 responses
        11 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-20-2026, 11:10 AM
        0 responses
        23 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-13-2026, 10:26 AM
        0 responses
        37 views
        0 reactions
        Last Post SEQadmin2  
        Working...