Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • mtiwaridros
    Member
    • Feb 2014
    • 20

    #1

    Candidate gene selection

    Hi

    My query is generic in nature and pertains to the downstream analysis of RNA-seq results.
    Usually, most differential expression (DE) analyses (of RNA-seq data) with different tools (edgeR/RSEM/DESeq etc) results in a list of differentially expressed genes (with logFC and corresponding FDR's).
    To be able to impart biological significance to such a list, one ususally identifies some candidates on which to work in vivo or in vitro.
    Enrichment with GO terms is the ususal approach that is employed to refine such a list.
    My concerns are two-fold:
    1. I have a long list of DE genes (~ 2500), and am having difficulty refining it (to 5-10 candidates).
    2. The 'post-DE-list' selection with a pre-determined target (GO terms) shall not yield novel genes that might otherwise be important, but are DE.
    Can anyone kindly suggest ways to get a pre-refined list of DE genes by specifying mapping parameters or without the need for GO selection?

    Any ideas/suggestions are appreciated.

    Thanks.
    Best
    M
  • mbblack
    Senior Member
    • Aug 2009
    • 245

    #2
    How exactly did you define the final list of genes that were considered to be "differentially expressed" in your selection of genes for enrichment? You will get by far your most robust list by simultaneously applying a statistical threshold (commonly, for example, FDR<0.05) and a magnitude threshold (e.g. {Log2FC>+1 OR Log2FC<-1}.

    If you select DGE gene lists based solely on a statistical threshold, or solely on a magnitude threshold, those lists will typically have fewer genes that will validate independently (e.g. say, by qPCR) than a list derived from both thresholds applied simultaneously.

    I usually use a threshold setting of (FDR<0.05 AND [log2ratio>+0.5849625 OR Log2ratio<-0.5849625]). Genes passing that threshold are my Differentially Expressed Genes, regardless of what is in that list. That list can then be used for enrichment to see which groups within the list appear to have some common functional role.

    But your list of DGE is whatever it is, based on your thresholds for selecting it - that determines what was or was not differentially expressed. Then you can look to see if there is any functional signal in that list, and whether that functional signal has any relevance to your particular study.

    I would never go into an enrichment analysis already excluding any ontology category outright, as that pre-supposes you will see enrichment for that selected category, and a hypergeometric analysis of lists is dependent on what is in the actually initial lists. So, you should define your list of differentially expressed genes, use them for your various enrichment analyses, and only then sort those enrichment results to see if your target ontology element(s) was actually enriched.

    When you find the ontology element you are interested in or expected to see, then you can look to see which of your differentially expressed genes actually contributed to the significant enrichment of that list. Those then would become your target genes for further work.
    Michael Black, Ph.D.
    ScitoVation LLC. RTP, N.C.

    Comment

    • mtiwaridros
      Member
      • Feb 2014
      • 20

      #3
      Hi Michael

      Thanks for your reply.
      My route to get DE genes has been: Paired-end reads>Bowtie>HTSeq>DESeq.
      On the normalized data, I have simultaneously applied the statistical (FDR<0.05) and magnitude thresholds (FC>2x).
      Subsequently, I get around 2200 DE genes.

      My original query meant to prune-down this list of DE genes, as most of the normal enrichment tools that I have used (DAVID, STRING etc) either do not work with such a large list or give vague results. Consequently, I am not able to figure out the functional relevance from this list.

      Can you suggest some relevant (1) GSEA methods, and/or (2) ways to prune down the original list (pre-/post-mapping).

      Thanks.
      Best
      M

      Comment

      • mbblack
        Senior Member
        • Aug 2009
        • 245

        #4
        With just a big list from a single condition or contrast, there are the obvious things you can do which is to increase your threshold stringency. Try an FDR<0.025 or FDR<0.01, or increase the minimum fold change that you applied. All such thresholds are subjective anyhow, so you are free to push or relax stringency as you wish, as long as you can defend the choice you end up using. Certainly, if increasing the statistical or the magnitude of change threshold yields informative results, nobody will criticize you for using higher stringency cutoffs.

        You can also rank order the list you have now by fold change and pick the top 5% or 10% from the up and down regulated ends of that rank ordered list. Or use any other creative means to sub-sample that list by some rational process.

        GSEA is another option, as long is there is a suitable gene set to compare to - or you can create your own if there is nothing already available that you think a suitable gene set to compare your list with.
        Michael Black, Ph.D.
        ScitoVation LLC. RTP, N.C.

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
          by SEQadmin2



          CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

          Despite this, “CRISPR helped turn genome editing from a specialized technique into
          ...
          07-31-2026, 11:01 AM
        • SEQadmin2
          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
          by SEQadmin2


          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

          The systematic characterization of the human proteome has
          ...
          07-20-2026, 11:48 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, 08-11-2026, 10:35 AM
        0 responses
        10 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 08-06-2026, 07:41 AM
        0 responses
        29 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 08-03-2026, 10:13 AM
        0 responses
        48 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-31-2026, 02:55 AM
        0 responses
        48 views
        0 reactions
        Last Post SEQadmin2  
        Working...