Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • vanillasky
    Member
    • Mar 2014
    • 42

    Blastp for identifying genes in contigs

    I have used NCBI blastp as a first step to identify predicted ORFs in my assembled contigs. My question is how to filter these results based on % identity of the protein and e-value scores. I picked a loose protein cut-off of 20% but I'm not sure how one goes about selecting the best filtering options.
    Any input would be appreciated. Thnx.
  • maubp
    Peter (Biopython etc)
    • Jul 2009
    • 1544

    #2
    I would output the BLAST results in tabular format, and then filter using a simple script. Tabular output is quite easy to work with. You could even do this in Excel if you really prefer.

    Comment

    • vanillasky
      Member
      • Mar 2014
      • 42

      #3
      Cut-off values

      Thank you for your suggestion. However, my question is a fundamental one. How does one decide on a cut-off value (s) with which to filter out the blasp results? For example based on greatest % identity for a protein match or an e-value score?

      Comment

      • rhinoceros
        Senior Member
        • Apr 2013
        • 372

        #4
        Check the discussion here.
        savetherhino.org

        Comment

        • maubp
          Peter (Biopython etc)
          • Jul 2009
          • 1544

          #5
          I don't have access to the paper right now (I'm away from the office), but I think reading this would be useful:

          Punta and Ofran (2008) The Rough Guide to In Silico Function Prediction,
          or How To Use Sequence and Structure Information To Predict Protein
          Function. PLoS Comput Biol 4(10): e1000160.
          Protein structure prediction,Protein structure,Protein structure comparison,Protein structure databases,Sequence motif analysis,Structural genomics,Protein domains,Sequence alignment

          Comment

          • GenoMax
            Senior Member
            • Feb 2008
            • 7142

            #6
            @vanillasky: Many things informaticians do can be automated/done in high throughput mode but once you get to "annotation" it may be best to slow down and do things the right way (that is why "finishing" a genome takes so long).

            You could choose a value ("n") for the % cut-off and get a good approximation of the identities of proteins in your dataset. Afterwards, only a fraction (more or less, would depend on how good your blast results were) of those identities may turn out to be correct on more rigorous examination (as an example http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2673347/).

            In general you can be reasonably sure about the validity of a protein blast if the e value is < 10^-3 and sequence identity is >= 25%.

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              Today, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM
            • SEQadmin2
              Cancer Drug Resistance: The Lingering Barrier to Rising Survival
              by SEQadmin2



              Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

              There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
              07-08-2026, 05:17 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Today, 11:10 AM
            0 responses
            7 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-13-2026, 10:26 AM
            0 responses
            29 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-09-2026, 10:04 AM
            0 responses
            38 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-08-2026, 10:08 AM
            0 responses
            25 views
            0 reactions
            Last Post SEQadmin2  
            Working...