Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • exo
    Member
    • Dec 2012
    • 26

    Counting specific sequence stretches in fastq files

    Hey all,

    For a while ive been trying to find a way how to count specific sequence stretches directly from fastq files. For example i would like to see how many reads contain the following sequence: "TGACCCGATAGA". It seems to work with grep -c when i use fasta files but not fastq files.

    Is there anything im missing?

    Best,
    Exo
  • dpryan
    Devon Ryan
    • Jul 2011
    • 3478

    #2
    You can still use grep -c, but pipe the file through awk first:

    Code:
    awk '{if(NR%4 == 2) print $0}' sample.fastq | grep -c ...

    Comment

    • Brian Bushnell
      Super Moderator
      • Jan 2014
      • 2709

      #3
      This is a perfect application for BBDuk. I would not recommend grep for this approach with multiline fasta as it does not distinguish between lines and reads, and cannot handle newlines in the middle of the target sequence. And I have no idea why grep would fail with fastq files; that should not be a problem. But anyway -

      bbduk.sh in=reads.fq literal=TGACCCGATAGA k=12 mm=f rcomp=f int=f

      That will tell you exactly how many reads contain that sequence. If you are also interested in reads containing the reverse-complement of the sequence, remove the flag "rcomp=f". The "int=f" flag tells it to treat reads individually rather than as pairs, if your reads are interleaved. k is the kmer length, which in this case is the length of your sequence, since you want an exact match. Incidentally, BBDuk can also tell you how many reads contain a sequence that is 1 substitution away from the target with the flag "hdist=1", and so forth. It can also handle degenerate symbols like 'N' if you add the flag "copyundefined", which is occasionally useful when dealing with amplification primers.
      Last edited by Brian Bushnell; 06-12-2016, 09:19 AM.

      Comment

      • rhinoceros
        Senior Member
        • Apr 2013
        • 372

        #4
        Originally posted by dpryan View Post
        You can still use grep -c, but pipe the file through awk first:

        Code:
        awk '{if(NR%4 == 2) print $0}' sample.fastq | grep -c ...
        Returns at max 1 hit per sequence though..
        savetherhino.org

        Comment

        • dpryan
          Devon Ryan
          • Jul 2011
          • 3478

          #5
          Originally posted by rhinoceros View Post
          Returns at max 1 hit per sequence though..
          Yes, though that's typically what's desired.

          Comment

          • dariober
            Senior Member
            • May 2010
            • 311

            #6
            Originally posted by rhinoceros View Post
            Returns at max 1 hit per sequence though..
            Hi- Is that a problem? The OP asks for the number of reads containing pattern. So if a read contains the pattern twice it should still be counted as 1, which is what grep -c does, isn't it?

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM
            • SEQadmin2
              Cancer Drug Resistance: The Lingering Barrier to Rising Survival
              by SEQadmin2



              Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

              There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
              07-08-2026, 05:17 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, 07-20-2026, 11:10 AM
            0 responses
            13 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-13-2026, 10:26 AM
            0 responses
            31 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-09-2026, 10:04 AM
            0 responses
            42 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-08-2026, 10:08 AM
            0 responses
            27 views
            0 reactions
            Last Post SEQadmin2  
            Working...