Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • exo
    Member
    • Dec 2012
    • 26

    #1

    Counting specific sequence stretches in fastq files

    Hey all,

    For a while ive been trying to find a way how to count specific sequence stretches directly from fastq files. For example i would like to see how many reads contain the following sequence: "TGACCCGATAGA". It seems to work with grep -c when i use fasta files but not fastq files.

    Is there anything im missing?

    Best,
    Exo
  • dpryan
    Devon Ryan
    • Jul 2011
    • 3478

    #2
    You can still use grep -c, but pipe the file through awk first:

    Code:
    awk '{if(NR%4 == 2) print $0}' sample.fastq | grep -c ...

    Comment

    • Brian Bushnell
      Super Moderator
      • Jan 2014
      • 2709

      #3
      This is a perfect application for BBDuk. I would not recommend grep for this approach with multiline fasta as it does not distinguish between lines and reads, and cannot handle newlines in the middle of the target sequence. And I have no idea why grep would fail with fastq files; that should not be a problem. But anyway -

      bbduk.sh in=reads.fq literal=TGACCCGATAGA k=12 mm=f rcomp=f int=f

      That will tell you exactly how many reads contain that sequence. If you are also interested in reads containing the reverse-complement of the sequence, remove the flag "rcomp=f". The "int=f" flag tells it to treat reads individually rather than as pairs, if your reads are interleaved. k is the kmer length, which in this case is the length of your sequence, since you want an exact match. Incidentally, BBDuk can also tell you how many reads contain a sequence that is 1 substitution away from the target with the flag "hdist=1", and so forth. It can also handle degenerate symbols like 'N' if you add the flag "copyundefined", which is occasionally useful when dealing with amplification primers.
      Last edited by Brian Bushnell; 06-12-2016, 09:19 AM.

      Comment

      • rhinoceros
        Senior Member
        • Apr 2013
        • 372

        #4
        Originally posted by dpryan View Post
        You can still use grep -c, but pipe the file through awk first:

        Code:
        awk '{if(NR%4 == 2) print $0}' sample.fastq | grep -c ...
        Returns at max 1 hit per sequence though..
        savetherhino.org

        Comment

        • dpryan
          Devon Ryan
          • Jul 2011
          • 3478

          #5
          Originally posted by rhinoceros View Post
          Returns at max 1 hit per sequence though..
          Yes, though that's typically what's desired.

          Comment

          • dariober
            Senior Member
            • May 2010
            • 311

            #6
            Originally posted by rhinoceros View Post
            Returns at max 1 hit per sequence though..
            Hi- Is that a problem? The OP asks for the number of reads containing pattern. So if a read contains the pattern twice it should still be counted as 1, which is what grep -c does, isn't it?

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              New Genomics Technologies Take Aim at Long-Standing Limits
              by SEQadmin2


              Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

              We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
              ...
              Yesterday, 10:25 AM
            • SEQadmin2
              How Immunogenomics Decodes Immunity’s Genetic Blueprint
              by SEQadmin2




              The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

              This convergence of genetics, immunology, and computation...
              09-01-2026, 05:41 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Today, 09:51 AM
            0 responses
            9 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-25-2026, 09:06 AM
            0 responses
            32 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-23-2026, 11:05 AM
            0 responses
            27 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 09-18-2026, 11:37 AM
            1 response
            47 views
            0 reactions
            Last Post pekgio
            by pekgio
             
            Working...