Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • adrian
    Member
    • Oct 2009
    • 90

    #1

    RNA-Seq sequence counts fastq and BAM files

    Hi:
    I am a bit confused about total number of reads obtained from RNA-Seq run.

    In the case of a paired end run fastq should both R1 and R2 reads be counted to get total number of reads.

    Example:

    Sample-S was run on two lanes as 2x100 PE reads configuration.

    Sample-S - R1 in lane 1 - 30Mil reads (as per FASTQ file)
    Sample-S - R2 in lane 1 - 30 Mil reads

    Sample-S -R1 in "lane 2" - 35Mil reads
    Sample-S -R2 in "lane 2" - 35Mil reads.


    If I want to know total # of reads I sequenced for Sample-S

    Is it 130Mil. reads or 65Mil reads?

    Question 2:
    I see difference close to double between FASTQ file and reads from samtools flagstat total reads. Why is this - is this because in a paired end BAM file both R1 and R2 reads are mapped and counted.
    In this case should one count both R1 and R2 reads in fastq file.

    Appreciate your help.

    Adrian
  • dpryan
    Devon Ryan
    • Jul 2011
    • 3478

    #2
    1) I would say "65 million read pairs", rather than "130 million reads", since the latter is always a bit ambiguous.
    2) samtools flagstat counts both reads in a pair, so you should expect to see double. The reason for this is to allow people to mix paired and single-end reads and still get meaningful metrics. Note also that there are separate metrics printed specifically for read1 and read2 in a pair (these numbers should more closely the numbers from the fastq files).

    Comment

    • GenoMax
      Senior Member
      • Feb 2008
      • 7142

      #3
      CASAVA reports the stats as "X million reads" through in reality it is "X/2" M read pairs per sample per lane for a paired-end run.

      In similar vein, one terabase of sequence from a HiSeq 2500 counts output of *two* flowcells from one instrument.

      Comment

      • bruce01
        Senior Member
        • Mar 2011
        • 160

        #4
        Originally posted by adrian View Post
        Hi:
        In the case of a paired end run fastq should both R1 and R2 reads be counted to get total number of reads.
        When you make the PE library there is one 'insert' or sequence between two primers, say R1 and R2. Each primer can hybridise to the flowcell. One end (R1) hybridises first. The read1 sequence is then 100bp into the insert from that direction. With PE sequencing the R2 end is then hybridised to the flowcell and read2 sequence is 100bp into the insert form the other direction. Thus it is really only a single sequence with an unknown stretch in between, (often, confusingly, called an insert) and so should be counted as such, as dpryan says.

        As a toy example, if you had a 10bp PE library sequenced, with original insert of 30bp:

        R1: ACTGACTGAC----------ACTGACTGAC :R2

        This is also more likely to align to a single position in the transcriptome, which is why it is a good sequencing strategy.

        Comment

        • adrian
          Member
          • Oct 2009
          • 90

          #5
          Thanks for replies. I got it.

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
            by SEQadmin2



            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

            Despite this, “CRISPR helped turn genome editing from a specialized technique into
            ...
            07-31-2026, 11:01 AM
          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, 08-03-2026, 10:13 AM
          0 responses
          20 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-31-2026, 02:55 AM
          0 responses
          34 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-24-2026, 12:17 PM
          0 responses
          24 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-23-2026, 11:41 AM
          0 responses
          21 views
          0 reactions
          Last Post SEQadmin2  
          Working...