Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • north_zeb
    Junior Member
    • Sep 2012
    • 9

    #1

    bowtie paired-end versus single-end

    Hi guys,

    I'm totally new to NGS and have 2 fastq files corresponding to a paired-end illumina chipseq experiment. I am using bowtie to align the fastq : first i align only one fastq file then both of them. The % of reads with at least 1 reported alignment is so different for each case: 83% when i used only 1 fastq compared to ONLY 47% when i used both of them (as i should have). Anyone can explain me why ?

    here are the results:


    /bowtie -t -p 4 --sam --chunkmbs 1000 hg19/hg19 reads/V400can_chipseq_1.fastq > results/v400.sam

    # reads processed: 15749589
    # reads with at least one reported alignment: 13044920 (82.83%)
    # reads that failed to align: 2704669 (17.17%)
    Reported 13044920 alignments to 1 output stream(s)
    Time searching: 00:32:23

    AS OPPOSED TO:

    ./bowtie -t -p 4 --sam --chunkmbs 1000 -m 1 hg19/hg19 -1 reads/V400can_chipseq_1.fastq -2 reads/V400can_chipseq_2.fastq > results/v400_paired.sam

    # reads processed: 15749589
    # reads with at least one reported alignment: 7813517 (49.61%)
    # reads that failed to align: 7480588 (47.50%)
    # reads with alignments suppressed due to -m: 455484 (2.89%)
    Reported 7813517 paired-end alignments to 1 output stream(s)
    Time searching: 01:44:03


    Thanks a mil in advance,

    NZ
  • swbarnes2
    Senior Member
    • May 2008
    • 910

    #2
    I don't use Bowtie much, but there's a setting for expected insert size, and I think Bowtie behaves very badly with pairs that are too far from that insert size. Crank up the maximum insert size, and try again

    Comment

    • biznatch
      Senior Member
      • Nov 2010
      • 124

      #3
      Originally posted by swbarnes2 View Post
      I don't use Bowtie much, but there's a setting for expected insert size, and I think Bowtie behaves very badly with pairs that are too far from that insert size. Crank up the maximum insert size, and try again
      Max insert size -X, default is only 250 (ie. total size including both reads + insert has to be 250 or less). I set mine at 1000 so nothing should be excluded.

      Comment

      • swbarnes2
        Senior Member
        • May 2008
        • 910

        #4
        Did you try inspecting reads that mapped the first time, and not the second?

        Again, I don't use bowtie, but the first time you used 1 fastq, and the second time you used 2, is it normal for the # of total reads to be the same?

        Also, rather than

        > results/v400_paired.sam
        consider

        | samtools view -bSh - > v400_paired.bam
        You can convert a subset of that file to .sam later to eyeball things.

        Comment

        • north_zeb
          Junior Member
          • Sep 2012
          • 9

          #5
          Thanks for all replies !
          i will have a look at the insert size story. True, the total nb of reads is given the same in both cases, and this is the output of bowtie. Can it be that bowtie counts each pair when reports on paired-end reads ?

          Comment

          • JackieBadger
            Senior Member
            • Mar 2009
            • 385

            #6
            What are you aligning to, a full genome or genomic scaffolds?
            It makes sense that if you map PE data to scaffolds (which are not a continuous fragment) then a lot of sequences will fail to map if your insert size causes them to fall off the end of the fragment that the first PE maps to.

            If you do not care about your insert size i.e. not trying to re-sequence large regions of the genome, and have genomic scaffolds I would concatenate the PEs and map in single end mode

            Comment

            • jbrwn
              Member
              • Mar 2011
              • 37

              #7
              honestly, that seems about right. bowtie2 made improvements to paired-end, so you may want to check that out. paired-end specific options: http://bowtie-bio.sourceforge.net/bo...ed-end-options

              Comment

              • north_zeb
                Junior Member
                • Sep 2012
                • 9

                #8
                I align against indexed hg19 downloaded from the bowtie website. i'll read that link. look at the beginning of the sam file bowtie gives me. Does anybody know what the 0 in the insert size position mean ?

                SRR424618.6 HWIUSI-EAS523_0001:5:1:999:17802 77 * 0 0 * * 0 0 NGGCTTTAGTCAAAGTACAGAAGACATTAGAAGAAAATTGCAGAAACAGGCTGGGTTTGCANGCATGAATNCGNCA #''''52)+.88633AAAAAAAAAAAA7AA7AAA7A72A8AAAAAA7AA########################### XM:i:1
                SRR424618.6 HWIUSI-EAS523_0001:5:1:999:17802 141 * 0 0 * * 0 0 NCAAACACCTGGTTGGCTATCTCCAATAACTGTGACGTATTCATGCCTGCAAACCCAGCNNNNNNNNNCANNNNNC #***('**+'::4:20*523AAA7AAAAAA############################################## XM:i:1
                SRR424618.10 99 chr20 42794368 255 76M = 42794395 103 NATGGAACCACCTCAGGGCCTTGGTATTGCTGTTCCCTCTACCTGTAATGCCCTTCCTCCAGATACCTACNTGGCT #'**'0.0..AAAAA8AA77::85:AAAAA############################################## XA:i:1 MD:Z:0C69A5 NM:i:2
                SRR424618.10 147 chr20 42794395 255 76M = 42794368 -103 TNNNNNTCNNNNNNNNNGTAATGCCCTTCCTCCAGATACCTACATGGCTCACCCTCTTGCCGTCTTCAAGCCTTTN ############################################################################ XA:i:1 MD:Z:1G0C0T0G0T2C0C0T0C0T0A0C0C0T58A0 NM:i:15
                SRR424618.9 163 chr13 99753904 255 76M = 99753933 105 NAGACCAGCCGGAGCAACAAAAAATTAGCTAGGCATGGTGGTGCATGCCAGTGGTCCCANNNNNNNNNGANNNNNG #''**00222AAAAAAAAAAA27*7626667AAAA######################################### XA:i:1 MD:Z:0G58G0C0T0A0C0T0T0T0G2G0G0G0T0G0A0 NM:i:16
                SRR424618.9 83 chr13 99753933 255 76M = 99753904 -105 TAGGCNTGGTGGTGCATGCCAGTGGTCCCAGCTACTTTGGAGGGTGAGATGTGAAGATCCCCTGAGCCCAGGAGTN ##################AAAA7AAA896:820*+*7AAAAAAAAAAAAAAAAAA8AAAAAAAAAA20.),*'*'# XA:i:1 MD:Z:5A69T0 NM:i:2

                Comment

                • swbarnes2
                  Senior Member
                  • May 2008
                  • 910

                  #9
                  Did you look at the binary flags? 77 means that neither read of the pair mapped.

                  141 means the same thing. Notice how neither has a mapping position either? the quality turns to junk in the end, that might be part of the problem.

                  Comment

                  • north_zeb
                    Junior Member
                    • Sep 2012
                    • 9

                    #10
                    oh, thanks for that actually, i have started to figure out some of the flags numbers but these are new to me. If i align only the first of the fast files , with the -m 1 option, it gives: reads with at least 1 alignment: 70,66%
                    The second fast file gives 59.07% reads with at least 1 alignment.

                    Comment

                    Latest Articles

                    Collapse

                    • SEQadmin2
                      Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                      by SEQadmin2



                      CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                      Despite this, “CRISPR helped turn genome editing from a specialized technique into
                      ...
                      07-31-2026, 11:01 AM
                    • SEQadmin2
                      Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                      by SEQadmin2


                      Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                      The systematic characterization of the human proteome has
                      ...
                      07-20-2026, 11:48 AM

                    ad_right_rmr

                    Collapse

                    News

                    Collapse

                    Topics Statistics Last Post
                    Started by SEQadmin2, 08-06-2026, 07:41 AM
                    0 responses
                    18 views
                    0 reactions
                    Last Post SEQadmin2  
                    Started by SEQadmin2, 08-03-2026, 10:13 AM
                    0 responses
                    33 views
                    0 reactions
                    Last Post SEQadmin2  
                    Started by SEQadmin2, 07-31-2026, 02:55 AM
                    0 responses
                    43 views
                    0 reactions
                    Last Post SEQadmin2  
                    Started by SEQadmin2, 07-24-2026, 12:17 PM
                    0 responses
                    26 views
                    0 reactions
                    Last Post SEQadmin2  
                    Working...