Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • puggie
    Member
    • Nov 2011
    • 52

    #1

    BWA PE Problem

    Dear forum,

    I am experiencing problems during BWA paired-end mapping on my local setup. I have two fastq files for each sample. BWA aln step seems okay, but then at sampe the alignment may fail after outputting 1 Mb or even 1 Gb of SAM data. Could this happen to be a pre-processing issue? Fastq --> rename --> trimming ?

    BWA is executed under Galaxy platform and I dont see the log file, so I may try to run it on command line. I see it fails at sampe via htop command.

    I have also indexed the reference.
  • swbarnes2
    Senior Member
    • May 2008
    • 910

    #2
    Pasting the output when it goes bad might help.

    Is it possible that the fastqs got corrupted part-way through?

    Renaming and trimming shouldn't hurt things, as long as the millionth read in fastq 1 really is the mate of the millionth read of fastq 2, and so on.

    Comment

    • puggie
      Member
      • Nov 2011
      • 52

      #3
      With the latter scenario I would have expected BWA to "hang" for several hours/days? Im currently running an alignment with newly transferred fastqs, and will report back if that doesnt solve the issue.

      Comment

      • swbarnes2
        Senior Member
        • May 2008
        • 910

        #4
        What does bwa sample tell you are the approximate insert sizes? If they are crazy huge, it can take the software a long time to decide that, and move on.

        But no, it should not hang. It should be outputting its progress every couple of seconds.

        Comment

        • Kennels
          Senior Member
          • Feb 2011
          • 149

          #5
          As mentioned above, you need to make sure that your processed paired end reads are paired correctly in read 1 and read 2 files. I had BWA stall forever when my reads weren't paired properly after preprocessing (filtering by quality will get rid of different reads in read 1 and read 2 file if you don't use the appropriate software or re-pair after processing).

          Comment

          • puggie
            Member
            • Nov 2011
            • 52

            #6
            Ok thanks for your replies. I tried with new fastq files, and still got the problem, only one sample passed through to the complete SAM file. I also tried mapping without any preprocessing, and same error BWA stops at sampe.

            I want to run it in command line outside of Galaxy, but Im not super-familiar with linux. Right now Im indexing the genome. Should all indexing files be in the same directory as my input fastqs and reference? Or can I specify the location somehow? I just installed bwa by apt-get on a clean Ubuntu disk image. But where exactly does it install? Only thing I know, is that bwa is in the PATH.

            And my apologies for such basic linux questions.

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
              by SEQadmin2



              CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

              Despite this, “CRISPR helped turn genome editing from a specialized technique into
              ...
              07-31-2026, 11:01 AM
            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Yesterday, 10:13 AM
            0 responses
            14 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-31-2026, 02:55 AM
            0 responses
            27 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-24-2026, 12:17 PM
            0 responses
            20 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-23-2026, 11:41 AM
            0 responses
            19 views
            0 reactions
            Last Post SEQadmin2  
            Working...