Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • etwatson
    Member
    • Jun 2012
    • 18

    #1

    BWA insert size distribution too large

    I'm trying to align pe reads to our reference genome (Tribolium), which has a couple thousand unmapped contigs. The reference has 10 chromosomes , each around 10 - 30 Mb in size, while the rest of the contigs are around 5kb in size.

    I ran this command on the reference with unknown contigs:
    bwa mem -k 5 -t 10 -Ma Ref.fa reads_1.fastq reads_2.fastq > map.sam

    [M::main_mem] read 1164592 sequences (100000019 bp)...
    [M::mem_pestat] # candidate unique pairs for (FF, FR, RF, RR): (23, 23, 12, 19)
    [M::mem_pestat] analyzing insert size distribution for orientation FF...
    [M::mem_pestat] (25, 50, 75) percentile: (2763, 6434, 8256)
    [M::mem_pestat] low and high boundaries for computing mean and std.dev: (1, 19242)
    [M::mem_pestat] mean and std.dev: (5558.52, 3049.04)
    [M::mem_pestat] low and high boundaries for proper pairs: (1, 24735)
    and without the unknown contains:
    bwa mem -k 5 -t 10 -Ma Ref_noUnknown.fa reads_1.fastq reads_2.fastq > map_noUnknown.sam

    [M::main_mem] read 1164592 sequences (100000019 bp)...
    [M::mem_pestat] # candidate unique pairs for (FF, FR, RF, RR): (64, 83, 49, 61)
    [M::mem_pestat] analyzing insert size distribution for orientation FF...
    [M::mem_pestat] (25, 50, 75) percentile: (69, 198, 546)
    [M::mem_pestat] low and high boundaries for computing mean and std.dev: (1, 1500)
    [M::mem_pestat] mean and std.dev: (244.98, 238.31)
    [M::mem_pestat] low and high boundaries for proper pairs: (1, 1977)
    Why is the insert size distribution different between these alignments? How can the insert size distribution be so large for the alignment that contains 2000+ contigs?
  • etwatson
    Member
    • Jun 2012
    • 18

    #2
    Problem solved.

    I ran fastq_quality_trimmer on my raw reads, and the output PE files did not match up, causing large inferred insert sizes.

    Comment

    Latest Articles

    Collapse

    • SEQadmin2
      How Immunogenomics Decodes Immunity’s Genetic Blueprint
      by SEQadmin2




      The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

      This convergence of genetics, immunology, and computation...
      Today, 05:41 AM
    • SEQadmin2
      Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
      by SEQadmin2



      CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

      Despite this, “CRISPR helped turn genome editing from a specialized technique into
      ...
      07-31-2026, 11:01 AM

    ad_right_rmr

    Collapse

    News

    Collapse

    Topics Statistics Last Post
    Started by SEQadmin2, 08-24-2026, 10:32 AM
    0 responses
    42 views
    0 reactions
    Last Post SEQadmin2  
    Started by SEQadmin2, 08-20-2026, 11:17 AM
    0 responses
    48 views
    0 reactions
    Last Post SEQadmin2  
    Started by SEQadmin2, 08-18-2026, 10:05 AM
    0 responses
    55 views
    0 reactions
    Last Post SEQadmin2  
    Started by SEQadmin2, 08-13-2026, 12:22 PM
    0 responses
    50 views
    0 reactions
    Last Post SEQadmin2  
    Working...